Libraries, methods, and kits for index color balancing nucleic acids

Artificial nucleic acid molecules with specific patterns and sequences address the challenge of signal misinterpretation in two-channel sequencing systems, improving accuracy and coverage by ensuring consistent signal detection across channels.

WO2026096864A1PCT designated stage Publication Date: 2026-05-07SEQWELL INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SEQWELL INC
Filing Date
2025-10-31
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Two-channel nucleic acid sequencing systems face challenges in accurate signal detection due to potential misinterpretation of dark signals, leading to errors in base calling, particularly in homopolymeric regions and low-complexity sequences, necessitating improved calibration methods to enhance sequencing accuracy and coverage.

Method used

The use of artificial nucleic acid molecules with specific nucleotide patterns and sequences, including an attachment region, calibration regions, and control regions, to create a color-balanced sequencing library that ensures consistent signal detection across channels, minimizing signal misinterpretation and improving sequencing accuracy.

Benefits of technology

The artificial nucleic acid molecules provide accurate color balancing, reducing errors in base calling and enhancing sequencing accuracy, especially in low-plexity libraries, by maintaining consistent signal intensities throughout the sequencing run.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000021_0001
    Figure IMGF000021_0001
  • Figure 00000029_0000
    Figure 00000029_0000
Patent Text Reader

Abstract

The invention provides compositions that include artificial nucleic acids that are useful for color balancing detected signals during a sequencing reaction, and methods of using the same such as methods of sequencing nucleic acids and methods of preparing a color balanced nucleic acid sample for improving accuracy of acquired nucleic acid sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PATENT

[0002] ATTORNEY DOCKET NO. 51178-020W02

[0003] LIBRARIES, METHODS, AND KITS FOR INDEX COLOR BALANCING NUCLEIC ACIDS

[0004] SEQUENCE LISITING

[0005] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on October 29, 2025, is named “51178-020W02_Sequence_Listing_10_29_25” and is 4,425 bytes in size.

[0006] FIELD OF THE INVENTION

[0007] The present invention relates generally to compositions including artificial nucleic acids such as color-balancing artificial nucleic acids that serve as calibration and validation controls for methods of sequencing, methods of preparing and sequencing nucleic acid samples that include the artificial nucleic acids, and kits including the same.

[0008] BACKGROUND

[0009] Nucleic acid sequencing methods (e.g., DNA sequencing methods) rely on a variety of means for detecting signals, including changes in conductivity, emission of fluorescence or radiation, and differences in mass. Sequencing-by-extension (SBE) involves template-dependent incorporation of nucleotides into a growing DNA chain by DNA polymerase. Typical sequencing chemistries rely on reversibly, fluorescently labeled nucleotides to identify the terminal base. To identify the terminal base (i.e., A, C, G, or T) present on the DNA, one chemistry combination requires four different fluorophores, one for each base. Another version only requires three different fluorophores, one for each base except one, which is known as the “dark base” because it has no fluorophore. Fewer fluorophores reduce the cost of the optical systems used to excite and detect them. Two fluorophores may also be used in detection.

[0010] However, utilizing fewer fluorophores, specifically in a two-channel detection system where only two fluorophores are used, poses significant challenges for accurate sequencing. In such systems, two dyes are used to distinguish four bases, with a binary code assigned to each nucleotide such that one fluorophore corresponds to a first base, the second fluorophore corresponds to a second base, both fluorophores correspond to a third base, and no signal (e.g., a dark signal) corresponds to a fourth base. One key challenge in this approach is that a dark signal of one base, which corresponds to the absence of a fluorescent signal, may be misinterpreted as a weak signal from a fluorophore if the detection system is not adequately calibrated. This challenge is compounded over the course of the sequencing run when signal intensities may shift due to photobleaching or internal detection fluctuations due to laser power drift or changes in the flow cell of the sequencing platform. Such challenges can reduce the accuracy and depth of the sequencing data and are particularly problematic in homopolymeric regions or in low-complexity sequences, where the lack of contrast in signals can lead to errors in base calling. Thus, there is a need for improved calibration of sequencing detection systems during data acquisition to improve sequencing accuracy and coverage. PATENT

[0011] ATTORNEY DOCKET NO. 51178-020W02

[0012] SUMMARY OF THE INVENTION

[0013] The present invention provides compositions and methods of use thereof, e.g., for nucleic acid library preparation and sequencing.

[0014] In one aspect, the invention features plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules includes from its 5’ end to its 3’ end the following:

[0015] (a) an attachment region for linking the artificial nucleic acid molecule to a surface;

[0016] (b) a first calibration region having a nucleotide sequence in a repeating pattern of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from adenine (A), thymine (T), guanine (G), or cytosine (C), and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases; and

[0017] (c) a control region, wherein the first calibration region and the control region both have a different nucleotide sequence in each artificial nucleic acid molecule.

[0018] In a second aspect, the invention features a plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules includes from its 5’ end to its 3’ end the following:

[0019] (a) an attachment region for linking the artificial nucleic acid molecule to a surface; and

[0020] (b) a calibration region having a nucleotide sequence in a repeating pattern of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from A, T, G, or C, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases; wherein the first calibration region of each artificial nucleic acid molecule has a different nucleotide sequence.

[0021] In some embodiments, each of the artificial nucleic acid molecules further includes a control region at its 3’ end, and the control region has a different nucleotide sequence in each artificial nucleic acid molecule.

[0022] In some embodiments, the attachment region includes a first priming region. In some embodiments, the first priming region includes a nucleotide sequence of 5’- AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1) or 5’- CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2).

[0023] In some embodiments, each artificial nucleic acid molecule further includes a second priming region 3’ of the control region. In some embodiments, the second priming region includes a nucleotide sequence of the reverse complement of 5’-AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1) or 5’-CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2).

[0024] In some embodiments, each artificial nucleic acid molecule further includes a first read sequence between the first calibration region and the control region. In some embodiments, the read sequence includes 5’-TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or 5’-GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4).

[0025] In some embodiments, the first calibration region is 5 to 20 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length). PATENT

[0026] ATTORNEY DOCKET NO. 51178-020W02

[0027] In some embodiments, the attachment region is 5 to 20 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length).

[0028] In some embodiments, each artificial nucleic acid molecule further includes a second calibration region 3’ of the control region, wherein the second calibration region has a nucleotide sequence in a repeating pattern of Y’X across the length of the second calibration region, wherein Y’ represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase complementary to the third nucleobase of Y, wherein the second calibration region has a different nucleotide sequence in each artificial nucleic acid molecule.

[0029] In some embodiments, the second calibration region is 5 to 20 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length).

[0030] In some embodiments, each artificial nucleic acid molecule further includes a second read sequence between the control region and the second calibration region. In some embodiments, the read sequence includes the reverse complement of 5’-TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or 5’- GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4).

[0031] In some embodiments, the artificial nucleic acid molecules are 100 to 1 ,000 nucleotides in length (100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, 400 to 450, 450 to 500, 500 to 550, 550 to 600, 600 to 650, 650 to 700, 700 to 750, 750 to 800, 800 to 850, 850 to 900, 900 to 950, or 950 to 1 ,000 nucleotides in length).

[0032] In some embodiments, the control region is a portion of an artificial sequence. In some embodiments, the control region is a portion of a genome of an organism. In some embodiments, the organism is PhiX174.

[0033] In some embodiments, the first calibration region is 10 nucleotides in length.

[0034] In some embodiments, the surface is a flow cell for nucleic acid sequencing. In some embodiments, plurality of nucleic acid molecules is linked to the surface via hybridization of the attachment region and a complementary oligonucleotide.

[0035] In some embodiments, the artificial nucleic acid molecules include DNA.

[0036] In some embodiments, the two nucleobases of X of the calibration region are A and T. In some embodiments, the third nucleobase of Y is G.

[0037] In a third aspect, the invention features a method of sequencing a plurality of nucleic acid molecules, wherein the method includes:

[0038] (a) adding to a sequencing library of nucleic acid molecules a plurality of artificial nucleic acid molecules to produce a color balanced sequencing library, wherein each of the artificial nucleic acid molecules includes from its 5’ end to its 3’ end the following:

[0039] (i) an attachment region for linking the artificial nucleic acid molecule to a surface;

[0040] (ii) a first calibration region having the order of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from A, T, G, or C, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases, wherein the first calibration region has a different nucleotide sequence in each artificial nucleic acid molecule; and PATENT

[0041] ATTORNEY DOCKET NO. 51178-020W02

[0042] (iii) a control region, wherein the control region has a different nucleotide sequence in each artificial nucleic acid molecule;

[0043] (b) applying the color balanced sequencing library to the surface; and

[0044] (c) sequencing with the color balanced sequencing library.

[0045] In a fourth aspect, the invention features method of preparing a color balanced sequencing library, the method includes adding to a sequencing library of nucleic acid molecules a plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules includes from its 5’ end to its 3’ end the following:

[0046] (i) an attachment region for linking the artificial nucleic acid molecule to a surface;

[0047] (ii) a first calibration region having the order of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from A, T, G, or C, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases; and

[0048] (iii) a control region, wherein the first calibration region and the control region both have a different nucleotide sequence in each artificial nucleic acid molecule, thereby producing a color balanced sequencing library.

[0049] In some embodiments of the fourth aspect, the method further includes sequencing the color balanced sequencing library.

[0050] In some embodiments of any one of the foregoing methods, each artificial nucleic acid molecule further includes a second calibration region 3’ of the control region, wherein the second calibration region has a nucleotide sequence in a repeating pattern of Y’X across the length of the second calibration region, wherein Y’ represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase complementary to the third nucleobase of Y, wherein the second calibration region has a different nucleotide sequence in each artificial nucleic acid molecule.

[0051] In some embodiments, the color balanced sequencing library includes the plurality of artificial nucleic acid molecules at a molar ratio of less than or equal to 20%. In some embodiments, the color balanced sequencing library includes a molar ratio of about 5%, 4.5%, 4%, 3.5%, 3%, 2.5%, 2%, 1 .5%, 1%, or 0.1% of the plurality of artificial nucleic acid molecules. In some embodiments, the color balanced sequencing library includes a molar ratio of about 3.5% of the plurality of artificial nucleic acid molecules. In some embodiments, the color balanced sequencing library includes a molar ratio of about 3% of the plurality of artificial nucleic acid molecules. In some embodiments, the color balanced sequencing library includes a molar ratio of about 1% of the plurality of artificial nucleic acid molecules.

[0052] In some embodiments, the method of sequencing includes an optical or imaging-based detection system.

[0053] In some embodiments, the detection system includes a first channel that detects a first positive signal, and a second channel that detects a second positive signal. In some embodiments, the first positive signal corresponds to one of the two nucleobases of X, and the second positive signal corresponds to the other nucleobase of the two nucleobases of X. In some embodiments, the PATENT

[0054] ATTORNEY DOCKET NO. 51178-020W02 first positive signal and the second positive signal correspond to the third nucleobase of Y. In some embodiments, the two nucleobases are A and T. In some embodiments, the third nucleobase is C.

[0055] In some embodiments, the method of sequencing includes next generation sequencing (NGS).

[0056] In a fifth aspect, the invention features a method of preparing a color balanced sequencing library, the method comprising adding to a sequencing library of nucleic acid molecules a plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules comprises from its 5’ end to its 3’ end the following:

[0057] (i) an attachment region for linking the artificial nucleic acid molecule to a surface;

[0058] (ii) a first calibration region having the order of XY across the length of the calibration region, wherein X represents one of two nucleobases detected in a single channel of a two channel imagingbased detection system, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases and is detected in both channels of the two channel imaging-based detection system; and

[0059] (iii) a control region, wherein the first calibration region and the control region both have a different nucleotide sequence in each artificial nucleic acid molecule, thereby producing a color balanced sequencing library.

[0060] Definitions

[0061] The term “about”, as applied to a numeric value, includes ± 10% of the recited value.

[0062] Unless otherwise specified, the terms “a” or “an” mean “one or more” throughout this application.

[0063] As used herein, the term “amplify” or “amplification” refers to the act or method of generating copies (i.e., amplicons) of a nucleic acid molecule. Methods of nucleic acid amplification are known in the art and include polymerase chain reaction (PCR), ligase chain reaction (LCR), looping-based amplification cycles (MALBAC), and multiple displacement amplification (MDA). In some instances, PCR may be performed using one or more primers.

[0064] As used herein, the terms “complement,” “complementary,” or “complementarity” in reference to nucleic acid sequences means that a sequence of a first nucleic acid in relation to that of a second nucleic acid will form hydrogen bonds with that of an opposing (i.e., antiparallel) nucleic acid strand. The 5’ end of the first nucleic acid when aligned to the 3’ end of the second nucleic acid, and vice versa to each other will have complementary structures following a lock-and-key principle (i.e., A will be paired with U or T and G will be paired with C). Further, a “reverse complement” in reference to a nucleic acid sequence means that the sequence direction is reversed, and each base is replaced with its complement. As an example, the reverse complement of a DNA sequence having the nucleotide sequence of 5’-TAGC-3’ is 5’-GCTA-3’.

[0065] As used herein, the term “color balancing” or variations thereof, in reference to a method of nucleic acid sequencing means preparing the sequence run and / or a nucleic acid sample for PATENT

[0066] ATTORNEY DOCKET NO. 51178-020W02 sequencing such that positively detected signals (e.g., fluorescence) of a detector system of a sequencing platform are accurately calibrated and interpreted for the duration of the sequence run to minimize errors (e.g., miscalled nucleobases) in acquired sequencing data. Generally, color balancing refers to balancing fluorophore signal intensities of each fluorophore detected by the detector system (e.g., a two-channel detector system, a three-channel detector system, or a four-channel detector system) of a sequencing platform. Methods of color balancing may include applying correction factors in data analysis following data acquisition. Methods of color balancing, such as the methods discussed herein, may include ensuring adequate nucleotide sequence diversity such that every position sequenced in a nucleic acid population over the course of a sequencing run leads to a positively detected signal in each channel of a detection system.

[0067] As used herein, the term “degenerate” in reference to a nucleic acid or portion of a nucleic acid, refers to a nucleic acid sequence or a portion thereof that contains random nucleotides, such that a population of degenerate nucleic acids includes a degenerate sequence region in each degenerate nucleic acid that differs in one or more nucleotides. In some embodiments, degenerate nucleic acids of a given length include nucleic acids having combinatorial sequences of the given length (e.g., comprising a large number of or all of the possible canonical nucleotide sequences of the given length (e.g., >50%, >60%, >70%, >80%, >90%, or 100% of the possible sequences).

[0068] As used herein, the term “flank” refers to the relative positions of three nucleic acid regions. A first and second nucleic acid region is said to flank a third nucleic acid region if the first and second regions lie upstream and downstream of the third nucleic acid region, respectively. The first and second nucleic acid regions may be directly adjacent to the third nucleic acid region.

[0069] As used herein, the term “homologous” refers to having substantially the same sequence. Homologous sequences may differ by up to one third of nucleotide bases. For example, two sequences that are nine bases in length may differ at most by 3, at most by 2, at most by 1 , or at most by 0 nucleotide bases, and remain homologous to one another.

[0070] As used herein, the term “hybridization” refers to a process in which two single-stranded nucleic acids bind non-covalently by base pairing to form a stable double-stranded nucleic acid. Hybridization may occur for the entire lengths of the two nucleic acids, or only for a portion or subregion of one or both of the nucleic acids. The resulting double-stranded nucleic acid molecule or region is a “duplex.”

[0071] As used herein, the term “index,” “index sequence,” or “barcode” refers to a short, predetermined nucleotide sequence that is appended 3’ and / or 5’ of nucleic acid molecules of a sample (e.g., in a nucleic acid library) typically to identify the source of the nucleic acid molecules (e.g., during sequencing). An index may be 5 to 15 nucleotides in length and serves as an identifying sequence to distinguish nucleic acid molecules that derive from distinct sources (e.g., different sample preparations or different sample sources) that are sequenced simultaneously. Typically, two different indices are attached to the samples. An index may be added to a nucleic acid via known methods in the art (e.g., tagmentation, ligation of an adapter nucleic acid that includes an index, amplification, among other known methods). Exemplary index sequences known in the art include i5 and I7 indices. PATENT

[0072] ATTORNEY DOCKET NO. 51178-020W02

[0073] As used herein, the terms “library” or “fragment library” refers to a collection of nucleic acids (e.g., engineered nucleic acids) derived from one or more nucleic acid samples, in which fragments of nucleic acid have been modified, generally by incorporating terminal adapter sequences comprising one or more domains to which one or more primers can bind and / or indices.

[0074] As used herein, the term “molar ratio” refers to the number or proportion of molecules of one species relative to other species in a mixture or a sample. For example, a molar ratio expressed as 5% means that 5 molecules in every 100 molecules total is a particular type of molecule.

[0075] As used herein, the term “nucleic acid” refers to a polymeric molecule of at least two linked nucleotides. The terms include, for example, deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), as well as hybrids and mixtures thereof. A nucleic acid may be double-stranded, single-stranded, or contain a mix of regions or portions of both single-stranded or double-stranded sequences. The nucleotides in a nucleic acid are usually linked by phosphodiester bonds, though “nucleic acid” may also refer to other molecular analogs having other types of chemical bonds or backbones, including, but not limited to, phosphoramide, phosphorothioate, phosphorodithioate, O-methyl phosphoramidate, morpholino, locked nucleic acid (LNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), and peptide nucleic acid (PNA) linkages or backbones. Nucleic acids may contain any combination of deoxyribonucleotides, ribonucleotides, or non-natural analogs thereof. Examples of nucleic acids include, but are not limited to, a gene, a gene fragment, a genomic gap, an exon, an intron, intergenic DNA (including, without limitation, heterochromatic DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, small interfering RNA (siRNA), miRNA, small nucleolar RNA (snoRNA), cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of a sequence, isolated RNA of a sequence, nucleic acid probes, and primers.

[0076] As used herein, a “nucleic acid sample” or an “analyte nucleic acid” refers to any nucleic acid (e.g., DNA) of interest that is selected for a method of the invention (e.g., color balancing a nucleic acid sample for massively parallel or next-generation sequencing) by combining the nucleic acid sample with a plurality of artificial nucleic acid molecules (e.g., artificial nucleic acid molecules described herein). The present methods can be carried out using nucleic acid samples (e.g., DNA samples) pooled from more than one source. It is to be understood that a nucleic acid sample may be DNA or RNA, for example. In some instances, RNA may be converted to cDNA prior to being selected for a method of the invention (e.g., color balancing a nucleic acid sample for massively parallel or next-generation sequencing).

[0077] As used herein, the term “nucleotide,” “nt,” or “nucleobase” refers to any deoxyribonucleotide, ribonucleotide, non-standard nucleotide, modified nucleotide, or nucleotide analog. Nucleotides include adenine, thymine, cytosine, guanine, and uracil. Examples of modified nucleotides include, but are not limited to, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5- carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D- galactosylqueosine, inosine, N6-isopentenyladenine, 1 -methylguanine, 1 -methylinosine, 2,2- dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6- adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D- PATENT

[0078] ATTORNEY DOCKET NO. 51178-020W02 mannosylqueosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6- isopentenyladenine, uracil-5-oxyacetic acid, wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5- methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, 5- methyl-2-thiouracil, and 3-(3-amino-3-N-2-carboxypropyl) uracil.

[0079] As used herein, the term “oligonucleotide” refers to a nucleic acid up to 500 nucleotides in length. Oligonucleotides may be synthetic. Oligonucleotides may contain one or more chemical modifications, whether on the 5’ end, the 3’ end, or internally. Examples of chemical modifications include, but are not limited to, addition of functional groups (e.g., biotins, amino modifiers, alkynes, thiol modifiers, phosphates, or azides), fluorophores (e.g., quantum dots or organic dyes), spacers (e.g., C3 spacer, dSpacer, photo-cleavable spacers), modified bases, or modified backbones.

[0080] As used herein, the term “PhiX174” refers to a bacteriophage that has a single-stranded DNA genome of 5,386 base pairs (Sanger et al., Nature. 265(5596) :687-695, 1977, which is hereby incorporated by reference). The sequence of the PhiX174 genome is available via the United States National Center for Biotechnology Information (Accession No. NC_001422).

[0081] As used herein the terms "a portion," “a part,” and / or grammatical equivalents thereof can refer to any fraction of a whole amount. For example, "a portion" can refer to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 99.9% or 100% of a whole amount.

[0082] As used herein, the term “primer” refers to an oligonucleotide, either natural or synthetic, that is capable of forming a duplex with a nucleic acid template and then acting as a point of initiation of nucleic acid synthesis for extension from its 3’ end along the template nucleic acid so that an extended duplex is formed, e.g., for nucleic acid amplification or sequencing by synthesis. The sequence of nucleotides added during the extension process is determined by the sequence of the template polynucleotide. Typically, primers are extended by a DNA polymerase. Primers usually have a length in the range of between 5 to 36 nucleotides (e.g., 5-12, 12-18, 18-24, 24-30, or 30-36 nucleotides). In certain aspects, primers are universal primers or non-universal primers. Pairs of primers can flank a sequence of interest or a set of sequences of interest. Primers can be degenerate in sequence to increase the number of regions that are hybridized in a nucleic acid sample or template, e.g., for non-sequence-specific nucleic acid amplification or sequencing. Primers may be universal primers, or primers that have nucleotide sequences that are well-known and commonly used in the art for sequencing or amplification methods. A nucleic acid (e.g., an artificial nucleic acid molecule described herein), may include a priming region at or near its 5’ and 3’ ends, in which the priming region of the 5’ end corresponds to a primer sequence and the priming region of the 3’ end corresponds to the reverse complement of the primer.

[0083] As used herein, the term “priming region” refers to a sequence to which a primer can bind or to a sequence that corresponds to a primer sequence (i.e., a reverse complement of a primer binding sequence).

[0084] As used herein, the terms “artificial,” “recombinant,” “engineered,” or “non-naturally occurring,” when used with reference to a cell, nucleic acid, or polypeptide, refers to a material, or a material corresponding to the natural or native form of the material, that has been produced by human PATENT

[0085] ATTORNEY DOCKET NO. 51178-020W02 intervention. In some instances, the cell, nucleic acid, or polypeptide may be modified in a manner that would not otherwise exist in nature. In some instances, the cell, nucleic acid, or polypeptide is identical to a naturally occurring cell, nucleic acid, or polypeptide, but is produced or derived from synthetic materials and / or by manipulation using recombinant techniques. Non-limiting examples include synthesized oligonucleotides such as nucleic acid molecules that are suitable for color balancing a sequencing reaction described herein or a nucleic acid fragment library generated from a nucleic acid sample.

[0086] As used herein, the term “redundant” in reference to a nucleic acid, a fragment thereof (e.g., a fragment generated for nucleic acid library preparation), or a nucleic acid adapter refers to multiple copies of a particular nucleotide sequence in a given sample (e.g., a generated sample for nucleic acid sequencing). Redundancy may arise due to identical nucleic acid adapters being appended to a particular nucleic acid fragment. In contrast, a “non-redundant” nucleic acid, fragment thereof, or a nucleic acid adapter refers to a nucleic acid with a unique identity (i.e., does not have an identical nucleic acid sequence) compared to other nucleic acids, fragments thereof, or nucleic acid adapters that are represented in a given sample (e.g., a generated sample for nucleic acid sequencing).

[0087] BRIEF DESCRIPTION OF THE DRAWINGS

[0088] FIG. 1A shows a design of the i5 index read positions with mixed bases in which N represents one of A, C, G, or T. The portion of the DNA sequence that corresponds to the 10-nucleotide i5 index read positions is indicated by the bases shown within the box. The sequence 5’ to the I5 index is the P5 attachment region (SEQ ID NO: 1), and the sequence 3’ to the I5 index is the Readl read sequence (SEQ ID NO: 3).

[0089] FIG. 1 B shows a design of the i7 index read positions with mixed bases in which N represents a mixture of A, C, G, or T. The portion of the DNA sequence that corresponds to the 10-nucleotide i7 index read positions is indicated by the bases shown within the box. The sequence 5’ to the i7 index is the P7 attachment region (SEQ ID NO: 2), and the sequence 3’ to the i7 index is the Read2 read sequence (SEQ ID NO: 4).

[0090] FIG. 1C shows a design of the i5 index read positions with mixed bases in which W represents a mixture of A or T, and D represents a mixture of A, T, or G. The i5 index is read in the reverse complement direction, so any G nucleobases in the index positions on the forward strand are read as C nucleobases. The portion of the DNA sequence that corresponds to the 10-nucleotide I5 index read positions is indicated by the bases shown within the box. The sequence 5’ to the i5 index is the P5 attachment region (SEQ ID NO: 1), and the sequence 3’ to the I5 index is the Readl read sequence (SEQ ID NO: 3).

[0091] FIG. 1D shows a design of the i7 index read positions with mixed nucleobases where W represents a mixture of A or T, and where H represents a mixture of A, T or C. The portion of the DNA sequence that corresponds to the 10-nucleotide i7 index read positions is indicated by the bases shown within the box. The sequence 5’ to the i7 index is the P7 attachment region (SEQ ID NO: 2), and the sequence 3’ to the i7 index is the Read2 read sequence (SEQ ID NO: 4). PATENT

[0092] ATTORNEY DOCKET NO. 51178-020W02

[0093] FIG. 2A shows the average Q30 score by position within the sequencing run without spike in of a PhiX174 color balancing library.

[0094] FIG. 2B shows the average Q30 score by position within the sequencing run with spike-in of a PhiX174 color balancing library.

[0095] FIG. 3A shows the average % of NoCalls (i.e., percent of uncalled nucleobases) by position within the sequencing run without spike-in of a PhiX174 color balancing library.

[0096] FIG. 3B shows the average % of NoCalls (i.e., percent of uncalled nucleobases) by position within the sequencing run with spike-in of a PhiX174 color balancing library.

[0097] DETAILED DESCRIPTION

[0098] Described herein is a plurality of artificial nucleic acid molecules (e.g., artificial DNA) that is suitable for use with methods of sequencing, e.g., massively parallel or next generation sequencing methods. Additionally, provided herein are methods of color balancing a sequence reaction by adding the plurality of artificial nucleic acid molecules to a prepared nucleic acid sample (e.g., a nucleic acid library for sequencing) for a method of sequencing, which leads to a proportional detection of positively detected signals over the course of the sequencing run and improves the accuracy and coverage of sequencing data. The artificial nucleic acid molecules described herein fulfill the color balancing requirements for low-plexity libraries on all sequencing systems but are especially useful for color balancing low-plexity libraries on two-color sequencing systems. Ensuring accurate color balancing during sequencing is critical to minimize signal misinterpretation, especially when sequencing combinatorial indexes in low-diversity libraries. Color balancing maintains the integrity of demultiplexing and improves the fidelity of downstream data analysis.

[0099] Color balancing in the context of nucleic acid sequencing (e.g., DNA sequencing) refers to the process of ensuring that the intensities of fluorescent signals emitted by different fluorophores are accurately calibrated and maintained throughout the sequencing run. This calibration is crucial for correctly interpreting the signals corresponding to each nucleotide or index base, especially in systems with limited signal channels, such as sequencing platforms with a two-channel detection system. Proper color balancing promotes accurate detection and interpretation of signals, reducing errors in base calling and improving sequencing accuracy. Color balancing is critical because any imbalance can lead to misinterpretation of signals, thus reducing the accuracy and depth of the sequencing results. Such imbalance may be prevalent in the presence of a dark signal in a two- channel detection system, since one base corresponds to the absence of a fluorescent signal. Improper calibration via lack of color balancing or insufficient color balancing can lead to miscalling of bases, especially in regions of low diversity where the presence of a single type of base might predominate and produce signals that lack contrast. Surprisingly, the artificial nucleic acids allow for sequencing of indices, which may have low plexity relative to the sequence of the inserts, with few or no miscalled bases. In particular, a relatively low amount of the artificial nucleic acids can be added to provide proper color balancing during sequencing, even with indices that would otherwise fail to color balance. PATENT

[0100] ATTORNEY DOCKET NO. 51178-020W02

[0101] I. Artificial Nucleic Acid Molecules

[0102] Provided herein are a plurality of artificial nucleic acid molecules that include an attachment region, a calibration region, and a control region. The artificial nucleic acids may be added to a nucleic acid sample (e.g., DNA) for sequencing to achieve proper color balancing for the duration of the sequencing run, especially when sequencing indices.

[0103] An artificial nucleic acid molecule may be 100 to 1 ,000 nucleotides in length. In some embodiments, an artificial nucleic molecule is 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, 400 to 450, 450 to 500, 500 to 550, 550 to 600, 600 to 650, 650 to 700, 700 to 750, 750 to 800, 800 to 850, 850 to 900, 900 to 950, or 950 to 1 ,000 nucleotides in length. In some embodiments, an artificial nucleic acid molecule is double-stranded. In some embodiments, an artificial nucleic acid molecule is single-stranded. In some embodiments, an artificial nucleic acid molecule has from its 5’ end to its 3’ end the following components: an attachment region and one or more calibration regions, e.g., for use as an adapter to attach to an insert. In some embodiments, an artificial nucleic acid molecule includes a control region 3’ of the calibration region. Additional regions may also be included in the artificial nucleic acids as described herein.

[0104] A. Attachment Region

[0105] An artificial nucleic acid molecule of the invention includes an attachment region at or near its 5’ end surface. The attachment region of the artificial nucleic acid molecule is a nucleotide sequence that can be linked to a surface to maintain signal integrity throughout the sequencing run. In some embodiments, the surface is a flow cell for a sequencing platform. In some embodiments, the surface is a plate such as a microwell plate. In some embodiments, the surface is a bead. In some embodiments, the attachment region links the artificial nucleic acid molecule to a surface by hybridizing to an oligonucleotide on the surface that is complementary to the attachment region or a contiguous portion thereof. Alternative attachment methods, such as biotin-streptavidin linkage, amine-reactive chemistry, or covalent binding strategies, are also options for linking the artificial nucleotide to sequencing surfaces.

[0106] In some embodiments, an attachment region is 5 to 20 nucleotides in length. In some embodiments, the attachment region is 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.

[0107] In some embodiments, the attachment region includes a priming region (e.g., a nucleic acid sequence of a primer or the reverse complement thereto (e.g., a domain for primer binding)) for amplification during a sequencing reaction. In some embodiments, a priming region corresponds to the nucleic acid sequence of at least one known nucleotide sequence such as a universal primer or the reverse complement thereto (e.g., the domain for universal primer binding). In some embodiments, the attachment region includes the nucleotide sequence 5’- AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1), which corresponds to universal primer P5, or the reverse complement thereto. In some embodiments, the attachment region includes the nucleotide sequence 5’CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2), which PATENT

[0108] ATTORNEY DOCKET NO. 51178-020W02 corresponds to universal primer P7, or the reverse complement thereto.

[0109] In some embodiments, an artificial nucleic acid molecule includes a second attachment region in addition to the first attachment region. Like the first attachment region, all or a portion of the second attachment region may be used for priming. In some embodiments, the second attachment region is downstream of a control region (e.g., 3’ to a control region). In some embodiments, the second attachment region is downstream of a read sequence (e.g., 3’ to a read sequence). In some embodiments, the second attachment region is downstream of a second calibration region (e.g., 3’ to a second calibration region). In some embodiments, the second attachment region is at the 3’ end of the artificial nucleic acid molecule.

[0110] In some embodiments, the second attachment region includes the nucleotide sequence 5’-AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1), which corresponds to universal primer P5, or the reverse complement thereto. In some embodiments, the second priming region of the includes the nucleotide sequence 5’-CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2), which corresponds to universal primer P7, or the reverse complement thereto. In preferred embodiments, the second attachment region is not the reverse complement of the first priming region of the artificial nucleic acid molecule to prevent circularization of the artificial nucleic acid molecule.

[0111] In some embodiments, each artificial nucleic acid molecule in a plurality of artificial nucleic acid molecules includes an attachment region that has one of two possible nucleic acid sequences.

[0112] B. Calibration Region

[0113] An artificial nucleic acid molecule of the invention includes one or more calibration regions, such as a first calibration region and a second calibration region, in which the nucleotide sequence of the one or more calibration regions differs among the artificial nucleic acid molecules in a plurality of artificial nucleic acid molecules. In some embodiments, the first calibration region is downstream of the attachment region (e.g., 3’ of the attachment region). A calibration region of an artificial nucleic acid molecule includes nucleobases that may be used to ensure that a detection system of a sequencing platform registers proportionate positive signals during each cycle of sequencing for the duration of a sequencing run. For example, in a two-channel detection system (e.g., an optical or imaging-based detection system), a calibration region includes nucleobases that correspond to a positive signal in one or both channels to minimize lack of detection of a signal (e.g., a nucleobase that is detected by the absence of a positive signal, e.g., a fluorophore). Furthermore, when used in a plurality, the combined sequences of the calibration regions can ensure that signal is detected in both channels during sequencing.

[0114] In some embodiments, for each nucleotide in the calibration region, a plurality provides approximately equal color signal in each channel, for example, in a two channel system, approximately 50%, such as 40-60% signal in each channel per nucleotide; in a three channel system, approximately 33%, such as 25-45% signal in each channel per nucleotide; or in a four channel system, approximately 25%, such at 20-30% signal in each channel per nucleotide.

[0115] The sequences of the calibration regions can be selected in any appropriate manner. In one embodiment, the plurality includes a set of known sequences selected to provide the desired color PATENT

[0116] ATTORNEY DOCKET NO. 51178-020W02 signal at each nucleotide. Such a set may include 10 or more sequences, e.g., 10-10,000 known sequences. In another embodiment, the sequences are randomly generated, e.g., by including a mixture of allowed nucleotides for each position, e.g., approximately 50:50 (e.g., between 40-60%) of two nucleotides or 33:33:33 (between 25-45%) of three nucleotides, the sequencing of which produced color signal.

[0117] For any plurality, the calibration region may exclude nucleotides that, when sequenced, would provide no signal, e.g., “dark bases,” such as G. (Notably, the dark base is always allowed when the calibration region is sequenced as its reverse complement.) In some embodiments, a percentage of “dark base” may be included in the calibration regions (sequenced in the sense direction) in the plurality, so long as the aggregate color signal is maintained. For example, in some embodiments, the “dark base” may be present at a percentage of 0 - 25%, e.g., 0 - 15 %, 0 - 10%, 0 - 5%, or 0 - 1%.

[0118] In some embodiments, a first calibration region has a nucleotide sequence consisting of nucleobases X and Y across its length, wherein nucleobase X represents one of two nucleobases selected from A, T, G, and C, and nucleobase Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase that is different from the first two nucleobases.

[0119] In some embodiments, nucleobase X represents A or T, and nucleobase Y represents A, T, or G. In some embodiments, nucleobase X represents A or T, and nucleobase Y represents A, T, or C. In some embodiments, nucleobase X represents A or G, and nucleobase Y represents A, G, or T. In some embodiments, nucleobase X represents A or G, and nucleobase Y represents A, G, or C. In some embodiments, nucleobase X represents A or C, and nucleobase Y represents A, C, or T. In some embodiments, nucleobase X represents A or C, and nucleobase Y represents A, C, or G. In some embodiments, nucleobase X represents T or G, and nucleobase Y represents T, G, or C. In some embodiments, nucleobase X represents T or G, and nucleobase Y represents T, G, or A. In some embodiments, nucleobase X represents T or C, and nucleobase Y represents T, C, or G. In some embodiments, nucleobase X represents T or C, and nucleobase Y represents T, C, or A. In some embodiments, nucleobase X represents G or C, and nucleobase Y represents G, C, or T. In some embodiments, nucleobase X represents G or C, and nucleobase Y represents G, C, or A. The specific combination of nucleobases may be dependent on a detection system of a sequencing platform and / or availability of nucleobases that correspond to a positive signal (e.g., fluorescent nucleobases).

[0120] In some embodiments, nucleobase X represents one of two nucleobases that correspond to a positive signal in a detection system of a sequencing platform. In some embodiments, the two nucleobases represented by X each correspond to a positive signal in one channel of a detection system (e.g., a two-channel detection system). In some embodiments, nucleobase X represents A or T at a given position of the first calibration region.

[0121] In some embodiments, nucleobase Y represents one of three nucleobases, wherein each of the three nucleobases corresponds to a positive signal in a detection system of a sequencing platform. In some embodiments, two of the three nucleobases of nucleobase Y each correspond to a positive signal in one channel of a detection system (e.g., a two-channel detection system), and the third nucleobase of nucleobase Y corresponds to a positive signal in two channels of the detection PATENT

[0122] ATTORNEY DOCKET NO. 51178-020W02 system. In some embodiments, nucleobase Y represents A, T, or C at a given position of the first calibration region. In some embodiments, nucleobase Y represents A, T, or G at a given position of the first calibration region.

[0123] While the sequence of the first calibration region may be completely randomly generated, the sequence may also be generated according to a pattern, either as a defined sequence or semirandomly. In some embodiments, the first calibration region has a nucleotide sequence in a repeating pattern of XY across its length. In some embodiments, the first calibration region has a nucleotide sequence in a repeating pattern of XXY across its length. In some embodiments, the first calibration region has a nucleotide sequence in a repeating pattern of XYY across its length. In some embodiments, the first calibration region has a nucleotide sequence of two or more of XY, XXY, XYY, YX, YXX, YYX, XYX, YXY, XX, and YY. An additional X or Y may also be appended to the 5’ or 3’ end of any first calibration sequence herein.

[0124] In some embodiments, the first calibration region is 5 to 30 nucleotides in length. In some embodiments, the first calibration region is 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length.

[0125] In some embodiments, an artificial nucleic acid molecule includes a second calibration region. In some embodiments, the second calibration region is 3’ of the control region. The second calibration region shares the properties of the first calibration region; however, it may be sequenced from its reverse complement.

[0126] In some embodiments, a second calibration region has a nucleotide sequence including nucleobases X and Y’ across its length, wherein nucleobase Y’ represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase complementary to the third nucleobase of Y. Here, X or Y’ may include the “dark base” when the reverse complement is sequenced.

[0127] In some embodiments, the second calibration region has a nucleotide sequence in a repeating pattern of Y’X across its length. In some embodiments, the second calibration region has a nucleotide sequence in a repeating pattern of Y’Y’X across its length. In some embodiments, the second calibration region has a nucleotide sequence in a repeating pattern of Y’XX across its length. In some embodiments, the second calibration region has a nucleotide sequence in a repeating pattern of Y’Y’XX across its length. In some embodiments, the second calibration region has a nucleotide sequence of two or more of XY’, XXY’, XY’Y’, Y’X, Y’XX, Y’Y’X, XY’X, Y’XY’, XX, and Y’Y’. An additional X or Y’ may also be appended to the 5’ or 3’ end of any second calibration sequence herein.

[0128] In some embodiments, the second calibration region is 5 to 30 nucleotides in length. In some embodiments, the second calibration region is 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length.

[0129] In certain embodiments, a plurality includes 10 - 100,000 different first calibration region sequences, e.g., at least 50, 100, 250, 500, 750, 1000, 1500, 2000, 5000, 7500, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, or more. The upper limit is determined by the length PATENT

[0130] ATTORNEY DOCKET NO. 51178-020W02 of the calibration region and the number of nucleotides employed. For a region of 10 nucleotides with three nucleotides possible at each position, the maximum number is about 60,000.

[0131] C. Control Region

[0132] In some embodiments, an artificial nucleic acid molecule includes a control region downstream of a first calibration region (e.g., 3’ to the first calibration region).

[0133] The control region includes a nucleic acid sequence of a known source to validate acquired sequencing data. The control region may be included in an artificial nucleic acid molecule (e.g., an artificial nucleic acid molecule that includes an attachment region and one or more calibration regions) by any suitable method (e.g., a method for nucleic library preparation) known in the art, including tagmentation, PCR amplification, isothermal amplification, and / or ligation.

[0134] In some embodiments, the control region is an artificial nucleotide sequence or a portion thereof (e.g., a portion of an artificial or synthetic plasmid).

[0135] In some embodiments, the control region is a portion of a genome of an organism such that a plurality of artificial nucleic acid molecules forms a library that derives from the genome or a portion thereof of the organism. In some embodiments, the genome is a naturally occurring genome or an engineered or artificial genome. In some embodiments, the genome is a selected portion of the genome.

[0136] In some embodiments, the genome is small and is about 3 kilobases to 6 kilobases in size. In some embodiments, the genome is about 3 kb, about 3.1 kb, about 3.2 kb, about 3.3 kb, about 3.4 kb, about 3.5 kb, about 3.6 kb, about 3.7 kb, about 3.8 kb, about 3.9 kb, about 4 kb, about 4.1 kb, about 4.2 kb, about 4.3 kb, about 4.4 kb, about 4.5 kb, about 4.6 kb, about 4.7 kb, about 4.8 kb, about 4.9 kb, about 5 kb, about 5.1 kb, about 5.2 kb, about 5.3 kb, about 5.4 kb, about 5.5 kb, about 5.6 kb, about 5.7 kb, about 5.8 kb, about 5.9 kb, or about 6 kb in size.

[0137] In some embodiments, the organism is a bacteriophage. In some embodiments, the bacteriophage is a bacteriophage of the Microviridae family. In some embodiments, the bacteriophage is an alphatrevirus (e.g., alpha3, ID21 , ID32, ID62, NC28, NC29, NC35, PhiK, St1 , or WA45), a gequatrovirus (e.g., G4, ID52, or Gequatrovirus talmos), a sinsheimervirus (e.g., PhiX174), a bdellomicrovirus (e.g., MH2K or MAC1), a chlamydiamicrovirus (e.g., Chp1 , Chp2, CPAR39, or CPG1), an enterogokushovirus (e.g., EC6098), or a spiromicrovirus (e.g., SpV4). In preferred embodiments, the bacteriophage is PhiX174. In some embodiments, a control region of an artificial nucleic acid molecule is 20 to 900 nucleotides in length. In some embodiments, the control region of the artificial nucleic acid molecule is 20 to 50, 50 to 100, 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, 400 to 450, 450 to 500, 500 to 550, 550 to 600, 600 to 650, 650 to 700, 700 to 750, 750 to 800, 800 to 850, or 850 to 900 nucleotides in length (e.g., 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 625, 650, 675, 700, 725, 750, 775, 800, 825, 850, 875, or 900 nucleotides in length). The control region of the nucleic acid molecule may derive from random portions of a genome or one or more specific regions of a genome. PATENT

[0138] ATTORNEY DOCKET NO. 51178-020W02

[0139] D. Read Sequences

[0140] In some embodiments, an artificial nucleic acid molecule includes one or more read sequences that may be priming regions for amplifying and sequencing a control region and / or a calibration region. In some embodiments, a read sequence includes the Readl nucleotide sequence 5’-TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or the reverse complement thereto. In some embodiments, a read sequence includes the Read2 nucleotide sequence 5’-GTCTCGTGGGCTCGG- 3’ (SEQ ID NO: 4) or the reverse complement thereto.

[0141] In some embodiments, an artificial nucleic acid molecule includes a first read sequence between a first calibration region and a control region. In some embodiments, an artificial nucleic acid molecule includes a second read sequence between a control region and a second calibration region. In some embodiments an artificial nucleic acid molecule includes a second read sequence between a control region and a second attachment region.

[0142] E. Exemplary Embodiments of the Nucleic Acids

[0143] The invention includes a plurality of artificial nucleic acid molecules in which each artificial nucleic acid molecule comprises from its 5’ end to its 3’ end an attachment region that optionally includes a first priming region, a first calibration region, a first read sequence, a control region, a second read sequence, a second calibration region, and a second attachment region that optionally includes a second priming region.

[0144] In some embodiments, the first or second attachment region of each artificial nucleic acid molecule links each artificial nucleic acid molecule to a surface (e.g., a surface of a flow cell). In some embodiments, the first or second attachment region links to the surface via hybridization by an oligonucleotide sequence that is the reverse complement to the nucleotide sequence of the first or second attachment region or a contiguous portion thereof. In some embodiments, the first or second attachment region links to the surface via an oligonucleotide sequence that is complementary to the first priming region. In some embodiments, the first attachment region includes the P5 nucleotide sequence 5’-AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1) or the reverse complement thereto. In some embodiments, the first attachment region of the includes the P7 nucleotide sequence 5’-CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2) or the reverse complement thereto. In preferred embodiments, the first attachment region includes the P5 nucleotide sequence 5’-AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1).

[0145] In some embodiments, the first calibration region of each artificial nucleic acid molecule is about 10 nucleotides in length and has a nucleotide sequence that has a repeating pattern of XY across its length, wherein nucleobase X is A or T, and nucleobase Y is either A, T, or C. When a plurality of artificial nucleic acids is employed, the first calibration region of each artificial nucleic acid molecule has a different nucleotide sequence. It will, however, be understood that when the calibration region is generated randomly that duplicates of a particular calibration sequence may be present in the library as a whole. So long as the library includes a plurality of artificial nucleic acids with different calibration sequences, then the presence of duplicates does not affect the color PATENT

[0146] ATTORNEY DOCKET NO. 51178-020W02 balancing properties of the library.

[0147] In some embodiments, the first read sequence includes the nucleotide sequence 5’- TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or the reverse complement thereto. In some embodiments, the first read sequence includes the nucleotide sequence 5’-GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4) or the reverse complement thereto. In preferred embodiments, the first read sequence includes the nucleotide sequence 5’-TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3).

[0148] In some embodiments, the control region of each artificial nucleic acid molecule is about 20 to about 900 nucleotides in length and has a nucleotide sequence that corresponds to a portion of the genome of bacteriophage PhiX174. In some embodiments, the control region of each artificial nucleic acid molecule in a plurality has a different nucleotide sequence. As with randomly generated calibration sequences, randomly generated control sequences may also be duplicated in the library without adverse effect.

[0149] In some embodiments, the second read sequence includes the nucleotide sequence 5’- TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or the reverse complement thereto. In some embodiments, the second read sequence includes the nucleotide sequence 5’-GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4) or the reverse complement thereto. In preferred embodiments, the second read sequence includes the nucleotide sequence that is the reverse complement of 5’-GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4).

[0150] In some embodiments, the second calibration region of each artificial nucleic acid molecule is about 10 nucleotides in length and has a nucleotide sequence that has a repeating pattern of Y’X across its length, wherein nucleobase X is A or T, and nucleobase Y’ is either A, T, or G. When a plurality of artificial nucleic acids is employed, the second calibration region of each artificial nucleic acid molecule has a different nucleotide sequence. It will, however, be understood that, when the calibration region is generated randomly, duplicates of a particular calibration sequence may be present in the library as a whole. So long as the library includes a plurality of artificial nucleic acids with different calibration sequences, then the presence of duplicates does not affect the color balancing properties of the library.

[0151] In some embodiments, the second attachment region includes the P5 nucleotide sequence 5’- AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1) or the reverse complement thereto. In some embodiments, the second attachment region of the includes the P7 nucleotide sequence 5’- CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2) or the reverse complement thereto. In preferred embodiments, the second attachment region includes the reverse complement of the P7 nucleotide sequence 5’-CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2).

[0152] II. Methods

[0153] Methods of the invention include color balancing a sequencing reaction or preparing a color balanced nucleic acid sample for sequencing by adding to a sample that includes analyte nucleic acid molecules (e.g., DNA) a plurality of artificial nucleic acid molecules (i.e., a plurality of nucleic acid molecules described herein). Analyte nucleic acid molecules (e.g., DNA including DNA prepared from reverse transcription of RNA)) may be obtained from one or more cells or one or more tissues from an PATENT

[0154] ATTORNEY DOCKET NO. 51178-020W02 organism (e.g., a prokaryote or a eukaryote). In some embodiments, the plurality of artificial nucleic acid molecules is added to a sample as a plurality of double-stranded artificial nucleic acid molecules (i.e., a plurality of artificial nucleic acid molecules and the complementary strands thereof).

[0155] Color balancing methods may be used with massively parallel or next-generation sequencing platforms that utilize sequencing by synthesis (SBS) technology (e.g., XLEAP-SBS™ technology). For systems that utilize XLEAP-SBS™ technology, when detected, cytosine (C) is represented by a combination of blue and green signals, adenine (A) by a blue signal only, thymine (T) by a green signal only, and guanine (G) remains a "dark base" with no signal. Because of these assignments, it is crucial that both the blue and green channels are active during every cycle of indexing (i.e., detecting one or more calibration regions), minimizing the risk of signal registration issues. In such systems, it is recommended to avoid nucleotide sequence combinations in a calibration region (e.g., the first and / or the second calibration region) that result in only blue signals (e.g., A or A and G) or no signal at all (e.g., G) in any given cycle. Additionally, it is recommended that at least one of the first two bases in a calibration region (e.g., the first and / or the second calibration region) is not a G, as starting with two G nucleobases can prevent proper signal detection and cause registration errors. It will be understood that the sequence of the calibration regions of the present artificial nucleic acids can be adjusted for other color assignments in two channel systems. In some instances, the combination of blue and green signals may correspond to different nucleobases. As a non-limiting example, a two channel sequencing platform may be configured such that a blue signal corresponds to cytosine (C), a green signal corresponds to guanine (G), simultaneous blue and green signals correspond to thymine (T), and adenine is the “dark base” with no signal. In some instances, a two channel system may be configured to detect color combinations distinct from blue and green signals (e.g., green and yellow, red and green, blue and yellow, yellow and red, etc.). In such instances, the sequence of the one or more calibration regions of the artificial nucleic acids will reflect the specific color configuration such that each channel is proportionally active during each indexing cycle for the duration of a sequencing reaction.

[0156] Exemplary sequencing platforms that are compatible with methods of color balancing (i.e., color balancing with a plurality of artificial nucleic acid molecules described herein) are the ILLUMINA™ NEXTSEQ™ 2000 and the ILLUMINA™ NOVASEQ™ X sequencing platforms.

[0157] In some embodiments, preparing a color balanced nucleic acid sample includes adding a plurality of artificial nucleic acid molecules to analyte nucleic acid molecules (e.g., DNA) to produce a color balanced sequencing library. In some embodiments, a color balanced sequencing library includes a plurality of nucleic acid molecules at an amount less than or equal to a molar ratio (of the plurality relative to the sample) of 20% (e.g., 0.1% to 20%, 1 % to 20%, 2% to 20%, 3% to 20%, 4% to 20%, 5% to 20%, 6% to 20%, 7% to 20%, 8% to 20%, 9% to 20%, 10% to 20%, 1 1 % to 20%, 12% to 20%, 13% to 20%, 14% to 20%, 15% to 20%, 16% to 20%, 17% to 20%, 18% to 20%, or 19% to 20%) relative to the analyte nucleic acid molecules. In some embodiments, the analyte nucleic acid molecules include combinatorial dual indices in which one or both indices have low-plexity, in which a method of sequencing without preparing a color balanced nucleic acid sample yields a higher rate of miscalling nucleobases of the analyte nucleic acid molecules as compared to a color balanced nucleic PATENT

[0158] ATTORNEY DOCKET NO. 51178-020W02 acid sample. In some embodiments, the analyte nucleic acid molecules include unique dual indices (UDIs) with a low number of indices such that the total combination of indices gives a higher propensity for miscalling nucleobases in a method of sequencing without preparing a color balanced nucleic acid sample, as compared to a color balanced nucleic acid sample. For example, the sample may include 10 or fewer UDIs. In certain embodiments, the sample being sequenced has fewer than 10,000, 1 ,000, 100, or 10 distinct index sequences in the aggregate. In certain embodiments, the sample may include a set of indexes that do not color balance in the absence of a plurality of the artificial nucleic acids described herein.

[0159] In some embodiments, the color balanced sequencing library includes a plurality of artificial nucleic acid molecules at a molar ratio of 0.01% to 5%, 0.01% to 4.5% , 0.01% to 4%, 0.01 to 3.5%, 0.01 to 3%, 0.01 to 2.5%, 0.01 to 2%, 0.01 to 1 .5%, 0.01 to 1%, 0.01 to 0.5%, 0.1% to 5%, 0.1% to 4.5%, 0.1% to 4%, 0.1% to 3.5% , 0.1 to 3%, 0.1 to 2.5%, 0.1 to 2%, 0.1 to 1 .5%, 0.1 to 1%, 0.1 to 0.5%, 1% to 5%, 1% to 4.5%, 1% to 4%, 1% to 3.5%, 1 to 3%, 1 to 2.5%, 1 to 2%, 1 to 1 .5%, relative to the analyte nucleic acid molecules. In some embodiments, the color balanced sequencing library includes a molar ratio of about 5%, 4.5%, 4%, 3.5%, 3%, 2.5%, 2%, 1 .5%, 1%, 0.5%, or 0.1 %of the plurality of artificial nucleic acid molecules relative to the analyte nucleic acid molecules. In some embodiments, the molar ratio of artificial nucleic acid molecules is validated prior to sequencing. In some embodiments, the molar ratio is validated via a quantitative method such as quantitative PCR (qPCR) or fluorometrically.

[0160] In some embodiments, preparing a color balanced nucleic acid sample includes adding an amount of artificial nucleic acid molecules to analyte nucleic acid molecules such that a certain amount of a surface (e.g., a flow cell) of a sequencing platform is covered by a plurality of artificial nucleic acid molecules. In some embodiments, 5% or less, 4% or less, 3% or less, 2% or less, or 1% or less of the covered surface of a sequencing platform is covered by a plurality of artificial nucleic acid molecules.

[0161] In some embodiments, preparing a color balanced nucleic sample includes adding an amount of artificial nucleic acid molecules to analyte nucleic acid molecules such that 10% or fewer (e.g. 10% or fewer, 9% or fewer, 8% or fewer, 7% or fewer, 6% or fewer, 5% or fewer, 4% or fewer, 3% or fewer, 2% or fewer, or 1% or fewer) of the acquired sequencing reads correspond to the plurality of artificial nucleic acid molecules.

[0162] In some embodiments, color balancing a sequencing reaction (e.g., by one of the preparation methods described above) yields sequencing data that has fewer than 15% (e.g., fewer than 15%, fewer than 14%, fewer than 13%, fewer than 12%, fewer than 11 %, fewer than 10%, fewer than 9%, fewer than 8%, fewer than 7%, fewer than 6%, fewer than 5%, fewer than 4%, fewer than 3%, fewer than 2%, or fewer than 1%) miscalled nucleobases (i.e., nucleobases that could not be accurately assigned). The artificial nucleic acids of the invention may also be used as an internal control even when the samples being sequenced have sufficient plexity to provide color balance. Artificial nucleic acids may also be employed in color balancing in systems using three or four channel detection. PATENT

[0163] ATTORNEY DOCKET NO. 51178-020W02

[0164] EXAMPLES

[0165] The invention is described by the following non-limiting examples.

[0166] Example 1 : Indexing nucleic acid libraries for sequencing

[0167] Ninety-six pUC19 plasmid libraries were prepared in one 96-well plate using an EXPRESSPLEX™ 2.0 library preparation kit (seqWell Inc., Beverly, MA). The batch of 96 samples received 96 well-specific i7 indexes, which after sequencing allow identification of the samples originating from each well, and a single i5 index (Indexing Plate 2004). This general indexing approach is known as combinatorial dual indexing (CDI), where in this example, there are 96 I7 indexes x 1 i5 indexes, providing a total of 96 index combinations of the plasmid libraries. In this example, the i5 index “2004” was deliberately chosen based on its incompatibility with Illumina’s color balancing recommendations for XLEAP sequencing chemistry. The resultant library was loaded to two separate NEXTSEQ™ 2000 XLEAP-SBS™ sequencing runs: one without a color balancing library (control) and one with an indexed PhiX174 color balancing library (~5% of the total library loaded) prepared with the sequences shown in FIGS. 1C-1D. The results from the two sequencing runs are shown in Table 1, FIGS. 2A-2B, and FIGS. 3A-3B after de-multiplexing against the expected 96 known barcode combinations.

[0168] Table 1 : Sequencing results of plasmid libraries with or without a color balance index

[0169] The good performance observed with the PhiX174 color balancing library was not observed with the non-indexed PhiX library (Illumina) control. The high percentages of Undetermined Reads (59.52%) and One Mismatch Barcode (40.48%) in the control sequencing run would be a wholly unacceptable result to any person with ordinary skill in the art.

[0170] When the PhiX174 color balancing library was spiked in at approximately 5% of the total library mass loaded onto the sequencer, a dramatic improvement in the demultiplexing metrics were observed. The percentage of Perfect Barcode (index reads) increased from 0% to 97.18%. The percentage of undetermined reads decreased from 59.52% down to 6.46%, a 9.2-fold improvement. PATENT

[0171] ATTORNEY DOCKET NO. 51178-020W02

[0172] Example 2: Color balancing a nucleic acid sample with a plurality of artificial nucleic acid molecules

[0173] This example demonstrates that color balancing a nucleic acid sample may be achieved with a plurality of artificial nucleic acid molecules in which each artificial nucleic acid molecule includes one or more calibration regions selected from a group of calibration regions that consist of pre-determined nucleotide sequences.

[0174] An analyte nucleic acid sample (e.g., a DNA library) may be color balanced in a nextgeneration sequencing reaction by adding an amount of a plurality of artificial nucleic acid molecules. Each artificial nucleic acid molecule includes from its 5’ end to its 3’ end an attachment region that optionally includes a first priming region, a first calibration region, a first read sequence, a control region (e.g., a portion of DNA from a known nucleotide sequence such as a contiguous portion of the PhiX174 genome), a second read sequence, a second calibration region, and a second attachment region that optionally includes a second priming region.

[0175] Each artificial nucleic acid molecule includes a first calibration region selected from a group consisting of 5 to 20 possible calibration regions (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 calibration regions), in which each of the possible calibration regions has a different nucleotide sequence. In some embodiments, each artificial nucleic acid molecule includes a second calibration region that is a reverse complement of a different calibration region than the first calibration region.

[0176] A calibration region of the possible calibration regions may be 5 to 30 nucleotides in length. A calibration region of the possible calibration regions may have a pre-determined nucleotide sequence that is characterized by any one of the repeating patterns described herein (see Section IB). In some embodiments, the number of possible calibration regions and the nucleotide sequence therein is prepared such that each position across the possible calibration regions corresponds to a positive signal in the detection system of the sequencing platform (e.g., a two-channel detection system).

[0177] The plurality of artificial nucleic acid molecules is added to the analyte nucleic acid sample such that the color balanced sequencing library includes about 5%, 4.5%, 4%, 3.5%, 3%, 2.5%, 2%, 1 .5%, 1%, 0.5%, or 0.1 % molar ratio of the plurality of artificial nucleic acid molecules relative to the analyte nucleic acid molecules.

[0178] Preparing a color balanced sequencing library that includes a plurality of artificial nucleic acid molecules yields sequencing data that has fewer than 15% (e.g., fewer than 15%, fewer than 14%, fewer than 13%, fewer than 12%, fewer than 11%, fewer than 10%, fewer than 9%, fewer than 8%, fewer than 7%, fewer than 6%, fewer than 5%, fewer than 4%, fewer than 3%, fewer than 2%, or fewer than 1%) miscalled nucleobases (i.e., nucleobases that could not be accurately assigned).

[0179] Other Embodiments

[0180] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each independent publication or patent application was specifically and individually indicated to be incorporated by reference. PATENT

[0181] ATTORNEY DOCKET NO. 51178-020W02

[0182] While the invention has been described in connection with specific embodiments thereof, it will be understood that it is capable of further modifications and this application is intended to cover any variations, uses, or adaptations following, in general, the principles and including such departures from the invention that come within known or customary practice within the art to which the invention pertains and may be applied to the essential features hereinbefore set forth, and follows in the scope of the claims.

[0183] Other embodiments are within the claims.

Claims

PATENTATTORNEY DOCKET NO. 51178-020W02CLAIMS1 . A plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules comprises from its 5’ end to its 3’ end the following:(a) an attachment region for linking the artificial nucleic acid molecule to a surface;(b) a first calibration region having a nucleotide sequence in a repeating pattern of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from adenine (A), thymine (T), guanine (G), or cytosine (C), and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases; and(c) a control region, wherein the first calibration region and the control region both have a different nucleotide sequence in each artificial nucleic acid molecule.

2. A plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules comprises from its 5’ end to its 3’ end the following:(a) an attachment region for linking the artificial nucleic acid molecule to a surface; and(b) a calibration region having a nucleotide sequence in a repeating pattern of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from A, T, G, or C, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases; wherein the first calibration region of each artificial nucleic acid molecule has a different nucleotide sequence.

3. The plurality of artificial nucleic acid molecules of claim 2, wherein each of the artificial nucleic acid molecules further comprises a control region at its 3’ end and the control region has a different nucleotide sequence in each artificial nucleic acid molecule.

4. The plurality of artificial nucleic acid molecules of any one of claims 1 -3, wherein the attachment region comprises a first priming region.

5. The plurality of artificial nucleic acid molecules of claim 4, wherein the first priming region comprises a nucleotide sequence of 5’-AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1) or 5’-CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2).

6. The plurality of artificial nucleic acid molecules of any one of claims 1 -5, wherein each artificial nucleic acid molecule further comprises a second priming region 3’ of the control region.

7. The plurality of artificial nucleic acid molecules of claim 6, wherein the second priming region comprises a nucleotide sequence of the reverse complement of 5’-PATENTATTORNEY DOCKET NO. 51 178-020W02AATGATACGGCGACCACCGAGATCTACAC-3’ (SEQ ID NO: 1 ) or 5’- CAAGCAGAAGACGGCATACGAGAT-3’ (SEQ ID NO: 2).

8. The plurality of artificial nucleic acid molecules of any one of claims 1 and 3-7, wherein each artificial nucleic acid molecule further comprising a first read sequence between the first calibration region and the control region.

9. The plurality of artificial nucleic acid molecules of claim 8, wherein the read sequence comprises 5’-TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or 5’-GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4).

10. The plurality of artificial nucleic acid molecules of any one of claims 1 -9, wherein the first calibration region is 5 to 20 nucleotides in length.1 1 . The plurality of artificial nucleic acid molecules of any one of claims 1 -10, wherein the attachment region is 5 to 20 nucleotides in length.

12. The plurality of artificial nucleic acid molecules of any one of claims 1 and 3-10, wherein each artificial nucleic acid molecule further comprises a second calibration region 3’ of the control region, wherein the second calibration region has a nucleotide sequence in a repeating pattern of Y’X across the length of the second calibration region, wherein Y’ represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase complementary to the third nucleobase of Y, wherein the second calibration region has a different nucleotide sequence in each artificial nucleic acid molecule.

13. The plurality of artificial nucleic acid molecules of claim 12, wherein the second calibration region is 5 to 20 nucleotides in length.

14. The plurality of artificial nucleic acid molecules of claim 12 or 13, further comprising a second read sequence between the control region and the second calibration region.

15. The plurality of artificial nucleic acid molecules of claim 14, wherein the read sequence comprises the reverse complement of 5’-TCGTCGGCAGCGTC-3’ (SEQ ID NO: 3) or 5’- GTCTCGTGGGCTCGG-3’ (SEQ ID NO: 4).

16. The plurality of artificial nucleic acid molecules of any one of claims 1 -15, wherein the artificial nucleic acid molecules are 100 to 1 ,000 nucleotides in length.

17. The plurality of artificial nucleic acid molecules of any one of claims 1 and 3-16, wherein the control region is a portion of an artificial sequence or a portion of a genome of an organism.PATENTATTORNEY DOCKET NO. 51 178-020W0218. The plurality of artificial nucleic acid molecules of claim 17, wherein the organism is PhiX174.

19. The plurality of artificial nucleic acid molecules of any one of claims 9-18, wherein the first calibration region is 10 nucleotides in length.

20. The plurality of artificial nucleic acid molecules of any one of claims 1 -19, wherein the surface is a flow cell for nucleic acid sequencing.21 . The plurality of nucleic acid molecules of any one of claims 1 -20, wherein the plurality of nucleic acid molecules is linked to the surface via hybridization of the attachment region and a complementary oligonucleotide.

22. The plurality of artificial nucleic acid molecules of any one of claims 1 -21 , wherein the artificial nucleic acid molecules comprise DNA.

23. The plurality of artificial nucleic acid molecules of any one of claims 1 -22, wherein the two nucleobases of X of the calibration region are A and T.

24. The plurality of artificial nucleic acid molecules of claim 23, wherein, in Y, the third nucleobase is G.

25. A method of sequencing a plurality of nucleic acid molecules, wherein the method comprises:(a) adding to a sequencing library of nucleic acid molecules a plurality of artificial nucleic acid molecules to produce a color balanced sequencing library, wherein each of the artificial nucleic acid molecules comprises from its 5’ end to its 3’ end the following:(i) an attachment region for linking the artificial nucleic acid molecule to a surface;(ii) a first calibration region having the order of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from A, T, G, or C, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases, wherein the first calibration region has a different nucleotide sequence in each artificial nucleic acid molecule; and(iii) a control region, wherein the control region has a different nucleotide sequence in each artificial nucleic acid molecule;(b) applying the color balanced sequencing library to the surface; and(c) sequencing with the color balanced sequencing library.PATENTATTORNEY DOCKET NO. 51 178-020W0226. A method of preparing a color balanced sequencing library, the method comprising adding to a sequencing library of nucleic acid molecules a plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules comprises from its 5’ end to its 3’ end the following:(i) an attachment region for linking the artificial nucleic acid molecule to a surface;(ii) a first calibration region having the order of XY across the length of the calibration region, wherein X represents one of two nucleobases selected from A, T, G, or C, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases; and(iii) a control region, wherein the first calibration region and the control region both have a different nucleotide sequence in each artificial nucleic acid molecule, thereby producing a color balanced sequencing library.

27. The method of claim 26, further comprising sequencing the color balanced sequencing library.

28. The method of any one of claims 25-27, wherein each artificial nucleic acid molecule further comprises a second calibration region 3’ of the control region, wherein the second calibration region has a nucleotide sequence in a repeating pattern of Y’X across the length of the second calibration region, wherein Y’ represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase complementary to the third nucleobase of Y, wherein the second calibration region has a different nucleotide sequence in each artificial nucleic acid molecule.

29. The method of any one of claims 25-28, wherein the color balanced sequencing library comprises the plurality of artificial nucleic acid molecules at a molar ratio of less than or equal to 20%.

30. The method of claim 29, wherein the color balanced sequencing library comprises a molar ratio of about 5%, 4.5%, 4%, 3.5%, 3%, 2.5%, 2%, 1 .5%, 1 %, or 0.1 % of the plurality of artificial nucleic acid molecules.31 . The method of claim 30, wherein the color balanced sequencing library comprises a molar ratio of about 3.5% of the plurality of artificial nucleic acid molecules.

32. The method of claim 30, wherein the color balanced sequencing library comprises a molar ratio of about 3% of the plurality of artificial nucleic acid molecules.

33. The method of claim 32, wherein the color balanced sequencing library comprises a molar ratio of about 1 % of the plurality of artificial nucleic acid molecules.PATENTATTORNEY DOCKET NO. 51 178-020W0234. The method of any one of claims 25 and 27-33, wherein the method of sequencing comprises an imaging-based detection system.

35. The method of claim 34, wherein the detection system comprises a first channel that detects a first positive signal, and a second channel that detects a second positive signal.

36. The method of claim 35, wherein the first positive signal corresponds to one of the two nucleobases of X, and the second positive signal corresponds to the other nucleobase of the two nucleobases of X.

37. The method of claim 35 or 36, wherein the first positive signal and the second positive signal corresponds to the third nucleobase of Y.

38. The method of any one of claims 25-37, wherein the two nucleobases are A and T.

39. The method of claim 38, wherein the third nucleobase is C.

40. The method of any one of claims 25 and 27-39, wherein the method of sequencing comprises next generation sequencing (NGS).41 . A method of preparing a color balanced sequencing library, the method comprising adding to a sequencing library of nucleic acid molecules a plurality of artificial nucleic acid molecules, wherein each of the artificial nucleic acid molecules comprises from its 5’ end to its 3’ end the following:(i) an attachment region for linking the artificial nucleic acid molecule to a surface;(ii) a first calibration region having the order of XY across the length of the calibration region, wherein X represents one of two nucleobases detected in a single channel of a two channel imagingbased detection system, and Y represents one of three nucleobases consisting of the same two nucleobases of X and a third nucleobase, wherein the third nucleobase is a different nucleobase than the two nucleobases of X and is detected in both channels of the two channel imaging-based detection system; and(iii) a control region, wherein the first calibration region and the control region both have a different nucleotide sequence in each artificial nucleic acid molecule, thereby producing a color balanced sequencing library.

Citation Information

Patent Citations

  • Nucleic acid sequencing adapters and uses thereof

    US20170355984A1

  • Sequencing oligonucleotides and methods of use thereof

    US20240011020A1

  • Method for quantitating nucleic acid library

    WO2021155252A1