Universal short adapters for indexing of polynucleotide samples

By using index oligonucleotides to contact the target nucleic acid in large-scale parallel multiple sequencing, the index-target polynucleotides are generated and sequenced, the problem of index misallocation in the prior art is solved, and accurate identification and efficient sorting of sample sources are achieved.

CN120060440APending Publication Date: 2025-05-30ILLUMINA INC
View PDF 25 Cites 0 Cited by

Patent Information

Application Number
CN202510226948.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-06-23
Filing Date
2018-05-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has the problem of index error allocation in large-scale parallel multiple sequencing, resulting in the flux gain of sample multiplexing accompanied by the complexity and error rate of data analysis.

Method used

An index oligonucleotide configured to identify sample sources in large-scale parallel multiple sequencing is used to determine the sample source by contacting multiple index polynucleotides with the target nucleic acid, and obtaining the index sequence and reads of the target sequence through sequencing, and the sample source of the target read is determined using the index read.

Benefits of technology

Effectively identifying and sorting target nucleic acids from multiple samples reduces the incidence of index misallocation and improves the accuracy and analysis efficiency of sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120060440A_ABST
    Figure CN120060440A_ABST
Patent Text Reader

Abstract

The invention relates to a universal short adaptor for indexing of polynucleotide samples. The disclosed embodiments relate to indexed oligonucleotides configured to identify a source of a nucleic acid sample, as well as methods, apparatus, systems, and computer program products for identifying and preparing indexed oligonucleotides. In some embodiments, the indexed oligonucleotide comprises a set of indexed sequences, the Hamming distance between any two indexed sequences in the set of indexed sequences satisfying one or more criteria. Systems, devices, and computer program products for determining a sequence of interest using indexed oligonucleotides are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with the filing date of May 7, 2018, application number 201880042758.8, and invention title "Universal Short Adapters for Indexing of Polynucleotide Samples".

[0002] Cross - reference to related applications

[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 503,272, entitled UNIVERSAL SHORT ADAPTERS FOR INDEXING OF POLYNUCLEOTIDE, filed on May 8, 2017, under 35 U.S.C. § 119(e); this application also claims the benefit of U.S. Provisional Patent Application No. 62 / 524,390, entitled OPTIMAL INDEX SEQUENCES FOR MULTIPLEX MASSIVELY PARALLEL SEQUENCING, filed on June 23, 2017, under 35 U.S.C. § 119(e); for all purposes, the entire above prior applications are incorporated herein by reference in their entirety.

[0004] Incorporation by reference of the sequence listing

[0005] This application includes a sequence listing, which has been electronically submitted in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy created on May 3, 2018 is named ILMNP023_ST25.txt and is 29,527 bytes in size. Background

[0007] The present disclosure relates, among other things, to sequencing polynucleotides from multiple libraries; and more particularly, to increasing the likelihood that sequencing correctly identifies the library from which a polynucleotide originated.

[0008] Improvements in next-generation sequencing (NGS) technologies have greatly increased sequencing speed and data output, which has led to large-scale sample throughput on current sequencing platforms. Approximately 10 years ago, the Illumina Genome Analyzer was capable of generating up to 1 gigabyte of sequence data per run. Today, the Illumina NovaSeq TM series of systems is capable of generating up to 2 terabytes of data in two days, which represents an increase of greater than 2000x in capacity.

[0009] One aspect of implementing this increased capacity is multiplexing, which adds a unique sequence called an index to each DNA fragment during library preparation. This allows large numbers of libraries to be pooled and sequenced simultaneously during a single sequencing run. The throughput gain from multiplexing is accompanied by an additional layer of complexity because, prior to final data analysis, the sequencing reads from the pooled libraries need to be computationally identified and sorted in a process called demultiplexing. Incorrect assignment of indices between multiplexed libraries is a known problem that has affected NGS technology since the time of the development of sample multiplexing (Kircher et al., 2012, Nucleic Acids Res., Vol. 40, No. 1).

[0010] Overview

[0011] The disclosed embodiments relate to index oligonucleotides configured to identify the origin of samples in massively parallel multiplexed sequencing. Also provided are methods, devices, systems, and computer program products for preparing and using the index oligonucleotides.

[0012] One aspect of the present disclosure provides a method for sequencing target nucleic acids derived from multiple samples. The method includes: (a) contacting a plurality of index polynucleotides with target nucleic acids derived from multiple samples to generate a plurality of index-target polynucleotides, wherein the index polynucleotide contacting the target nucleic acid derived from each sample comprises an index sequence or a combination of index sequences uniquely associated with this sample, the index sequence or the combination of index sequences being selected from a set of index sequences, and the Hamming distance between any two index sequences in the set of index sequences is not less than a first criterion value, wherein the first criterion value is at least 2; (b) pooling the plurality of index-target polynucleotides; (c) sequencing the pooled index-target polynucleotides to obtain a plurality of index reads of the index sequences and a plurality of target reads of the target sequences, each target read being associated with at least one index read; and (d) using the index reads to determine the sample origin of the target reads.

[0013] In some embodiments, the set of index sequences comprises multiple pairs of color-balanced index sequences, wherein any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine (A) base or a cytosine (C) base, and (ii) a guanine (G) base, a thymine (T) base, or a uracil (U) base.

[0014] In some embodiments, the combination of index sequences is an ordered combination of index sequences.

[0015] In some embodiments, using indexed reads to determine the sample origin of target reads includes: for each indexed read, obtaining an alignment score for the indexed read with respect to a set of indexing sequences, where each alignment score indicates the similarity between the sequence of the indexed read and an indexing sequence in the set of indexing sequences; determining that a particular indexed read matches a particular indexing sequence based on the alignment score; and determining that the target reads associated with the particular indexed read originate from a sample uniquely associated with the particular indexing sequence.

[0016] In some embodiments, the plurality of indexed polynucleotides comprise a plurality of indexing primers, where the indexing primers comprise indexing sequences from the set of indexing sequences. In some embodiments, each indexing primer further comprises a flow cell amplification primer binding sequence. In some embodiments, the flow cell amplification primer binding sequence comprises a P5 sequence or a P7′ sequence.

[0017] In some embodiments, the target nucleic acids from the plurality of samples comprise nucleic acids having universal adapters covalently attached to one or both ends. In some embodiments, contacting the plurality of indexed polynucleotides with the target nucleic acids from the plurality of samples includes: hybridizing the plurality of indexing primers to the universal adapters covalently attached to one or both ends of the nucleic acid; and extending the plurality of indexing primers to obtain a plurality of indexed - adapter - target polynucleotides. In some embodiments, the universal adapter and the target nucleic acid are double - stranded, and hybridizing the plurality of indexing primers to the universal adapter includes hybridizing the plurality of indexing primers to only one strand of the universal adapter.

[0018] In some embodiments, the universal adapter comprises a double - stranded adapter. In some embodiments, the universal adapter comprises a Y - shaped adapter. In some embodiments, the universal adapter comprises a single - stranded adapter. In some embodiments, the universal adapter comprises a hairpin adapter. In some embodiments, each of the universal adapters comprises an overhang at one end to be attached to the nucleic acid before being attached to the nucleic acid. In some embodiments, each of the universal adapters comprises a blunt end to be attached to the nucleic acid before being attached to the nucleic acid.

[0019] In some embodiments, the universal adapter and the target nucleic acid are double - stranded, and hybridizing the plurality of indexing primers to the universal adapter includes hybridizing the plurality of indexing primers to both strands of the universal adapter. In some embodiments, a first indexing primer that hybridizes to the first strand of a particular universal adapter comprises a first indexing sequence selected from a first subset of the set of indexing sequences, and a second indexing primer that hybridizes to the second strand of the particular universal adapter comprises a second indexing sequence selected from a second subset of the set of indexing sequences.

[0020] In some embodiments, the first index sequence and the second index sequence are, respectively: the nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 or the reverse complement of SEQ ID NO:11; the nth 10-mer in SEQ ID NO:12 and the nth 10-mer in SEQ ID NO:13 or the reverse complement of SEQ ID NO:13; the nth 10-mer in SEQ ID NO:14 and the nth 10-mer in SEQ ID NO:15 or the reverse complement of SEQ ID NO:15; the nth 10-mer in SEQ ID NO:16 and the nth 10-mer in SEQ ID NO:17 or the reverse complement of SEQ ID NO:17; or the nth 10-mer in SEQ ID NO:18 and the nth 10-mer in SEQ ID NO:19 or the reverse complement of SEQ ID NO:19.

[0021] In some embodiments, the first subset comprises the index sequences listed in Table 1, and the second subset comprises the index sequences listed in Table 2.

[0022] In some embodiments, the index primer that hybridizes to both strands of the universal adaptor comprises an index sequence selected from the same subset of the set of index sequences. In some embodiments, this subset of index sequences is selected from one of the subsets of index sequences in Table 3.

[0023] In some embodiments, the method further comprises attaching a universal adaptor to one or both ends of the nucleic acid prior to step (a). In some embodiments, the attachment comprises attaching the universal adaptor by transposase-mediated fragmentation. In some embodiments, transposase-mediated fragmentation comprises: providing nucleic acid molecules obtained from a plurality of samples and a plurality of transpososome complexes, wherein each transpososome complex comprises a transposase and two transposon end compositions, and the transposon end compositions comprise the sequences of the universal adaptor; and obtaining a target nucleic acid, wherein the target nucleic acid comprises the sequences of the universal adaptor transposed from the transposon end compositions at one or both ends.

[0024] In some embodiments, the attachment comprises ligating the universal adaptor to one or both ends of the nucleic acid. In some embodiments, the ligation comprises enzymatic ligation or chemical ligation. In some embodiments, the chemical ligation comprises click chemical reaction ligation.

[0025] In some embodiments, the attachment is carried out by amplification with a target-specific primer comprising the sequence of the universal adaptor.

[0026] In some embodiments, the plurality of indexed polynucleotides comprise a sample-specific adaptor that comprises an index sequence from an index sequence set. In some embodiments, the sample-specific adaptor comprises an adaptor having two strands. In some embodiments, only one strand of the sample-specific adaptor comprises an index sequence. In some embodiments, each strand of the sample-specific adaptor comprises an index sequence. In some embodiments, the first strand of each sample-specific adaptor comprises a first index sequence selected from a first subset of the index sequence set, and the second strand of the sample-specific adaptor comprises a second index sequence selected from a second subset of the index sequence set.

[0027] In some embodiments, the first index sequence and the second index sequence are, respectively: the nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 or the reverse complement of SEQ ID NO:11; the nth 10-mer in SEQ ID NO:12 and the nth 10-mer in SEQ ID NO:13 or the reverse complement of SEQ ID NO:13; the nth 10-mer in SEQ ID NO:14 and the nth 10-mer in SEQ ID NO:15 or the reverse complement of SEQ ID NO:15; the nth 10-mer in SEQ ID NO:16 and the nth 10-mer in SEQ ID NO:17 or the reverse complement of SEQ ID NO:17; or the nth 10-mer in SEQ ID NO:18 and the nth 10-mer in SEQ ID NO:19 or the reverse complement of SEQ ID NO:19.

[0028] In some embodiments, the first subset comprises the index sequences listed in Table 1, and the second subset comprises the index sequences listed in Table 2.

[0029] In some embodiments, the first subset and the second subset are the same. In some embodiments, the subset of index sequences is selected from one of the subsets of index sequences in Table 3.

[0030] In some embodiments, each sample-specific adaptor comprises a flow cell amplification primer binding sequence. In some embodiments, the flow cell amplification primer binding sequence comprises a P5 sequence, a P5′ sequence, a P7 sequence, or a P7′ sequence.

[0031] In some embodiments, contacting a plurality of indexed polynucleotides with a target nucleic acid includes attaching a sample-specific adaptor to the target nucleic acid by transposome-mediated fragmentation. In some embodiments, transposome-mediated fragmentation includes: providing nucleic acid molecules obtained from a plurality of samples; providing a plurality of transposome complexes, wherein each transposome complex comprises a transposase and two transposon end compositions, and the transposon end compositions comprise sequences of the sample-specific adaptor; and obtaining a target nucleic acid, wherein the target nucleic acid comprises sequences of the sample-specific adaptor transposed from the transposon end compositions at one or both ends.

[0032] In some embodiments, contacting a plurality of indexed polynucleotides with a target nucleic acid includes ligating a sample-specific adaptor to the target nucleic acid. In some embodiments, ligation includes enzymatic ligation or chemical ligation. In some embodiments, chemical ligation includes click chemical reaction ligation.

[0033] In some embodiments, the sample-specific adaptor comprises a Y-shaped adaptor having complementary double-stranded regions and mismatched single-stranded regions. In some embodiments, each strand of the sample-specific adaptor comprises an index sequence at the mismatched single-stranded region. In some embodiments, only one strand of the sample-specific adaptor comprises an index sequence at the mismatched single-stranded region.

[0034] In some embodiments, the sample-specific adaptor comprises a single-stranded adaptor.

[0035] In some embodiments, the sample-specific adaptor comprises a hairpin adaptor.

[0036] In some embodiments, contacting a plurality of indexed polynucleotides with a target nucleic acid includes attaching a plurality of indexed polynucleotides to both ends of the target nucleic acid.

[0037] In some embodiments, contacting a plurality of indexed polynucleotides with a target nucleic acid includes attaching a plurality of indexed polynucleotides to only one end of the target nucleic acid.

[0038] In some embodiments, the method further includes amplifying the pooled indexed-target polynucleotides before sequencing the pooled indexed-target polynucleotides.

[0039] In some embodiments, the method further includes fragmenting the nucleic acid molecules obtained from a plurality of samples to obtain target nucleic acids before step (a). In some embodiments, fragmentation includes transposome-mediated fragmentation. In some embodiments, transposome-mediated fragmentation includes: providing nucleic acid molecules and a plurality of transposon complexes, wherein each transposon complex comprises a transposase and two transposon end compositions; and obtaining target nucleic acids, wherein the target nucleic acids comprise sequences transposed from the transposon end compositions at one or both ends.

[0040] In some embodiments, fragmentation includes contacting multiple PCR primers that target a sequence of interest to obtain a target nucleic acid that contains the sequence of interest.

[0041] In some embodiments, the set of index sequences includes multiple non-overlapping subsets of index sequences, and the Hamming distance between any two index sequences in any subset is not less than a second standard value. In some embodiments, the second standard value is greater than the first standard value. In some embodiments, the first standard value is 4, and the second standard value is 5. In some embodiments, the first standard value is 3. In some embodiments, the first standard value is 4.

[0042] In some embodiments, the edit distance between any two index sequences in the set of index sequences is not less than a third standard value. In some embodiments, the edit distance is a modified Levenshtein distance where terminal gaps are not assigned a penalty. In some embodiments, the third standard value is 3. In some embodiments, each index sequence in the set of index sequences has 10 bases; the first standard value is 4; and the third standard value is 3.

[0043] In some embodiments, the set of index sequences includes the 10-mers in SEQ ID NO:9.

[0044] In some embodiments, each index sequence has 32 or fewer bases. In some embodiments, each index sequence has 10 or fewer bases. In some embodiments, each index sequence has 6 to 8 bases.

[0045] In some embodiments, the set of index sequences does not include index sequences that are empirically determined to have poor performance in indexing the origin of nucleic acid samples in multiplexed massively parallel sequencing. In some embodiments, the index sequences that are not included contain the sequences in Table 4.

[0046] In some embodiments, the set of index sequences does not include any subsequence of the sequence of an adaptor or primer in a sequencing platform, or the reverse complement of the subsequence. In some embodiments, the sequence of an adaptor or primer in a sequencing platform includes SEQ ID NO:1 (AGATGTGTATAAGAGACAG), SEQ ID NO:3 (TCGTCGGCAGCGTC), SEQ ID NO:5 (CCGAGCCCACGAGAC), SEQ ID NO:7 (CAAGCAGAAGACGGCATACGAGAT), and SEQ ID NO:8 (AATGATACGGCGACCACCGAGATCTACAC).

[0047] In some embodiments, each index sequence in the index sequence set has a guanine / cytosine (GC) content between 25% and 75%.

[0048] In some embodiments, the index sequence set comprises at least 12 different index sequences. In some embodiments, the index sequence set comprises at least 20 different index sequences. In some embodiments, the index sequence set comprises at least 24, at least 28, at least 48, at least 80, at least 96, at least 112, or at least 384 different index sequences.

[0049] In some embodiments, the index sequence set does not include any homopolymers having four or more consecutive identical bases.

[0050] In some embodiments, the index sequence set does not include index sequences that match or are reverse complementary to one or more sequencing primer sequences. In some embodiments, the sequencing primer sequences are included in the sequences of the plurality of index polynucleotides.

[0051] In some embodiments, the index sequence set does not include index sequences that match or are reverse complementary to one or more flow cell amplification primer sequences. In some embodiments, the flow cell amplification primer sequences are included in the sequences of the plurality of index polynucleotides.

[0052] In some embodiments, the index sequence set comprises index sequences having the same number of bases.

[0053] In some embodiments, each index sequence in the index sequence set has a combined number of guanine and cytosine bases that is not less than 2 and not greater than 6.

[0054] In some embodiments, the plurality of index polynucleotides comprises DNA or RNA.

[0055] Another aspect of the present disclosure relates to methods for sequencing target nucleic acids derived from multiple samples. The method includes: (a) providing a plurality of double-stranded nucleic acid molecules derived from multiple samples; (b) providing a plurality of transposome complexes, wherein each transposome complex comprises a transposase and two transposon end compositions; (c) incubating the double-stranded nucleic acid molecules with the transposome complexes to obtain double-stranded nucleic acid fragments, wherein the double-stranded nucleic acid fragments comprise sequences transposed from the transposon end compositions at one or both ends; (d) contacting a plurality of indexing primers with the double-stranded nucleic acid fragments to generate a plurality of index-fragment polynucleotides, wherein the indexing primers contacting the double-stranded nucleic acid fragments derived from each sample comprise an index sequence or a combination of index sequences uniquely associated with this sample, and the index sequence or the combination of index sequences is selected from a set of index sequences; (e) pooling the plurality of index-fragment polynucleotides; (f) sequencing the pooled index-fragment polynucleotides to obtain index reads of the index sequences and a plurality of target reads of the target sequences, each target read being associated with at least one index read; and (g) using the index reads to determine the sample origin of the target reads.

[0056] In some embodiments, the Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, wherein the first standard value is at least 2. In some embodiments, the set of index sequences comprises multiple pairs of color-balanced index sequences, wherein any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base.

[0057] In some embodiments, contacting the plurality of indexing primers with the double-stranded nucleic acid fragments includes: hybridizing the plurality of indexing primers to the sequences transposed from the transposon end compositions at one or both ends of the double-stranded nucleic acid fragments; and extending the plurality of indexing primers to obtain index-fragment polynucleotides. In some embodiments, the hybridization includes hybridizing the plurality of indexing primers to only one strand of the double-stranded nucleic acid fragments. In some embodiments, the hybridization includes hybridizing the plurality of indexing primers to both strands of the double-stranded nucleic acid fragments.

[0058] In some embodiments, a first indexing primer that hybridizes to a first strand of a specific double-stranded nucleic acid fragment comprises a first indexing sequence selected from a first subset of an indexing sequence set, and a second indexing primer that hybridizes to a second strand of the specific double-stranded nucleic acid fragment comprises a second indexing sequence selected from a second subset of the indexing sequence set. In some embodiments, the first indexing sequence and the second indexing sequence are, respectively: the nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 or the reverse complement of SEQ ID NO:11; the nth 10-mer in SEQ ID NO:12 and the nth 10-mer in SEQ ID NO:13 or the reverse complement of SEQ ID NO:13; the nth 10-mer in SEQ ID NO:14 and the nth 10-mer in SEQ ID NO:15 or the reverse complement of SEQ ID NO:15; the nth 10-mer in SEQ ID NO:16 and the nth 10-mer in SEQ ID NO:17 or the reverse complement of SEQ ID NO:17; or the nth 10-mer in SEQ ID NO:18 and the nth 10-mer in SEQ ID NO:19 or the reverse complement of SEQ ID NO:19. In some embodiments, the first subset comprises the indexing sequences listed in Table 1, and the second subset comprises the indexing sequences listed in Table 2.

[0059] In some embodiments, the first subset and the second subset are the same. In some embodiments, the first subset or the second subset of the indexing sequences is selected from one of the subsets of the indexing sequences in Table 3.

[0060] In some embodiments, the indexing sequence set comprises at least 6 different indexing sequences.

[0061] In some embodiments, each indexing primer comprises an amplification primer binding sequence.

[0062] In some embodiments, the flow cell amplification primer binding sequence comprises a P5 sequence or a P7′ sequence.

[0063] In some embodiments, at least one of the transposome complexes comprises a Tn5 transposase and a Tn5 transposon end composition.

[0064] In some embodiments, at least one of the transposome complexes comprises a Mu transposase and a Mu transposon end composition.

[0065] Also provided are systems, devices, and computer program products for identifying and preparing indexing oligonucleotides, and for determining the sequences of DNA fragments using the disclosed indexing sequences.

[0066] Additional aspects of the present disclosure relate to a computer program product comprising a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for sequencing target nucleic acids from multiple samples. The program code includes: (a) code for receiving a plurality of indexed reads and a plurality of target reads of target sequences obtained from target nucleic acids from multiple samples. Each target read comprises a target sequence obtained from a target nucleic acid of a sample from among the multiple samples. Each indexed read comprises an index sequence obtained from a target nucleic acid of a sample from among the multiple samples, the index sequence being selected from a set of index sequences. Each target read is associated with at least one indexed read. Each of the multiple samples is uniquely associated with one or more index sequences from the set of index sequences. The Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, where the first standard value is at least 2. The program code further comprises: (b) code for identifying, in the plurality of target reads, a subset of target reads associated with an indexed read that matches at least one index sequence uniquely associated with a particular sample among the multiple samples; and (c) code for determining the target sequence of the particular sample based on the identified subset of target reads.

[0067] In some embodiments, the set of index sequences comprises multiple pairs of color-balanced index sequences, where any two bases at corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base.

[0068] In some embodiments, the computer program product comprises a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement one or more of the methods described above.

[0069] Additional aspects of the present disclosure relate to a computer system comprising: one or more processors; a system memory; and one or more computer-readable storage media having stored thereon computer-executable instructions that cause the computer system to implement a method for sequencing nucleic acids in a plurality of samples. The instructions include: (a) receiving a plurality of indexed reads and a plurality of target reads of target sequences obtained from target nucleic acids derived from the plurality of samples. Each target read comprises a target sequence obtained from a target nucleic acid of a sample from the plurality of samples. Each indexed read comprises an index sequence obtained from a target nucleic acid of a sample from the plurality of samples, the index sequence selected from a set of index sequences. Each target read is associated with at least one indexed read. Each sample of the plurality of samples is uniquely associated with one or more index sequences from the set of index sequences. The Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, wherein the first standard value is at least 2. The instructions further include: (b) identifying, in the plurality of target reads, a subset of target reads associated with an indexed read that matches at least one index sequence uniquely associated with a particular sample of the plurality of samples; and (c) determining the target sequence of the particular sample based on the identified subset of target reads. In some embodiments, the set of index sequences comprises a plurality of pairs of color-balanced index sequences, wherein any two bases at corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base.

[0070] In some embodiments, one or more computer-readable storage media have stored thereon computer-executable instructions that cause the computer system to implement one or more of the methods described above.

[0071] Although the examples herein relate to humans and the language is primarily directed to human problems, the concepts described herein apply to nucleic acids from any virus, plant, animal, or other organism, and populations thereof (metagenomes, viromes, etc.). These and other features of the present disclosure will become more fully apparent from the following description with reference to the drawings and the appended claims, or may be learned by the practice of the present disclosure as set forth hereinafter.

[0072] Incorporation by reference

[0073] All patents, patent applications, and other publications mentioned herein (including all sequences disclosed within these references) are hereby expressly incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. All cited documents are incorporated by reference in their entirety into the relevant parts of this document for the purposes indicated by the context in which they are cited herein. However, the citation of any document should not be construed as an admission that it is prior art to the present disclosure. Brief Description of the Drawings

[0075] Figures 1A - 1C Illustrates an exemplary workflow for sequencing nucleic acid fragments using indexed oligonucleotides.

[0076] Figure 1D Illustrates a process for sequencing target nucleic acids derived from multiple samples according to some embodiments.

[0077] Figure 1E and Figure 1F Illustrates a process of performing transposome-mediated fragmentation and applying indexed primers to nucleic acids having double-stranded short universal adapters attached to both ends.

[0078] Figure 1G Illustrates the sequences of target nucleic acids having double-stranded short universal adapters attached to both ends according to some embodiments.

[0079] Figure 1H Illustrates the sequences of target nucleic acids having Y-shaped short universal adapters attached to both ends according to some embodiments.

[0080] Figure 1I Illustrates the sequences in the i7 indexed primers according to some embodiments.

[0081] Figure 1J Illustrates the sequences in the i5 indexed primers according to some embodiments.

[0082] Figure 1K Illustrates a process of adding indexed sequences to nucleic acids having Y-shaped short universal adapters at both ends according to some embodiments.

[0083] Figures 2A - 2D Illustrates multiple embodiments of indexed oligonucleotides.

[0084] Figure 3 Schematically illustrates an indexed sequence design that provides a mechanism for detecting errors that occur in the indexed sequences during the sequencing process.

[0085] Figures 4A - 4CSchematically illustrates a multi-well plate in which index oligonucleotides can be provided and an exemplary layout of the index oligonucleotides.

[0086] Figure 5 Illustrates a process for preparing index oligonucleotides such as indexed adapters.

[0087] Figure 6 Illustrates one embodiment of a dispersed system for generating calls or diagnoses from test samples.

[0088] Figure 7 The figure illustrates a computer system that can be used as a computing device according to certain embodiments. Detailed description

[0090] Numeric ranges include the numbers defining the range. It is expected that every maximum numeric limitation given throughout this specification will include every lower numeric limitation, as if such lower numeric limitations were expressly written herein. Every minimum numeric limitation given throughout this specification will include every higher numeric limitation, as if such higher numeric limitations were expressly written herein. Every numeric range given throughout this specification will include every narrower numeric range that falls within such broader numeric range, as if such narrower numeric ranges were expressly written herein.

[0091] The headings provided herein are not intended to limit the disclosure.

[0092] Unless defined otherwise herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Various scientific dictionaries including the terms contained herein are well known and available to those of skill in the art. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the embodiments disclosed herein, some methods and materials are described.

[0093] The terms defined immediately below are more fully described by reference to the entire specification. It should be understood that the disclosure is not limited to the specific methodologies, protocols, and reagents described, as these may vary depending on the context in which those skilled in the art use them.

[0094] Definitions

[0095] As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly indicates otherwise.

[0096] As used herein, and where context is appropriate and unless otherwise indicated, the term “include” encompasses the meanings of “comprise,” “consist of,” or “consist essentially of.”

[0097] Unless otherwise indicated, nucleic acids are written left to right in the 5′ to 3′ direction, and amino acid sequences are written left to right in the amino to carboxy direction, respectively.

[0098] Edit distance is a measure of how different two strings (e.g., words) are from each other by counting the minimum number of operations required to transform one string into the other. In bioinformatics, it can be used to quantify the similarity of DNA sequences, which can be viewed as strings of the letters A, C, G, and T.

[0099] Different forms of edit distance use different sets of string operations. The Levenshtein distance is a common type of edit distance. The string operations of the Levenshtein distance account for the number of character deletions, insertions, and substitutions in a string. In some embodiments, other variants of edit distance can be used. For example, other variants of edit distance can be obtained by restricting the set of operations. The longest common subsequence (LCS) distance is an edit distance that has insertions and deletions as the only two edit operations (both with unit cost). The Jaro-Winkler distance can be obtained from an edit distance that only allows transpositions. Similarly, by only allowing substitutions, the Hamming distance is obtained, which is restricted to strings of equal length. The Hamming distance between two strings of equal length is the number of positions where the corresponding symbols are different. In other words, it measures the minimum number of substitutions required to transform one string into the other, or the minimum number of errors required to transform one string into the other.

[0100] In some embodiments, for edit distance, different string operations can have different weights. For example, a substitution operation can have a weight of 3, while an insertion / deletion (indel) can have a weight of 2. In some embodiments, different types of matches can have different weights. For example, an A-A match can have twice the weight of a G-G match.

[0101] As used herein, the term "common sequence" refers to a sequence region shared by two or more nucleic acid molecules (e.g., adapter-target-adapter molecules), where the molecules also have sequence regions that are different from each other. A common sequence present in different members of a collection of molecules can allow for the capture of multiple different nucleic acids using a population of common capture nucleic acids that are complementary to a portion of the common sequence (e.g., a common extension primer binding site). Non-limiting examples of common extension primer binding sites include sequences that are identical or complementary to the P5 primer and the P7 primer. Similarly, a common sequence present in different members of a collection of molecules can allow for the replication or amplification of multiple different nucleic acids using a population of common primers that are complementary to a portion of the common sequence (e.g., a common primer binding site). Thus, common capture nucleic acids or common primers include sequences that can hybridize specifically to a common sequence. Target nucleic acid molecules can be modified, for example, to attach adapters at one or both ends of different target sequences, as described herein.

[0102] When referring to amplification primers, such as common primer extension primers, the terms "P5" and "P7" can be used. The terms "P5′" (P5 prime) and "P7′" (P7 prime) refer to the complements of P5 and P7, respectively. It will be understood that any suitable amplification primers can be used in the methods presented herein, and the use of P5 and P7 is merely an exemplary embodiment. The use of amplification primers such as P5 and P7 on flow cells is known in the art, as exemplified by the disclosures of WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957. For example, any suitable forward amplification primer, whether immobilized or in solution, can be useful for hybridizing to a complementary sequence and amplifying a sequence in the methods presented herein. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, can be useful for hybridizing to a complementary sequence and amplifying a sequence in the methods presented herein. Those skilled in the art will understand how to design and use primer sequences suitable for capturing and amplifying nucleic acids as presented herein.

[0103] The terms "upstream" and "5′-of" with respect to a position in a nucleic acid sequence are used interchangeably and refer to a relative position in the nucleic acid sequence that is further toward the 5′ end of the sequence.

[0104] The terms "downstream" and "3′-of" with respect to a position in a nucleic acid sequence are used interchangeably and refer to a relative position in the nucleic acid sequence that is further toward the 3′ end of the sequence.

[0105] One step in some embodiments of the methods of the present disclosure is to fragment and tag a target DNA using an in vitro transposition reaction to generate tagged DNA fragments. The in vitro transposition reaction requires a transposase, a transposon end composition, and suitable reaction conditions.

[0106] "Transposase" means an enzyme that is capable of forming a functional complex with a composition containing transposon ends (e.g., a transposon, a transposon end, a transposon end composition) and catalyzing the insertion or transposition of the composition containing transposon ends into a double-stranded target DNA incubated therewith in an in vitro transposition reaction. Transposases also include integrases from retrotransposons and retroviruses.

[0107] "Transposition reaction" is a reaction in which one or more transposon ends are inserted into target DNA at random or near-random sites. In some embodiments, the transposition reaction causes fragmentation of the target DNA or RNA at random positions. Important components in the transposition reaction are the transposase and DNA oligonucleotides that exhibit the nucleotide sequences of the transposon ends (including the transferred transposon end sequences and their complements, non-transferred transposon end sequences), as well as other components required to form a functional transpososome complex. The methods of the present invention are exemplified by using a transpososome complex formed by a hyperactive Tn5 transposase and Tn5-type transposon ends (Goryshin, I. and Reznikoff, W.S., J. Biol. Chem., 273:7367, 1998) or a transpososome complex formed by MuA transposase and Mu transposon ends comprising R1 and R2 end sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H, et al., EMBO J., 14:4893, 1995). However, any transposon system capable of inserting transposon ends in a random or near-random manner and having sufficient efficiency to 5'-tag and fragment the target DNA for its intended purpose can be used in the present invention.Examples of transposon systems known in the art that can be applied include, but are not limited to, Staphylococcus aureus Tn552 (Colegio O R et al., J Bacteriol., 183:2384-8, 2001; Kirby C et al., Mol Microbiol., 43:173-86, 2002), Ty1 (Devine S E and Boeke J D., Nucleic Acids Res., 22:3765-72, 1994 and International Patent Application No. WO 95 / 23875), transposon Tn7 (Craig, N L, Science, 271:1512, 1996; Craig, N L, Review in: Curr Top Microbiol Immunol., 204:27-48, 1996), Tn10 and IS10 (Kleckner N, et al., Curr Top Microbiol Immunol., 204:49-82, 1996), Mariner transposase (Lampe D J, et al., EMBO J., 15:5470-9, 1996), Tc1 (Plasterk R H, Curr Top Microbiol Immunol, 204:125-43, 1996), P element (Gloor, G B, Methods Mol Biol., 260:97-114, 2004), Tn3 (Ichikawa H, and Ohtsubo E., J Biol Chem. 265:18829-32, 1990), bacterial insertion sequences (Ohtsubo, F and Sekine, Y, Curr.Top.Microbiol.Immunol. 204:1-26, 1996), retroviruses (Brown P O, et al., Proc Natl Acad Sci USA, 86:2525-9, 1989) and yeast retrotransposons (Boeke J D and Corces V G, Annual Rev Microbiol. 43:403-34, 1989).

[0108] Methods for inserting transposon ends into target sequences can be performed in vitro using any suitable transposon system for which suitable in vitro transposition systems are available or can be developed based on knowledge in the art. Generally, suitable in vitro transposition systems for use in the methods of the invention require: a transposase having sufficient purity, sufficient concentration, and sufficient in vitro transposition activity; and transposon ends with which the transposase forms a functional complex that has transposase capable of catalyzing the transposition reaction for each. Suitable transposon end sequences that can be used in the present invention include, but are not limited to, wild-type transposon end sequences, derived transposon end sequences, or mutated transposon end sequences that form a complex with a transposase selected from wild-type transposase, a derived form of transposase, or a mutated form of transposase. Exemplary transposases include wild-type or mutated forms of Tn5 transposase and MuA transposase (although the EZ-Tn5 transposase is significantly more effective than an equivalent amount of MuA transposase in generating 5'-tagged DNA fragments in the methods of the present invention), but any other transposase for which compositions and conditions for efficient in vitro transposition for defined transposon ends are known or subsequently developed can be used in the methods of the present invention. Transposon end sequences recognized by wild-type or mutated forms of Tn5 transposase or MuA transposase are suitable in some embodiments, and those transposon end sequences that result in the highest transposition efficiency when complexed with the transposase, along with the corresponding optimally active transposase complexed therewith, are advantageous for some embodiments. In some embodiments, a transposon is selected in which the transposon end sequence required by the transposase for transposition is not too large and the transposon end sequence has the smallest possible size that functions well for the intended purpose and has a size sufficient such that the same sequence is rarely or not at all present in the target DNA or sample DNA. For example, the EZ-Tn5 TM transposon end sequence of the transposon end sequence contains only 19 nucleotides, while some other transposases require much larger end sequences for transposition (e.g., the MuA transposase requires a transposon end sequence of approximately 51 nucleotides).

[0109] Suitable in vitro transposition systems that can be used to insert transposon ends into target nucleic acids include, but are not limited to, those using the EZ-Tn5 TM Ultra-Active Tn5 transposase available from EPICENTRE Technologies, Madison, WI, or HyperMu TM Ultra-Active MuA transposase from EPICENTRE, or other in vitro transposition systems such as another MuA transposase available from Finnzymes Oy, Espoo, Finland.

[0110] In some embodiments, insertion of transposon ends into target DNA according to the invention can also be carried out in vivo. If transposition is carried out in vivo, transposition into target DNA is preferably achieved by electroporating a synaptic complex of a transposase and a suitable transposon end composition into a host cell, as described in U.S. Patent No. 6,159,736 (incorporated herein by reference). This transposition method is exemplified by the use of the following transposon complexes: a transposon complex formed by a hyperactive Tn5 transposase and a suitable Tn5-type transposon end composition using a method similar to that described by (Goryshin, I. and Reznikoff, W.S., J. Biol. Chem., 273:7367, 1998), or by HyperMu TM a transposon complex formed by a hyperactive MuA transposase (EPICENTRE, Madison, Wis.) and a suitable MuA transposon end composition that exhibits R1 end sequences and R2 end sequences recognized by the transposase. A suitable synaptic complex or "Transposome TM complex (EPICENTRE)" between the transposon end composition and the transposase can be prepared as described in U.S. Patent No. 6,159,736 and the related patents of Goryshin and Reznikoff, or as described in the product literature for the Tn5-type EZ-Tn5 TM Transposome TM complex or HyperMu TM MuA Transposome TM complex.

[0111] The term "transposon end" means double-stranded DNA that exhibits only the nucleotide sequences ("transposon end sequences") necessary to form a complex with a transposase or integrase that is functional in an in vitro transposition reaction. The transposon end forms a "complex" or "synaptic complex" or "transpososome complex" or "transpososome composition" with a transposase or integrase that recognizes and binds to the transposon end, and the complex is capable of inserting or transposing the transposon end into a target DNA with which it is incubated in an in vitro transposition reaction. The transposon end exhibits two complementary sequences consisting of a "transferred transposon end sequence" or "transfer strand" and a "non-transferred transposon end sequence" or "non-transfer strand". For example, with a hyperactive Tn5 transposase (e.g., EZ-Tn5) that is active in an in vitro transposition reaction TMA transposon end that forms a complex with a transposase (EPICENTRE Biotechnologies, Madison, Wis., USA) includes a transfer strand that exhibits a "transferred transposon end sequence" as follows:

[0112] 5′AGATGTGTATAAGAGACAG3′ (SEQ ID NO:1)

[0113] and a non-transfer strand that exhibits a "non-transferred transposon end sequence" as follows:

[0114] 5′CTGTCTCTTATACACATCT3′ (SEQ ID NO:2)

[0115] The term "pMETS" refers to a 19-base single-stranded transposon end oligonucleotide containing a 5′-phosphate that exhibits the EZ-Tn5 TM transposon end sequence:

[0116] 5′pAGATGTGTATAAGAGACAG3′ (SEQ ID NO:1)

[0117] The term "METS" refers to a 19-base single-stranded transposon end oligonucleotide that exhibits the EZ-Tn5 TM transposon end sequence:

[0118] 5′AGATGTGTATAAGAGACAG3′ (SEQ ID NO:1)

[0119] The term "pMENTS" refers to a 19-base single-stranded transposon end oligonucleotide containing a 5′-phosphate that exhibits the EZ-Tn5 TM transposon end sequence:

[0120] 5′pCTGTCTCTTATACACATCT3′ (SEQ ID NO:2)

[0121] The term "pMEDS" refers to a 19-base pair double-stranded EZ-Tn5 TM transposon end, where the two 5′ ends contain phosphates:

[0122] 5′pAGATGTGTATAAGAGACAG3′ (SEQ ID NO:1)

[0123] 3′TCTACACATATTCTCTGTCp5′ (SEQ ID NO:2)

[0124] pMEDS EZ-Tn5 TMThe transposon ends are prepared by annealing pMETS transposon end oligonucleotides to pMENTS transposon end oligonucleotides.

[0125] The designation “MEDS” refers to a 19 base pair double-stranded EZ-Tn5 TM transposon end, where only the non-transfer strand (pMENTS) contains a 5′-phosphate:

[0126] 5′AGATGTGTATAAGAGACAG3′ (SEQ ID NO:1)

[0127] 3′TCTACACATATTCTCTGTCp5′ (SEQ ID NO:2)

[0128] MEDS EZ-Tn5 TM The transposon ends are prepared by annealing METS transposon end oligonucleotides to pMENTS transposon end oligonucleotides.

[0129] The 3′ end of the transfer strand is ligated or transferred to the target DNA in an in vitro transposition reaction. The non-transfer strand that exhibits a transposon end sequence complementary to the transferred transposon end sequence is not ligated or transferred to the target DNA in the in vitro transposition reaction.

[0130] In some embodiments, the transfer strand and the non-transfer strand are covalently linked. For example, in some embodiments, the transfer strand sequence and the non-transfer strand sequence are provided on a single oligonucleotide, such as in the form of a hairpin configuration. Thus, although the free end of the non-transfer strand is not directly ligated to the target DNA by the transposition reaction, the non-transfer strand becomes indirectly attached to the DNA fragment because the non-transfer strand is linked to the transfer strand through the loop of the hairpin structure.

[0131] “Transposon end composition” means a composition comprising a transposon end (i.e., the smallest double-stranded DNA segment capable of interacting with a transposase to undergo a transposition reaction), optionally plus one or more additional sequences at the 5' of the transferred transposon end sequence and / or at the 3' of the non-transferred transposon end sequence. For example, a transposon end attached to a tag is a “transposon end composition”. In some embodiments, the transposon end composition comprises two transposon end oligonucleotides consisting of a “transferred transposon end oligonucleotide” or “transfer strand” and a “non-transfer strand end oligonucleotide” or “non-transfer strand” or is composed of the two transposon end oligonucleotides, the “transferred transposon end oligonucleotide” or “transfer strand” and the “non-transfer strand end oligonucleotide” or “non-transfer strand” combinatorially exhibit the sequence of the transposon end, and one or both strands contain additional sequences.

[0132] The terms "transferred transposon end oligonucleotide" and "transferred strand" are used interchangeably and refer to the transferred portion of both "transposon end" and "transposon end composition", i.e., regardless of whether the transposon end is attached to a tag or other moiety. Similarly, the terms "non-transferred transposon end oligonucleotide" and "non-transferred strand" are used interchangeably and refer to the non-transferred portion of both "transposon end" and "transposon end composition". In some embodiments, the transposon end composition is a "hairpin transposon end composition".

[0133] As used herein, a "hairpin transposon end composition" means a transposon end composition consisting of a single oligodeoxyribonucleotide that exhibits a non-transferred transposon end sequence at its 5' end, a transferred transposon end sequence at its 3' end, and an intervening arbitrary sequence between the non-transferred transposon end sequence and the transferred transposon end sequence, the intervening arbitrary sequence being long enough to allow intramolecular stem-loop formation such that the transposon end portion can function in a transposition reaction. In some embodiments, the 5' end of the hairpin transposon end composition has a phosphate group at the 5' position of the 5' nucleotide. In some embodiments, the intervening arbitrary sequence between the non-transferred transposon end sequence and the transferred transposon end sequence of the hairpin transposon end composition provides a tag (e.g., including one or more tag domains) for a particular use or application.

[0134] In some embodiments, the methods of the present disclosure produce tagged circular ssDNA fragments. In some embodiments, the tagged circular ssDNA fragments exhibit only the sequence of the transferred strand of the transposon end composition and do not exhibit the sequence of the non-transferred strand of the transposon end composition.

[0135] In some embodiments, the transposon end oligonucleotides used in the methods of the invention exhibit only the transposon end sequences required in the transposition reaction. However, in some embodiments, at least one of the transposon end oligonucleotides additionally exhibits one or more other nucleotide sequences 5′ of the transposon end sequence. Thus, in some embodiments, the method uses a transfer strand having a 3′ portion and a 5′ portion, wherein the 3′ portion exhibits the transposed transposon end sequence and the 5′ portion exhibits one or more additional sequences that do not participate in forming a functional complex with the transposase. There are no limitations on which additional sequences can be used for the one or more additional sequences in the 5′ portion of the transfer strand, and these sequences can be used to achieve any desired purpose. For example, in some embodiments, the 5′ portion of the transfer strand exhibits one or more additional tag sequences. In some embodiments, the tag sequence can be an index sequence associated with a particular sample. In some embodiments, the tag sequence allows capture by annealing to a specific sequence on a surface. In some embodiments, the tag sequence allows 5′-tagged target fragments to be captured on a flow cell substrate for next-generation sequencing; e.g., for P5 or P7′ tags captured on the flow cell of an Illumina sequencing platform, or for 454A or 454B tag sequences captured on beads for sequencing using a Roche 454 next-generation sequencer.

[0136] In some embodiments, the tag sequence can be one or more sequences for the identification, detection (e.g., fluorescence detection), or sorting of the products of the method. In some other embodiments, the 5′ portion of the transfer strand exhibits one or more additional nucleotides or sequences or chemical groups or moieties that comprise or consist of an affinity binding molecule, such as a tag sequence that allows capture by annealing to a specific sequence on a surface such as beads or a probe on a microchip or array. In some preferred embodiments, the size of the one or more additional sequences in the 5′ portion of the transfer strand is minimized to minimize the probability or frequency of the transfer strand inserting into itself during an in vitro transposase reaction. For example, in some embodiments, the 5′ portion of the transfer strand is less than about 150 nucleotides, less than about 100 nucleotides, less than about 75 nucleotides, less than about 50 nucleotides, less than about 25 nucleotides, or less than about 15 nucleotides in size.

[0137] In some embodiments, the 5′ end of the transfer strand has a 5′ monophosphate group. In some embodiments, both the transfer strand and the non-transfer strand have a 5′ monophosphate group. In some preferred embodiments, only the 5′ end of the non-transfer strand has a 5′ monophosphate group. In some other embodiments, there is no 5′ monophosphate group at the 5′ end of the transfer strand.

[0138] In some embodiments, the transposon end composition used in the methods of the present disclosure comprises a transposon end oligonucleotide that exhibits only the transposon end sequences that form a complex with a transposase or integrase and that are required for the transposition reaction; in these embodiments, the tags in the tagged circular ssDNA fragments generated using the methods exhibit only the transferred transposon end sequences. However, in some embodiments, the transposon end composition comprises or consists of at least one transposon end oligonucleotide that exhibits one or more additional nucleotide sequences in addition to the transposon end sequences. Thus, in some embodiments, the transposon end composition comprises a transfer strand that exhibits one or more additional nucleotide sequences 5′ of the transferred transposon end sequence, and the one or more additional nucleotide sequences are also exhibited by the tag. Thus, in addition to the transferred transposon end sequence, the tag can have one or more additional tag portions or tag domains.

[0139] As used herein, a "tag" is a nucleic acid sequence that is associated with or can be associated with one or more nucleic acid molecules.

[0140] As used herein, a "tag portion" or "tag domain" means a portion or domain of a tag that exhibits a sequence for a desired purpose or application. A tag portion or tag domain is a "transposon end domain" that exhibits the transferred transposon end sequence. In some embodiments, where the transfer strand also exhibits one or more additional nucleotide sequences 5' of the transferred transposon end sequence, the tag also has one or more additional "tag domains" in the 5' portion, each of the tag domains being provided for any desired purpose. For example, some embodiments of the present disclosure comprise or consist of a transposon end composition that comprises or consists of: (i) a transfer strand that exhibits one or more sequences 5′ of the transferred transposon end sequence, the one or more sequences comprising or consisting of a tag domain selected from one or more of a sample-specific index sequence, a primer binding sequence, a restriction site tag domain, a capture tag domain, a sequencing tag domain, an amplification tag domain, a detection tag domain, and a transcription promoter domain; and (ii) a non-transfer strand that exhibits a non-transferred transposon end sequence. The present disclosure includes embodiments of methods using any one or more of the transposon end compositions described above.

[0141] In some embodiments, the transposon end composition comprises a transfer strand that includes a primer binding sequence that is reverse complementary to a sequence in a PCR primer. In some embodiments, the PCR primer is an indexing primer that includes a sample-specific indexing sequence. In some embodiments, after the transfer strand is transposed and attached to the target polynucleotide, the sample-specific indexing primer hybridizes to the primer binding sequence in the transfer strand attached to the target polynucleotide.

[0142] As used herein, "restriction site tag domain" or "restriction site domain" means a tag domain that exhibits a sequence for the purpose of facilitating cleavage using a restriction endonuclease. For example, in some embodiments, the restriction site domain is used to generate di-tagged linear ssDNA fragments. In some embodiments, the restriction site domain is used to generate compatible double-stranded 5' ends in the tag domain such that the ends can be ligated to another DNA molecule using a template-dependent DNA ligase. In some preferred embodiments, the restriction site domain in the tag exhibits the sequence of a restriction site that is rarely present (if at all) in the target DNA (e.g., the restriction site of a rare-cutting restriction endonuclease such as NotI or AscI). In some preferred embodiments, the restriction site in the restriction site domain is for a type II restriction endonuclease, such as the FokI restriction endonuclease.

[0143] In some embodiments in which the transfer strand of the transposon end composition comprises one or more restriction site domains at the 5'- of the transferred transposon end sequence, the method further comprises: annealing an oligodeoxynucleotide complementary to the single-stranded restriction site of the tagged circular ssDNA fragment, and then cleaving the tagged circular ssDNA fragment at the restriction site using a restriction endonuclease that recognizes the restriction site. Thus, in some embodiments, the method includes linearizing the tagged circular ssDNA fragment to generate a di-tagged linear ssDNA fragment.

[0144] In some other embodiments in which the transfer strand of the transposon end composition comprises one or more restriction site domains at the 5'- of the transferred transposon end sequence, the transfer strand of the transposon end composition comprises a double-stranded hairpin that includes a restriction site, and the method further comprises the step of cleaving the tagged linear ssDNA fragment at the restriction site using a restriction endonuclease that recognizes the restriction site; however, in some embodiments, the method is not preferred because the double-stranded hairpin provides a site in the dsDNA into which the transposon end composition can be transposed by a transposase or integrase.

[0145] In some preferred embodiments, it includes (i) generating a double-stranded restriction site by annealing an oligodeoxynucleotide complementary to a single-stranded restriction site or by using a transfer strand comprising a double-stranded hairpin, and (ii) then cleaving the restriction site using a restriction endonuclease that recognizes the double-stranded restriction site. The method further includes the step of ligating the tagged linear ssDNA fragment cleaved by the restriction endonuclease to another DNA molecule having a compatible 3′ end.

[0146] As used herein, "capture tag domain" or "capture tag" means a tag domain of a sequence that serves the purpose of facilitating the capture of an ssDNA fragment to which the tag domain is to be ligated (e.g., providing an annealing site or an affinity tag for capturing a tagged circular ssDNA fragment or a double-tagged linear ssDNA fragment on a bead or other surface, such as where the annealing site of the tag domain sequence allows capture by annealing to a specific sequence on the surface, such as a probe on a bead, or on a microchip or microarray, or on a sequencing bead). In some embodiments, the "capture tag" contains a flow cell amplification primer binding sequence. In some embodiments, the flow cell amplification primer binding sequence includes a P5 sequence or a P7′ sequence. In some embodiments of the method, after the tagged circular ssDNA fragment or the double-tagged linear ssDNA fragment is captured by annealing to a complementary probe on the surface, the capture tag domain provides a site for initiating DNA synthesis using the tagged circular ssDNA fragment or the double-tagged linear ssDNA fragment (or the complement of the tagged circular ssDNA fragment or the double-tagged linear ssDNA fragment) as a template. In some other embodiments, the capture tag domain comprises the 5′ portion of a transfer strand, and the 5′ portion of the transfer strand is linked to a chemical group or moiety comprising or consisting of an affinity binding molecule (e.g., where the 5′ portion of the transfer strand is linked to a first affinity binding molecule, such as biotin, streptavidin, an antigen, or an antibody that binds the antigen, which allows the circular tagged ssDNA fragment or the double-tagged linear ssDNA fragment to be captured on a surface to which a second affinity binding molecule is attached, and the second affinity binding molecule forms a specific binding pair with the first affinity binding molecule).

[0147] As used herein, "sequencing tag domain", "sequencing tag", or "sequencing primer binding sequence" refers to a sequence that facilitates the sequencing of an ssDNA fragment to which the tag is attached (e.g., provides a priming site for synthesis sequencing, or provides an annealing site for ligation sequencing, or provides an annealing site for hybridization sequencing). For example, in some embodiments, the sequencing tag domain or sequencing primer binding sequence provides a site for priming DNA synthesis of the ssDNA fragment or the complement of the ssDNA fragment. In some embodiments, the sequencing tag domain or sequencing primer binding sequence comprises an SBS3 sequence, an SBS8' sequence, an SBS12' sequence, or an SBS491' sequence.

[0148] As used herein, "amplification tag domain" refers to a tag domain that exhibits a sequence for the purpose of facilitating the amplification of a nucleic acid to which the tag is attached. For example, in some embodiments, the amplification tag domain provides a priming site for a nucleic acid amplification reaction using a DNA polymerase (e.g., a PCR amplification reaction, a strand displacement amplification reaction, or a rolling circle amplification reaction), or a ligation template for ligating a probe using a template-dependent ligase in a nucleic acid amplification reaction (e.g., a ligase chain reaction).

[0149] As used herein, "detection tag domain" or "detection tag" refers to a tag domain that exhibits a sequence or a detectable chemical or biochemical moiety for the purpose of facilitating the detection of a tagged circular ssDNA fragment or a doubly-tagged linear ssDNA fragment (e.g., where the sequence or chemical moiety comprises a detectable molecule or is linked to a detectable molecule; such as a detectable molecule selected from: visible dyes, fluorescent dyes, chemiluminescent dyes, or other detectable dyes; an enzyme detectable in the presence of a substrate, e.g., alkaline phosphatase in the presence of NBT plus BCIP, or peroxidase in the presence of a suitable substrate; a detectable protein, e.g., green fluorescent protein; and an affinity binding molecule that binds to a detectable moiety or can form an affinity binding pair or a specific binding pair with another detectable affinity binding molecule; or any of many other detectable molecules or systems known in the art).

[0150] As used herein, "transcription promoter domain" or "promoter domain" refers to a tag domain that exhibits a sense promoter sequence or an antisense promoter sequence of an RNA polymerase promoter.

[0151] As used herein, "DNA fragment" means a portion, piece, or segment of target DNA that has been cleaved, released, or broken from a longer DNA molecule such that it is no longer attached to the parental molecule. A DNA fragment can be double-stranded ("dsDNA fragment") or single-stranded ("ssDNA fragment"), and the process of generating DNA fragments from target DNA is referred to as "fragmenting" the target DNA. In some preferred embodiments, the method is used to generate a "DNA fragment library" comprising a collection or population of tagged DNA fragments.

[0152] As used herein, "target DNA" refers to any DNA of interest that is subjected to treatment, e.g., for generating a library of tagged DNA fragments (e.g., 5'- and 3'-tagged or dual-tagged linear ssDNA or dsDNA fragments or tagged circular ssDNA fragments).

[0153] "Target DNA" can be derived from any in vivo or in vitro source (including from one or more cells, tissues, organs, or organisms, whether alive or dead), or from any biological or environmental source (e.g., water, air, soil). For example, in some embodiments, the target DNA comprises and / or consists of eukaryotic and / or prokaryotic dsDNA that originates from or is derived from a human, animal, plant, fungus (e.g., mold or yeast), bacterium, virus, viroid, mycoplasma, or other microorganism. In some embodiments, the target DNA comprises or consists of: genomic DNA, subgenomic DNA, chromosomal DNA (e.g., from an isolated chromosome or a portion of a chromosome, such as one or more genes or loci from a chromosome), mitochondrial DNA, chloroplast DNA, plasmid or other episomal source of DNA (or recombinant DNA contained therein), or double-stranded cDNA that is prepared by reverse transcribing RNA using an RNA-dependent DNA polymerase or reverse transcriptase to generate first-strand cDNA, and then extending a primer annealed to the first-strand cDNA to generate dsDNA. In some embodiments, the target DNA comprises multiple dsDNA molecules that are contained within or prepared from nucleic acid molecules (e.g., multiple dsDNA molecules in genomic DNA or cDNA prepared from RNA, the genomic DNA or the RNA being in a biological (e.g., cell, tissue, organ, organism) or environmental (e.g., water, air, soil, saliva, sputum, urine, feces) source or from a biological (e.g., cell, tissue, organ, organism) or environmental (e.g., water, air, soil, saliva, sputum, urine, feces) source). In some embodiments, the target DNA is from an in vitro source. For example, in some embodiments, the target DNA comprises and / or consists of dsDNA that is prepared in vitro from single-stranded DNA (ssDNA) or from single-stranded or double-stranded RNA (e.g., using methods well known in the art, such as primer extension using a suitable DNA-dependent and / or RNA-dependent DNA polymerase (reverse transcriptase)).In some embodiments, the target DNA comprises or consists of dsDNA, which is prepared from all or part of one or more double-stranded or single-stranded DNA or RNA molecules using any method known in the art, the methods including those for: DNA or RNA amplification (e.g., PCR or reverse transcriptase-PCR (RT-PCR), transcription-mediated amplification methods, and amplification of all or part of one or more nucleic acid molecules); molecular cloning of all or part of one or more nucleic acid molecules in a plasmid, fosmid, BAC, or other vector, followed by replication in a suitable host cell; or capture of one or more nucleic acid molecules by hybridization, such as by hybridization with DNA probes on an array or microarray (e.g., by "sequence capture"; e.g., using kits and / or arrays from ROCHE NIMBLEGEN, AGILENT, or FEBIT).

[0154] In some embodiments, "target DNA" means dsDNA or ssDNA that has been prepared or modified (e.g., using various biochemical or molecular biology techniques) prior to being used to generate a library of tagged DNA fragments (e.g., 5'- and 3'-tagged or double-tagged linear ssDNA or dsDNA fragments or tagged circular ssDNA fragments).

[0155] As used herein, "amplify", "amplifying", or "amplification reaction" and their derivatives generally refer to any action or process by which at least a portion of a nucleic acid molecule is copied or replicated into at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally comprises a sequence that is substantially the same as or substantially complementary to at least some portions of the template nucleic acid molecule. The template nucleic acid molecule can be single-stranded or double-stranded, and the additional nucleic acid molecule can independently be single-stranded or double-stranded. Amplification optionally includes linear or exponential replication of the nucleic acid molecule. In some embodiments, such amplification can be carried out using isothermal conditions; in other embodiments, such amplification can include thermal cycling. In some embodiments, the amplification is multiplex amplification that includes simultaneous amplification of multiple target sequences in a single amplification reaction. In some embodiments, "amplify" includes amplifying at least some portions of DNA- and RNA-based nucleic acids, either alone or in combination. The amplification reaction can include any amplification process known to those of ordinary skill in the art. In some embodiments, the amplification reaction includes polymerase chain reaction (PCR).

[0156] As used herein, the term "polymerase chain reaction" ("PCR") refers to the methods of Mullis U.S. Patent Nos. 4,683,195 and 4,683,202, which describe methods for increasing the concentration of segments of a polynucleotide of interest in a mixture of genomic DNA without cloning or purification. The process for amplifying a polynucleotide of interest consists of the steps of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide of interest, followed by a series of thermal cycles in the presence of a DNA polymerase. The two primers are complementary to their respective strands of the double-stranded polynucleotide of interest. The mixture is first denatured at a higher temperature and then the primers are annealed to complementary sequences within the polynucleotide molecule of interest. After annealing, the primers are extended with polymerase to form a new pair of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated many times (referred to as thermal cycles) to obtain a high concentration of the amplified segment of the desired polynucleotide of interest. The length of the amplified segment (amplicon) of the desired polynucleotide of interest is determined by the relative positions of the primers with respect to each other and, thus, this length is a controllable parameter. By virtue of repeating this process, the method is referred to as "polymerase chain reaction" (hereinafter "PCR"). Since the amplified segments of the desired polynucleotide of interest become the major nucleic acid sequences (in terms of concentration) in the mixture, they are referred to as "PCR-amplified". In a modification of the method discussed above, the target nucleic acid molecule can be PCR-amplified using multiple different primer pairs (in some cases, one or more primer pairs for each target nucleic acid molecule of interest), thereby forming a multiplex PCR reaction.

[0157] As defined herein, "multiplex amplification" refers to the selective and non-random amplification of two or more target sequences in a sample using at least one target-specific primer. In some embodiments, multiplex amplification is performed such that some or all of the target sequences are amplified in a single reaction vessel. The "plexy" or "plex" of a given multiplex amplification generally refers to the number of different target-specific sequences that are amplified during this single multiplex amplification. In some embodiments, the plex can be about 12-plex, 24-plex, 48-plex, 96-plex, 192-plex, 384-plex, 768-plex, 1536-plex, 3072-plex, 6144-plex or higher. It is also possible to detect the amplified target sequences by several different methodologies (e.g., gel electrophoresis followed by densitometry, quantification with a bioanalyzer or quantitative PCR, hybridization with a labeled probe; incorporation of biotinylated primers followed by avidin-enzyme conjugate detection; incorporation of 32P-labeled deoxynucleotide triphosphate into the amplified target sequence).

[0158] As used herein, the term "primer" and its derivatives generally refer to any polynucleotide that can hybridize to a target sequence of interest. Typically, a primer serves as a substrate upon which nucleotides can be polymerized by a polymerase; however, in some embodiments, a primer can become incorporated into a synthesized nucleic acid strand and provide a site to which another primer can hybridize to initiate the synthesis of a new strand complementary to the synthesized nucleic acid molecule. A primer can include any combination of nucleotides or their analogs. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide.

[0159] In a number of embodiments, a primer has a free 3'-OH group that can be extended by a nucleic acid polymerase. For a template-dependent polymerase, typically at least the 3' portion of the primer oligonucleotide is complementary to a portion of the template nucleic acid, and the oligonucleotide "binds" (or "complexes", "anneals", or "hybridizes") to the portion of the template nucleic acid through hydrogen bonding to the template and other molecular forces to give a primer / template complex for initiating synthesis by a DNA polymerase, and the primer oligonucleotide is extended (i.e., "primer extension") during DNA synthesis by the addition of covalently bound bases complementary to the template that are attached at its 3' end. A primer extension product results. Template-dependent DNA polymerases (including reverse transcriptases) generally require the complexing of an oligonucleotide primer with a single-stranded template to initiate DNA synthesis ("priming"), but RNA polymerases generally do not require a primer complementary to a DNA template for the synthesis of RNA (transcription).

[0160] A "template" is a nucleic acid molecule that is copied by a nucleic acid polymerase such as a DNA polymerase. Whether the nucleic acid molecule contains two strands (i.e., is "double-stranded") or only contains one strand (i.e., is "single-stranded"), the strand of the nucleic acid molecule that dictates the sequence of nucleotides exhibited by the synthesized nucleic acid is the "template" or "template strand". The nucleic acid synthesized by a nucleic acid polymerase is complementary to the template. Both RNA and DNA are always synthesized in the 5'-to-3' direction starting from the 3' end of the template strand, and the two strands of a nucleic acid duplex always match such that the 5' ends of the two strands are at opposite ends of the duplex (and necessarily, then the 3' ends are also). For both RNA templates and DNA templates, a primer is required to initiate synthesis by a DNA polymerase, but a primer is not required to initiate synthesis by a DNA-dependent RNA polymerase, which is generally simply referred to as an "RNA polymerase".

[0161] The terms "polynucleotide" and "oligonucleotide" are used interchangeably herein and refer to polymeric forms of nucleotides of any length and may include ribonucleotides, deoxyribonucleotides, analogs thereof, or mixtures thereof. In some instances, the term "polynucleotide" may refer to a nucleotide polymer having a relatively large number of nucleotide monomers, while the term "oligonucleotide" may refer to a nucleotide polymer having a relatively small number of nucleotide monomers. However, unless indicated otherwise, this distinction is not applicable herein. Rather, the terms "polynucleotide" and "oligonucleotide" should be understood to include analogs of DNA or RNA made from nucleotide analogs as equivalents and apply to single-stranded (such as sense or antisense) polynucleotides and double-stranded polynucleotides. The term as used herein also encompasses cDNA, i.e., complementary or copy DNA produced, for example, by the action of reverse transcriptase from an RNA template. The term refers only to the primary structure of the molecule. Thus, the term includes triple-stranded, double-stranded, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-stranded, double-stranded, and single-stranded ribonucleic acid ("RNA").

[0162] In addition, the terms "polynucleotide", "nucleic acid", and "nucleic acid molecule" are used interchangeably and refer to a covalently linked sequence of nucleotides (i.e., ribonucleotides for RNA and deoxyribonucleotides for DNA), wherein the 3' position of the pentose of one nucleotide is linked to the 5' position of the pentose of the next nucleotide by a phosphodiester group. Nucleotides include sequences of any form of nucleic acid, including but not limited to RNA molecules and DNA molecules, such as cell-free DNA (cfDNA) molecules. The term "polynucleotide" includes but is not limited to single-stranded polynucleotides and double-stranded polynucleotides.

[0163] As used herein, the terms "ligating", "ligation", and their derivatives generally refer to the process of covalently linking two or more molecules together, such as covalently linking two or more nucleic acid molecules to each other. In some embodiments, ligation includes joining a gap between adjacent nucleotides of a nucleic acid. In some embodiments, ligation includes forming a covalent bond between the end of a first nucleic acid molecule and the end of a second nucleic acid molecule. In some embodiments, ligation may include forming a covalent bond between the 5' phosphate group of one nucleic acid and the 3' hydroxyl group of a second nucleic acid, thereby forming a ligated nucleic acid molecule. Generally for the purposes of the present disclosure, an amplified target sequence may be ligated to an adaptor to generate an adaptor-ligated amplified target sequence.

[0164] As used herein, "ligase" and its derivatives generally refer to any agent that can catalyze the ligation of two substrate molecules. In some embodiments, the ligase includes an enzyme that can catalyze the ligation of nicks between adjacent nucleotides of a nucleic acid. In some embodiments, the ligase includes an enzyme that can catalyze the formation of a covalent bond between the 5' phosphate of one nucleic acid molecule and the 3' hydroxyl of another nucleic acid molecule, thereby forming a ligated nucleic acid molecule. Suitable ligases can include, but are not limited to, T4 DNA ligase, T4 RNA ligase, and Escherichia coli (E. coli) DNA ligase.

[0165] As used herein, the term "adapter" generally refers to any linear oligonucleotide that can be ligated to a nucleic acid molecule to generate a nucleic acid product that can be sequenced on a sequencing platform such as various Illumina sequencing platforms. In some embodiments, the adapter includes two reverse complementary oligonucleotides that form a double-stranded structure. In some embodiments, the adapter includes two oligonucleotides that are complementary at one portion and mismatched at another portion, forming a Y-shaped adapter or a fork-shaped adapter that is double-stranded at the complementary portion and has two loose overhangs at the mismatched portion. Since Y-shaped adapters have complementary double-stranded regions, they can be considered a special form of double-stranded adapters. When the present disclosure compares Y-shaped adapters and double-stranded adapters, the term "double-stranded adapter" is used to refer to an adapter having two strands that are completely complementary, substantially (e.g., greater than 90% or 95%) complementary, or partially complementary.

[0166] In some embodiments, the adapter contains a sequence that binds to a sequencing primer (e.g., SEQ ID NO:3 and SEQ ID NO:5). In some embodiments, the adapter contains a sequence that binds to a flow cell oligonucleotide such as SEQ ID NO:7 and SEQ ID NO:8 (P7 sequence and P5 sequence) or its reverse complement.

[0167] In some embodiments, the adapter is substantially non-complementary to the 3' or 5' end of any target sequence present in the sample. Generally, the adapter can contain any combination of nucleotides and / or nucleic acids. In some aspects, the adapter can contain one or more cleavable groups at one or more positions. In another aspect, the adapter can contain a sequence that is substantially identical or substantially complementary to at least a portion of a primer (e.g., a universal primer). In some embodiments, the adapter can contain an index sequence (also referred to as a barcode or tag) to assist with downstream error correction, identification, or sequencing.

[0168] The term "universal adapter" is used to refer to an adapter of multiple sequencing adapters having one or more universal sequences. The multiple sequencing adapters are configured to be attached to (e.g., by ligation, hybridization, or transposition) a target nucleic acid to provide an adapter-target polynucleotide during a sequencing process for determining the sequence of the target nucleic acid. In some embodiments, all of the universal adapters during the sequencing process are the same and can be attached to target nucleic acids having different characteristics such as different sample sources, different organisms, different cell types, etc. During the sequencing process, oligonucleotides or polynucleotides (e.g., primers) having different index sequences are attached to the universal adapter before, during, or after the universal adapter is attached to the target nucleic acid. In some embodiments, a unique index sequence or a unique combination of index sequences is applied to and associated with nucleotides having the same characteristics (e.g., the same source, the same organism, or the same cell type). When target nucleic acids having different characteristics are combined / pooled in a sequencing reaction during the sequencing process, the index sequences provide a mechanism for determining or indexing the characteristics of the target nucleic acid.

[0169] An index sequence for indexing a sample source is a sample-specific sequence. An adapter containing a sample-specific index sequence is a sample-specific adapter and not a universal adapter. In the literature, some researchers use the term "universal adapter" in cases where the adapter is applied to libraries of more than one type of preparation, and the "universal adapter" still has a sample-specific index. The definition of such a "universal adapter" does not apply herein unless so specified.

[0170] As used herein, the term "flowcell" or "flow cell" refers to a chamber containing a solid surface through which one or more fluid reagents can flow. Examples of flowcells and associated fluid systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008); WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281; and US2008 / 0108082, each of which is incorporated herein by reference.

[0171] As used herein, the term "amplicon", when referring to a nucleic acid, means a product of copying a nucleic acid, wherein the product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be generated by any of a variety of amplification methods using the nucleic acid or an amplicon thereof as a template, the variety of amplification methods including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule that is a single copy of a specific nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatameric product of RCA). The first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies generated from the target nucleic acid or from the first amplicon after the first amplicon is generated. Subsequent amplicons can have a sequence that is substantially complementary to or substantially the same as the target nucleic acid.

[0172] The term "paired end reads" refers to reads obtained from paired end sequencing, which obtains a read from each end of a nucleic acid fragment. Paired end sequencing involves fragmenting DNA into sequences called inserts. In some protocols, such as some used by Illumina, reads from shorter inserts (e.g., on the order of tens to hundreds of bp) are referred to as short insert paired end reads or simply paired end reads. In contrast, reads from longer inserts (e.g., on the order of several thousand bp) are referred to as matepair reads. In the present disclosure, both short insert paired end reads and long insert matepair reads can be used and there is no distinction with respect to the process for determining the sequence of a DNA fragment. Thus, the term "paired end reads" can refer to both short insert paired end reads and long insert matepair reads, which will be further described below. In some embodiments, paired end reads include reads that are about 20 bp to 1000 bp. In some embodiments, paired end reads include reads that are about 50 bp to 500 bp, about 80 bp to 150 bp, or about 100 bp.

[0173] As used herein, the terms "alignment" and "aligning" refer to the process of comparing a read to a reference sequence and thereby determining whether the reference sequence contains the read sequence. As used herein, the alignment process attempts to determine whether a read can be mapped to the reference sequence, but does not always result in a read that matches the reference sequence. If the reference sequence contains the read, the read can be mapped to the reference sequence or, in some embodiments, to a specific location within the reference sequence. In some cases, the alignment simply determines whether the read is a member of a particular reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, an alignment of a read to a reference sequence of human chromosome 13 will determine whether the read is present in the reference sequence of chromosome 13.

[0174] Of course, alignment tools have many additional aspects and many other applications in bioinformatics that are not described in this application. For example, alignment can also be used to determine the degree of similarity between two DNA sequences from two different species, thereby providing a measure of how closely related they are on the evolutionary tree.

[0175] In some cases, the alignment additionally indicates the location within the reference sequence to which the read maps. For example, if the reference sequence is the entire human genome sequence, the alignment can indicate that the read is present on chromosome 13 and can also indicate that the read is at a particular strand and / or locus on chromosome 13. In some cases, alignment tools are imperfect because a) not all valid alignments are found and b) some of the alignments obtained are invalid. This occurs for various reasons, e.g., the read can contain errors and the sequenced read can differ from the reference genome due to haplotype differences. In some applications, alignment tools include built-in mismatch tolerance that tolerates a certain degree of base pair mismatches and still allows the alignment of the read to the reference sequence. This can help identify valid alignments of reads that would otherwise be missed.

[0176] The term "mapping" as used herein refers to the assignment of a read sequence to a larger sequence, such as a reference genome, by alignment.

[0177] The term "test sample" as used herein refers to a sample typically derived from a biological fluid, cell, tissue, organ, or organism, which sample contains a nucleic acid or mixture of nucleic acids having at least one nucleic acid sequence to be analyzed. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions or fine needle biopsy samples, urine, peritoneal fluid, pleural fluid, etc. Although the sample is typically obtained from a human subject (e.g., a patient), the assay can be used for samples from any mammal (including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc.) as well as mixed populations (such as a microbial population from the wild or a viral population from a patient). The sample can be used directly as obtained from the biological source or after a pretreatment that modifies the characteristics of the sample. For example, such a pretreatment can include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods can also include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc. If such a pretreatment method is employed with respect to the sample, then such a pretreatment method typically results in the nucleic acid of interest remaining in the test sample, sometimes at a concentration proportional to the concentration in the untreated test sample (e.g., that is, a sample that has not been subjected to any such pretreatment method). With respect to the methods described herein, such a "treated" or "processed" sample is still considered a biological "test" sample.

[0178] The term "next-generation sequencing (NGS)" as used herein refers to a sequencing method that permits large-scale parallel sequencing of clonally amplified molecules and individual nucleic acid molecules. Non-limiting examples of NGS include sequencing by synthesis using reversible dye terminators and sequencing by ligation.

[0179] The term "read" refers to a sequence read from a portion of a nucleic acid sample. Typically, although not necessarily, a read represents a short sequence of contiguous base pairs in the sample. A read can be represented symbolically by the sequence of base pairs A, T, C, and G of the sample portion, along with an estimate of the probability of base correctness (quality score). It can be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read can be obtained directly from a sequencing device or indirectly from stored sequence information about the sample. In some cases, a read is a DNA sequence that is long enough (e.g., at least about 20 bp) to be used to identify a larger sequence or region, e.g., that can be aligned and mapped to a chromosome or genomic region or gene.

[0180] The terms "locus" and "alignment position" are used interchangeably and refer to a unique position on a reference genome (i.e., chromosome ID, chromosomal position, and orientation). In some embodiments, a locus can be the position of a residue, the position of a sequence tag, or the position of a segment on a reference sequence.

[0181] As used herein, the term "reference genome" or "reference sequence" refers to any particular known gene sequence (partial or complete) of any organism or virus, which can be used as a reference for a sequence identified from a subject. For example, reference genomes for human subjects and many other organisms can be found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. A "genome" refers to the complete genetic information of an organism or virus represented as a nucleic acid sequence. However, it should be understood that "complete" is a relative concept, as even the gold standard reference genomes are expected to contain gaps and errors.

[0182] In various embodiments, the reference sequence is significantly larger than the read segments aligned to it. For example, it can be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 5 times, or at least about 10 6 times, or at least about 10 7 times.

[0183] In one example, the reference sequence is the sequence of the full-length human genome. Such a sequence can be referred to as a genomic reference sequence. In another example, the reference sequence is limited to a specific human chromosome, such as chromosome 13. In some embodiments, the reference Y chromosome is the Y chromosome sequence from human genome version hg19. Such a sequence can be referred to as a chromosomal reference sequence. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (such as strands), etc. of any species.

[0184] When used in the context of a nucleic acid or a mixture of nucleic acids, the term "derived from" herein refers to the means by which the nucleic acid is obtained from its source of origin. For example, in one embodiment, a mixture of nucleic acids derived from two different genomes means nucleic acids, such as cfDNA, that are naturally released by cells via natural processes such as necrosis or apoptosis. In another embodiment, a mixture of nucleic acids derived from two different genomes means nucleic acids extracted from two different types of cells from a subject.

[0185] As used herein, the term "biological fluid" refers to a liquid obtained from a biological source and includes, for example, blood, serum, plasma, sputum, lavage fluid, cerebrospinal fluid, urine, semen, sweat, tears, saliva, etc. As used herein, the terms "blood", "plasma" and "serum" expressly cover their fractions or processed parts. Similarly, in cases where a sample is obtained from a biopsy, swab, smear, etc., the "sample" expressly covers the processed fractions or parts derived from the biopsy, swab, smear, etc.

[0186] As used herein, the term "chromosome" refers to the genetic carrier of a living cell that carries genetic characteristics and is derived from chromatin strands containing DNA and protein components (especially histones). A conventional internationally recognized individual human genome chromosome numbering system is adopted herein.

[0187] Introduction and background

[0188] Next-generation sequencing (NGS) technologies have rapidly evolved, providing new tools for advancing research and science as well as healthcare and services that rely on genetic and related biological information. NGS methods are performed in a massively parallel manner, providing increasing speeds for determining biomolecular sequence information. Index sequences have been used in the art to tag samples for multiplex NGS sequencing or to identify their origin. However, many NGS methods and related sample manipulation techniques introduce errors, resulting in relatively high error rates in the resulting sequences, ranging from one error in several hundred base pairs to one error in several thousand base pairs. When such errors occur in the reads of the sample index sequences, the reads cannot be correctly associated with the origin of the sample and can cause incorrect associations between the reads and the origin of the sample.

[0189] One source of sequencing errors is related to index hopping. Index hopping, or index jumping, is observed when the DNA library molecules being sequenced contain index sequences different from those present in the library adapters during library preparation. Index hopping can occur during sample preparation or during cluster amplification of pooled multiplex libraries. One mechanism that causes index hopping involves the presence of free unligated adapter molecules that are present after library preparation.

[0190] Without wishing to be bound by theory, the problem of index hopping has multiple modes, some of which involve the presence of residual unligated adapter molecules left over from library preparation. One class of index hopping can be caused by free unligated adapter molecules present in the library pool that have a specific universal primer extension sequence (e.g., P7'), which can contribute to the formation of a library with swapped indices. This problem can be prevented by using a 5' exonuclease for degradation that specifically targets the P7' adapter strand. Such measures solve index hopping using biochemical methods. Some embodiments correct index hopping using bioinformatics methods as further described herein.

[0191] Some sequencing platforms use one color (e.g., a green laser) to sequence two base types (e.g., G / T), and use another color (e.g., a red laser) to sequence two other base types (e.g., A / C). On some of these platforms, in each cycle, at least 1 of the 2 nucleotides in each color channel needs to be read to ensure proper image registration. Importantly, the color balance of each base of the index read being sequenced is maintained; otherwise index read sequencing may fail due to registration failure. This can be especially a problem during low multiplexing sequencing, where a relatively small number of index sequences makes it more likely that all nucleotides in the read cycle will activate one color.

[0192] In various applications, it is desirable to layout the index plate such that the user can select groups of 3 rows across (i.e., a quarter of the rows) or groups of 4 columns down (i.e., half of the columns) or other rearrangements such as 6 - fold, 8 - fold, and 9 - fold without sacrificing the color balance of the oligonucleotides in the flow cell.

[0193] Various embodiments provide at least some of the following advantages.

[0194] Some embodiments using short universal adapters and index primers can be easily scaled up to high sample numbers without the need for new adapter designs. Only new index primers with new index sequences are required.

[0195] Some embodiments using short universal adapters and index primers are cost - effective, involving 1 complex instead of 40 complexes for 384 samples in a 16x24 well plate with combinatorial index pairs.

[0196] Since oligonucleotide purification is expensive, the handling of shorter length (e.g., 33bp compared to 70bp) adapters is cheaper and provides higher yields. In addition, high - performance liquid chromatography (HPLC) purification columns for different oligonucleotides can be shared across projects.

[0197] In some embodiments, the universal adaptor can be prepared with a simpler manufacturing process. In embodiments involving Y-shaped adaptors and blunt-ended adaptors, annealing of 2 oligonucleotides is required to prepare one universal adaptor. In embodiments involving double-stranded blunt-ended A&B version adaptors, annealing of 4 oligonucleotides is required to prepare two universal adaptors.

[0198] Some embodiments provide a simpler quality control (QC) process. Various existing processes and tools are fully functional for adaptors, including gravimetric analysis (weighing of oligonucleotides), OD, mass spectrometry, and purity determination for indexing primers.

[0199] The assay performance can be improved due to the smaller adaptor size, resulting in increased ligation efficiency and more effective clearance to remove dimers (e.g., using SPRI beads).

[0200] Workflow for sequencing nucleic acid fragments using index sequences

[0201] Figures 1A - 1C The figures illustrate exemplary workflows 100 and 120 for sequencing nucleic acid fragments using index sequences. Workflows 100 and 120 are merely illustrative of some embodiments. It should be understood that some embodiments employ workflows with additional operations not illustrated herein, while other embodiments may skip some of the operations illustrated herein. For example, workflow 120 is used for whole genome sequencing. In some embodiments involving targeted sequencing, the operation steps of hybridizing and enriching certain regions can be applied between operation 122 and operation 128. Additionally, the workflows illustrate the application of index sequences by ligation of sample-specific barcoded adaptors. Transposase-mediated adaptors can be applied. Additionally, universal adaptors without sample-specific sequences can be applied alternatively or additionally.

[0202] Operation 102 applies oligonucleotides to both ends of nucleic acid fragments (or target fragments) of multiple samples, the oligonucleotides comprising index sequences for identifying the sources of the multiple samples. In some embodiments, the index sequences are selected from an index sequence set comprising at least 6 different index sequences, and each subset of the multiple subsets of oligonucleotides comprises multiple index sequences of the index sequence set. In some embodiments, the Hamming distance between any two index sequences in the index sequence set is not less than a first standard value, where the first standard value is at least 2. The index sequence set contains multiple pairs of color-balanced index sequences, where any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine (A) base or a cytosine (C) base, and (ii) a guanine (G) base, a thymine (T) base, or a uracil (U) base.

[0203] In some embodiments, operation 102 attaches to each end of a double-stranded target fragment separated from a source to produce an adaptor-target-adaptor molecule. The attachment can be performed by using standard library preparation techniques for ligation, or by tagging with a transposase complex (Gunderson et al., WO2016 / 130704). In some embodiments, the attachment can be performed by Figure 1B the ligation process 120 shown in

[0204] Process 120 involves fragmenting nucleic acid of multiple samples. In some embodiments, the fragments are double-stranded DNA of a size, for example, less than 1000 bp. For example, DNA fragments can be obtained by fragmenting genomic DNA, collecting naturally fragmented DNA (e.g., cfDNA or ctDNA), or synthesizing DNA fragments from RNA. In some embodiments, to synthesize DNA fragments from RNA, messenger RNA or non-coding RNA is first purified using poly A selection or depletion of ribosomal RNA, then the selected mRNA is chemically fragmented using random hexamer priming and converted into single-stranded cDNA. The complementary strand of the cDNA is generated to produce double-stranded cDNA, which is ready for library construction. To obtain double-stranded DNA fragments from genomic DNA (gDNA), the input gDNA is fragmented, for example, by hydrodynamic shearing, nebulization, enzymatic fragmentation, etc., to generate fragments of an appropriate length, such as about 1000 bp, 800 bp, 500 bp, or 200 bp. For example, nebulization can break down DNA into fragments less than 800 bp within a short period of time. This process generates double-stranded DNA fragments.

[0205] In some embodiments, fragmented or damaged DNA can be processed without additional fragmentation. For example, formalin-fixed, paraffin-embedded (FFPE) DNA or certain cfDNA are sometimes fragmented sufficiently such that no additional fragmentation step is required.

[0206] Figure 1C Shown are the DNA fragments / molecules and adaptors employed in the initial steps of workflow 120 in Figure 1B Although only one double-stranded fragment is illustrated in Figure 1C , thousands to millions of fragments of a sample can be prepared simultaneously in the workflow. DNA fragmentation by physical methods produces heterogeneous ends, including a mixture of 3'-overhangs, 5'-overhangs, and blunt ends. The overhangs will have varying lengths, and the ends may or may not be phosphorylated. An example of a double-stranded DNA fragment obtained by fragmenting genomic DNA in operation 122 is shown as fragment 133 in Figure 1C .

[0207] Fragment 133 has both a 3' overhang at the left end and a 5' overhang as shown at the right end. If the DNA fragment is generated by physical means, workflow 120 proceeds to end repair operation 124, which produces blunt-ended fragments with 5'-phosphorylated ends. In some embodiments, this step uses T4 DNA polymerase and Klenow enzyme to convert the overhangs generated from fragmentation into blunt ends. The 3' to 5' exonuclease activity of these enzymes removes the 3' overhangs, and the 5' to 3' polymerase activity fills in the 5' overhangs. Additionally, T4 polynucleotide kinase phosphorylates the 5' ends of the DNA fragments in this reaction. Figure 1C Fragment 135 in is an example of an end-repaired blunt-ended product.

[0208] After end repair, workflow 120 proceeds to operation 126 to adenylate the 3' ends of the fragments, which is also known as A-tailing or dA-tailing because a single dATP is added to the 3' ends of the blunt-ended fragments to prevent them from ligating to each other during adapter ligation reactions. Figure 1C The double-stranded molecule 137 of shows an A-tailed fragment with blunt ends, which has 3'-dA overhangs and 5'-phosphate ends. As Figure 1C As seen in entry 139 of, a single "T" nucleotide at the 3' end of each of the two sequencing adapters provides an overhang complementary to the 3'-dA overhangs at each end of the insert fragment for ligating the two adapters to the insert fragment.

[0209] After adenylating the 3' ends, workflow 120 proceeds to operation 128 to ligate oligonucleotides (such as adapters) to both ends of the fragments of multiple samples. The oligonucleotides include index sequences for identifying the origin of the multiple samples.

[0210] Figure 1CEntry 139 shows two adaptors to be ligated to a double-stranded fragment, the adaptor comprising two index sequences i5 and i7. The index sequences provide a means for identifying the origin of multiple samples, thus allowing multiplexing of multiple samples on a sequencing platform. Other index sequences can be applied. The P5 oligonucleotide and the P7' oligonucleotide are complementary to the amplification primers bound to the surface of the flow cell of the Illumina sequencing platform and are also referred to as amplification primer binding sites. They allow the adaptor-target-adaptor library to undergo bridge amplification. Other designs of adaptors and sequencing platforms can be used in various embodiments. Adaptors and sequencing techniques are further described in the following sections. The adaptor also contains two sequence primer binding sequences SP1 (e.g., the SBS3 primer of Illumina, for reading the i5 index sequence) and SP2 (e.g., SBS12'). Other sequencing primer binding sequences can be included in adaptors for different reactions and platforms.

[0211] Return to Figure 1A , process 100 continues to pool nucleic acid fragments from multiple samples for a sequencing reaction. See module 104. Index oligonucleotides containing index sequences are attached to the fragments, the index sequences being applied in a manner specific to the origin of the sample. Various techniques for pooling samples are further described below.

[0212] In some embodiments, the product of the ligation reaction is purified and / or size selected by agarose gel electrophoresis or magnetic beads. The size-selected DNA is then PCR amplified to enrich for fragments with adaptors on both ends. See module 106. As mentioned above, in some embodiments, operations of hybridizing and enriching certain regions of DNA fragments can be applied to target regions for sequencing.

[0213] Then, workflow 100 continues to perform cluster amplification of the PCR product, for example, on the Illumina platform. See operation 108. By clustering the PCR product, the library can be pooled for multiplexing, for example, with 96 or more samples per lane, using different index sequences on the adaptors to track different samples. 1536 multiplexing techniques are envisioned.

[0214] After cluster amplification, sequencing reads can be obtained by synthetic sequencing on the Illumina platform. See operation 110. The obtained reads contain reads of the target sequence and the index sequence. Although the adapter and sequencing process described here are based on the Illumina platform, other sequencing technologies, especially NGS methods, can also be used instead of the Illumina platform or in addition to the Illumina platform. Finally, the workflow 100 determines the sample source of the target sequence based on the index sequence associated with the sample. See operation 112.

[0215] Figure 1D The figure shows a process 150 for sequencing target nucleic acids derived from multiple samples. Process 150 involves applying index sequences to target nucleic acids of multiple samples by contacting multiple index polynucleotides with target nucleic acids derived from the samples to generate multiple index-target polynucleotides. In some embodiments, the multiple index polynucleotides include DNA or RNA. Each sample is associated with a unique index sequence or a unique combination of index sequences. See module 152. In some embodiments, the multiple index polynucleotides include sample-specific adapters that include index sequences uniquely associated with each sample. Figure 1C , Figure 2A and Figure 2B The figure illustrates an embodiment using sample-specific adaptors.In other embodiments, the plurality of index polynucleotides comprises index primers that can hybridize to universal adaptors attached to target nucleic acids. Figure 1C , Figure 1K , Figure 2C and Figure 2D The figure illustrates some embodiments using index primers and universal adaptors.

[0216] In some embodiments, applying index primers to target nucleic acids in multiple samples can be accomplished by Figure 1F The process 160 shown in FIG. Figure 1K The second half of process 199 shown in the figure is completed. Figure 1E and Figure 1F Shown is the process of performing transposome-mediated fragmentation and applying index primers to nucleic acids having double-stranded short universal adaptors attached to both ends.

[0217] Procedure 160 involves providing a plurality of double-stranded nucleic acid molecules from a plurality of samples. Double-stranded nucleic acid 166 (e.g., DNA) is a schematic illustration of one of the double-stranded nucleic acid molecules. Procedure 160 also involves providing a plurality of transposome complexes. Each transposome complex comprises a transposase and two transposon end compositions. Elements 161 - 165 form a transposome complex. Three transposome complexes 169a - 169c are illustrated here. Transposome complex 169a comprises a transposase 161 and two transposon end compositions. Transposon end sequence duplex 162 and 5'-tag 163 form one transposon end composition. Transposon end sequence duplex 164 and 5'-tag 165 form another transposon end composition. Transposon end sequence duplexes 162 and 164 comprise two sequence strands commonly referred to as MEDS. One strand of the MEDS duplex comprises the sequence of SEQ ID NO:1, which will be transferred from the transposome complex to the target DNA and is referred to as the transfer strand. The other strand of the MEDS comprises the transposon end sequence of SEQ ID NO:2, which is not transferred to the target nucleic acid and is referred to as the non-transfer strand. At the 5'-end of the transfer strand, the transposon end composition comprises 5'-tag 165. In some embodiments, the 5'-tag is the sequence primer binding sequence SP1, which provides a sequence binding site on the target nucleic acid after being transposed into the target nucleic acid. Transposon end duplex MEDS 162 and 5'-tag 163 form another transposon end composition. The 5'-end of the transfer strand of the transposon end composition is the 5'-tag sequence 163, which provides the sequence primer binding sequence SP2.

[0218] Similarly, transposome complexes 169b and 169c comprise the same components as transposome complex 169a. For example, transposome complex 169b comprises two transposon end compositions, one of which comprises transposon end duplex 162b and 5'-tag 163b (SP2).

[0219] Procedure 160 involves incubating the DNA fragments and the transposome complexes under conditions that permit a transposition reaction with a suitable concentration of the transposon complex and the DNA molecules. The transposase in the transposome complex digests the double-stranded nucleic acid 166 at random sites indicated by the black triangles 167a - 167f. The digestion divides the double-stranded nucleic acid molecule 166 into a plurality of fragments including fragments 168a - 168d.

[0220] The transposase also transposes the transfer strand of the MEDS duplex to the 5' end of the nucleic acid fragment at the digestion site (167a - 167f). After fragmentation and transposition, the 5' end of the top strand of fragment 168b has the transfer strand of the MED duplex (164) that has transposed and attached to the 5' end. At the 5' end of the transfer strand is a 5' tag (165) corresponding to the sequencing primer binding sequence SP1. At each 3' end of the double-stranded target fragment 168b, there is a gap between the non-transfer strand of the MEDS transposon end sequence and the target fragment. After fragmentation and transposition, four fragments (170a - 170d) are formed. Two of the four fragments, 170b and 170c, have MEDS duplexes with 5' tags at both ends. Two of the fragments (170a and 170d) have transposon end compositions only at one end, which are not processed in the downstream sequencing reaction. In some embodiments, after the target DNA fragment is formed and tagged, a DNA polymerase with strand displacement or 5'-to-3' exonuclease activity is added to extend the 3' end of the target nucleic acid.

[0221] Figure 1F Shown is an additional downstream process of the DNA fragment generated by transposome-mediated fragmentation to obtain a target nucleic acid fragment with double-stranded universal adapters at both ends. The figure also shows the addition of index sequences (i5 index sequence and i7 index sequence) and flow cell amplification primer binding sequences (P5 sequence and P7 sequence). After adding a polymerase with strand displacement or 5'-to-3' exonuclease activity, the 3' end of the target nucleic acid is extended and the non-transfer strand of the MEDS duplex is removed (see arrows 173a and 173b, which indicate the extension of the 3' end of the target nucleic acid fragment). The extension fills the gap between the 3' end of the target nucleic acid and the non-transfer strand of the MEDS duplex. The extension also generates nucleotides complementary to the 5' tag. As a result, a double-stranded target nucleic acid fragment flanked by MEDS sequences and sequencing primer binding sequences is formed, which has two complementary strands 174 and 175a. The double-stranded nucleic acid contains two double-stranded short universal adapters, each adapter containing a sequencing primer sequence and a MEDS sequence. In some embodiments, the double-stranded nucleic acid has Figure 1G the nucleotides shown in

[0222] Figure 1GShows the sequence of a target nucleic acid with double-stranded short universal adapters attached to both ends. The sequencing primer binding sequence SP1 has the sequence TCGTCGGCAGCGTC (SEQ ID NO:3) at the top strand and the reverse complement (GACGCTGCCGACGA (SEQ ID NO:4)) at the bottom strand. The MEDS duplex has the sequences of SEQ ID NO:1 and SEQ ID NO:2. The sequencing primer binding sequence (SP2) has the sequence CCGAGCCCACGAGAC (SEQ ID NO:5) at the top strand and the reverse complement GTCTCGTGGGCTCGG (SEQ ID NO:6) at the bottom strand.

[0223] Figure 1I Shows the sequence in the i7 index primer. In some embodiments, the i7 index primer (e.g., 178a) has, from 5′ to 3′, the P7 flow cell amplification primer binding sequence CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO:7), the i7 index sequence, and the SP2 sequencing primer binding sequence GTCTCGTGGGCTCGG (SEQ ID NO:6).

[0224] Figure 1J Shows the sequence in the i5 index primer. In some embodiments, the i5 index primer (e.g., 176a) has, from 5′ to 3′, the P5 flow cell amplification primer binding sequence AATGATACGGCGACCACCGAGATCTACAC (SEQ ID NO:8), the i5 index sequence, and the SP1 sequencing primer binding sequence TCGTCGGCAGCGTC (SEQ ID NO:3).

[0225] In other embodiments, a target insert with a Y-shaped universal adapter can be used in a process such as Figure 1K as shown. Figure 1H Shows the sequence of a target nucleic acid with a Y-shaped short universal adapter attached to both ends according to some embodiments. The Y-shaped universal adapter has the sequence TCGTCGGCAGCGTC (SEQ ID NO:3) at the 5′ arm and the sequence CCGAGCCCACGAGAC (SEQ ID NO:5) at the 3′ arm.

[0226] Procedure 160 also involves denaturing a double-stranded nucleic acid fragment having double-stranded short universal adapters at both ends. It also involves adding a primer (176a) and a nuclease that hybridize to the denatured nucleic acid fragment. As shown in the figure, the bottom strand 175a is further processed. The primer (176a) contains a P5 flow cell amplification primer binding site at the 5′ end, an i5 index sequence downstream of the P5 sequence, and SP1. This polynucleotide is also referred to as an index primer. The index primer hybridizes to the single-stranded nucleic acid 175b at the SP1 primer binding site. The polymerase extends the 3′ end of the index primer 176a to form an extended single-stranded nucleic acid fragment using the fragment 175b as a template. The resulting nucleic acid fragment is shown as 176b. Then, the procedure further adds a primer and a polymerase to further extend the fragment 176b. The primer added in this reaction contains a P7 flow cell amplification primer binding site at the 5′, an i7 index sequence at the 3′ of the P7 sequence, and an SP2 sequencing primer binding sequence. Then the 3′ end of the index primer sequence 178a is extended using the single-stranded nucleic acid 176b as a template. In addition, the 3′ end of the nucleic acid 176b is also extended using the index primer 178a as a template. As a result, a double-stranded nucleic acid fragment is formed, where one strand 176c extends from the fragment 176b and the other strand 178b extends from the index primer 178a. The final double-stranded nucleic acid fragment contains, in the 5′ to 3′ direction in the top strand (176c), a P5 flow cell amplification primer binding site, an i5 sequence, an SP1 sequencing primer binding sequence, a MEDS sequence, a target sequence, a MEDS sequence, an SP2 sequence, an i7 sequence, and a P7′ sequence. This final double-stranded nucleic acid fragment forms a library fragment for a sequencing platform such as the SBS platform of Illumina.

[0227] According to procedure 160, some embodiments provide a method for sequencing target nucleic acids derived from multiple samples. The method includes: (a) providing a plurality of double-stranded nucleic acid molecules derived from multiple samples; (b) providing a plurality of transpososome complexes, where each transpososome complex contains a transposase and two transposon end compositions; (c) incubating the double-stranded nucleic acid molecules with the transpososome complexes to obtain double-stranded nucleic acid fragments, where the double-stranded nucleic acid fragments contain sequences transposed from the transposon end compositions at one or both ends; (d) contacting a plurality of index primers with the double-stranded nucleic acid fragments to generate a plurality of index-fragment polynucleotides, where the index primers contacting the double-stranded nucleic acid fragments derived from each sample contain an index sequence or a combination of index sequences uniquely associated with that sample, and the index sequence or the combination of index sequences is selected from a set of index sequences; (e) pooling the plurality of index-fragment polynucleotides; (f) sequencing the pooled index-fragment polynucleotides to obtain index reads of the index sequences and a plurality of target reads of the target sequences, each target read being associated with at least one index read; and (g) using the index reads to determine the sample origin of the target reads.

[0228] In some embodiments, at least one transposome complex comprises a Tn5 transposase and a Tn5 transposon end composition. In some embodiments, at least one transposome complex comprises a Mu transposase and a Mu transposon end composition. Some embodiments include both a Tn5 transposase and a Mu transposase, as well as a transposon end composition.

[0229] Figure 1K Process 199 for adding an index sequence to a target nucleic acid having Y-shaped short universal adapters at both ends is shown. Process 199 is similar to Figure 1E process 160 thereof, which uses a target nucleic acid fragment having double-stranded short universal adapters attached to both ends. In process 199, a target nucleic acid having Y-shaped short adapters attached to both ends is used. Because the two strands of the Y-shaped adapter have two different sequencing primer binding sequences, both strands of the nucleic acid can be used to generate downstream fragments that can be sequenced on a sequencing platform. In contrast, in embodiments using double-stranded adapters, only one strand of the product of the double-stranded nucleic acid can be used for sequencing.

[0230] According to Figures 1H - 1J the embodiment illustrated therein, the nucleic acid sequences of the Y-shaped adapter and the index primer are shown. At the beginning of the process, a double-stranded nucleic acid having two Y-shaped short universal adapters attached to both ends is shown. The double-stranded nucleic acid comprises a top strand 190 and a bottom strand 191. A blocking portion 198 is shown at the 3′ end of the top strand 190, which blocks nucleic acid extension when polymerase is added. Although only one blocking group is shown in the figure, in some embodiments, additional blocking groups can be applied to other ends of the double-stranded nucleic acid.

[0231] Various blocking agents can be implemented. One possible form of blocking agent includes phosphorothioate (PS) bonds. Phosphorothioate (PS) bonds replace the non-bridging oxygen in the phosphate backbone of the oligonucleotide with a sulfur atom. Approximately 50% of the time (due to the 2 resulting stereoisomers that can be formed), PS modification makes the internucleotide bond more resistant to nuclease degradation. Therefore, it is recommended to include at least 3 PS bonds at the 5′ oligonucleotide end and the 3′ oligonucleotide end to inhibit exonuclease degradation. Including PS bonds throughout the oligonucleotide will also help reduce attack by endonucleases, but may also increase toxicity.

[0232] Another possible form of the blocker includes reverse dT and reverse ddT. Reverse dT can be incorporated at the 3′ end of the oligonucleotide, resulting in 3′-3′ bonding that will inhibit degradation by 3′ exonucleases and extension by DNA polymerase. In addition, placing a reverse 2′,3′ dideoxy-dT base (5′ reverse ddT) at the 5′ end of the oligonucleotide prevents false ligation and can protect against some forms of enzymatic degradation.

[0233] Another possible form of the blocker includes phosphorylation. Phosphorylation of the 3′ end of the oligonucleotide will inhibit degradation by some 3′ exonucleases.

[0234] Another possible form of the blocker includes LNA, where xGen locked nucleic acid modification prevents endonuclease digestion and exonuclease digestion.

[0235] The top strand 190 contains the SP1 sequence, the MEDS sequence, the target insert, the MEDS sequence, and the SP2 sequence from 5′ to 3′. The bottom strand 191 contains the SP2 sequence, the MEDS sequence, the target insert, the MEDS sequence, and the SP1 sequence from 3′ to 5′. Process 199 denatures the double-stranded nucleic acid and adds primers and polymerase to the nucleic acid. The index primer 192a hybridizes to the SP2 primer binding sequence and is extended using the single-stranded fragment 190 as a template. The 3′ end of the single-stranded nucleic acid 190 does not extend because it is blocked by the blocking group 198. After extension, a double-stranded structure containing the top strand 190 and the bottom strand 192b is obtained. Then, the double-stranded nucleic acid is denatured again. Process 199 adds primers and polymerase to the reaction mixture. The i5 index primer 194a is added, which hybridizes to SP1. The i5 index primer 194a contains the P5 sequence, the i5 index sequence, and the SP1 primer binding sequence from 5′ to 3′. The i5 index primer hybridizes to the SP1 sequence of the single-stranded nucleic acid 192c. Then, the PCR reaction extends the 3′ end of the i5 index primer 194a and the 3′ end of the single-stranded fragment 192c. After polymerase extension, a double-stranded nucleic acid is obtained that contains the top strand 194b and the bottom strand 192d. The top strand 194b contains the P5 flow cell amplification primer binding sequence, the i5 index sequence, the SP1 sequencing primer binding sequence, the MEDS sequence, the target sequence, the MEDS sequence, the SP2 sequencing primer binding sequence, the i7 index sequence, and the P7′ flow cell amplification primer binding sequence from 5′ to 3′. This double-stranded nucleic acid contains the sequences required for amplification and sequencing reactions on the Illumina sequencing platform.

[0236] Back to Figure 1D, Process 150 involves applying an index sequence to the target nucleic acids of multiple samples. In some embodiments, this is achieved by contacting multiple index polynucleotides with the target nucleic acids derived from multiple samples to generate multiple index-target polynucleotides. In some embodiments, the index polynucleotides contacted with the target nucleic acids derived from each sample contain an index sequence or a combination of index sequences uniquely associated with that sample. The index sequence or the combination of index sequences is selected from an index sequence set. The Hamming distance between any two index sequences in the index sequence set is not less than a first standard value, where the first standard value is at least 2.

[0237] In some embodiments, the index sequence set includes multiple pairs of color-balanced index sequences, where any two bases at the corresponding sequence positions of each pair of color-balanced index sequences include both: (i) an A base or a C base, and (ii) a G base, a T base, or a U base. In some embodiments, the index sequence set includes at least six different index sequences.

[0238] In some embodiments, the multiple index polynucleotides include index primers that can hybridize to a universal adapter. In some embodiments, the multiple index primers contain the index sequences of the index sequence set. In some embodiments, each index primer further contains a flow cell amplification primer binding sequence. In some embodiments, the flow cell amplification primer binding sequence includes a P5 sequence or a P7′ sequence. See Figure 2C and Figure 2D . In some embodiments, the target nucleic acids derived from multiple samples include nucleic acids having a universal adapter covalently attached to one or both ends. See Figure 1F the nucleic acid having a top strand 174 and a bottom strand 175a in Figure 1K and the nucleic acid having a top strand 190 and a bottom strand 191 in

[0239] In some embodiments, contacting the multiple index polynucleotides with the target nucleic acids derived from multiple samples includes: hybridizing the multiple index primers to the universal adapter covalently attached to one or both ends of the nucleic acid; and extending the multiple index primers to obtain multiple index-adapter-target polynucleotides. In some embodiments, the universal adapter and the target nucleic acid are double-stranded, and hybridizing the multiple index primers to the universal adapter includes hybridizing the multiple index primers to only one strand of the universal adapter.

[0240] In some embodiments, the universal adapter and the target nucleic acid are double-stranded, and hybridizing the multiple index primers to the universal adapter includes hybridizing the multiple index primers to both strands of the universal adapter. See Figure 1F and Figure 1K .

[0241] In some embodiments, the indexing primer that hybridizes to the first strand of the universal adaptor comprises an indexing sequence selected from a first subset of an indexing sequence set, and the indexing primer that hybridizes to the second strand of the universal adaptor comprises a sequence selected from a second subset of the indexing sequence set, wherein the first subset does not overlap with the second subset. In some embodiments, the first subset comprises the indexing sequences listed in Table 1, and the second subset comprises the indexing sequences listed in Table 2. In some embodiments, the indexing primer that hybridizes to both strands of the universal adaptor comprises an indexing sequence selected from the same subset of the indexing sequence set. In some embodiments, the subset of the indexing sequences is selected from one of the subsets of the indexing sequences in Table 3.

[0242] In some embodiments, the universal adaptor comprises a double-stranded adaptor. See, for example, Figure 2D . In some embodiments, the universal adaptor comprises a Y-shaped adaptor. See, for example, Figure 2C . In some embodiments, the universal adaptor comprises a single-stranded adaptor. In some embodiments, the universal adaptor comprises a hairpin adaptor. In some embodiments, each of the universal adaptors comprises a protruding end at one end to be attached to the nucleic acid before being attached to the nucleic acid. In some embodiments, the protruding end is a T protruding end. See Figure 1C and Figures 2A - 2C . In some embodiments, each of the universal adaptors comprises a blunt end to be attached to the nucleic acid before being attached to the nucleic acid. See Figure 2D .

[0243] In some embodiments, the method comprises attaching a universal adaptor to one or both ends of a nucleic acid before applying the indexing sequence to the target nucleic acid. In some embodiments, the attachment comprises attaching the universal adaptor by transposome-mediated fragmentation.

[0244] In some embodiments, the attachment comprises ligating the universal adaptor to one or both ends of the nucleic acid. In some embodiments, the ligation comprises enzymatic ligation or chemical ligation.

[0245] In some embodiments, the attachment is carried out by amplification with a target-specific primer comprising a universal adaptor sequence. In some embodiments, the universal adaptor sequence is at the end of the primer.

[0246] Some embodiments apply a plurality of indexed polynucleotides comprising sample-specific adaptors. The adaptor comprises an indexing sequence of an indexing sequence set. See Figure 1C and Figures 2A - 2C. In some embodiments, a sample-specific adapter comprises two chains. In some embodiments, only one chain comprises an index sequence. In some embodiments, each chain of a sample-specific adapter comprises an index sequence. In some embodiments, a first chain of a sample-specific adapter comprises an index sequence selected from a first subset of a set of index sequences, and a second chain of a sample-specific adapter comprises an index sequence selected from a second subset of a set of index sequences, and the first subset does not overlap with the second subset. In some embodiments, the first subset of index sequences includes the index sequences listed in Table 1, and the second subset includes the index sequences listed in Table 2. In some embodiments, the first chain and the second chain of a sample-specific universal adapter comprise index sequences selected from the same subset of a set of index sequences. In some embodiments, a subset of index sequences is selected from one of the subsets of index sequences in Table 3.

[0247] In some embodiments, each sample-specific adaptor comprises a flow cell amplification primer binding sequence. Figure 1C , Figure 2A and Figure 2B In some embodiments, the flow cell amplification primer binding sequence includes a P5 sequence or a P7' sequence.

[0248] In some embodiments, contacting a plurality of index polynucleotides with a target nucleic acid comprises attaching a sample-specific adapter to the target nucleic acid by transposome-mediated fragmentation. In some embodiments, contacting a plurality of index polynucleotides with a target nucleic acid comprises connecting a sample-specific adapter to the target nucleic acid. In some embodiments, the connection comprises enzymatic connection or chemical connection. In some embodiments, chemical connection comprises chemical reaction connection.

[0249] In some embodiments, the sample-specific adapter comprises a Y-shaped adapter having a complementary double-stranded region and a mismatched single-stranded region. In some embodiments, each chain of the sample-specific adapter comprises an index sequence at the mismatched single-stranded region. In some embodiments, only one chain of the sample-specific adapter comprises an index sequence at the mismatched single-stranded region. In some embodiments, the sample-specific adapter comprises a single-stranded adapter. In some embodiments, the sample-specific adapter comprises a hairpin adapter. In some embodiments, contacting a plurality of index polynucleotides with a target nucleic acid involves attaching a plurality of index polynucleotides to both ends of the target nucleotide.

[0250] In some embodiments, contacting a plurality of index polynucleotides with a target nucleic acid comprises attaching the plurality of index polynucleotides to both ends of the target nucleic acid. In some embodiments, contacting a plurality of index polynucleotides with a target nucleic acid comprises attaching the plurality of index polynucleotides to only one end of the target nucleic acid.

[0251] In some embodiments, the combination of index sequences uniquely associated with a sample is an ordered combination of index sequences.

[0252] In some embodiments, the set of index sequences includes multiple non-overlapping subsets of index sequences, and the Hamming distance between any two index sequences in any subset is not less than a second criterion value, where the second criterion value is greater than the first criterion value. In some embodiments, the first criterion value is 4, and the second criterion value is 5. In some embodiments, the first criterion value is 3. In some embodiments, the first criterion value is 4. Various other designs of the set of index sequences can be applied as further described herein.

[0253] In some embodiments, process 150 includes fragmenting nucleic acid molecules obtained from multiple samples to obtain target nucleic acids before applying the index sequences to the target nucleic acids of the multiple samples. In some embodiments, the fragmentation includes transposome-mediated fragmentation, such as Figure 1E the process shown in

[0254] In some embodiments, the fragmentation is contacting with multiple PCR primers targeting the sequence of interest to obtain target nucleic acids containing the sequence of interest.

[0255] Process 150 involves pooling multiple index-target polynucleotides after obtaining them. See module 154. In some embodiments, process 150 further includes amplifying the pooled index-target polynucleotides before sequencing the polynucleotides.

[0256] In some embodiments, process 150 further involves sequencing the pooled index-target polynucleotides to obtain multiple index reads of the index sequences and multiple target reads of the target sequences, with each target read associated with at least one index read. See module 156.

[0257] Process 150 further involves using the index reads to determine the sample origin of the target reads. In some embodiments, this is achieved by a method including the following steps: for each index read, obtaining an alignment score for the set of index sequences, each alignment score indicating the similarity between the sequence of the index read and the index sequences of the set of index sequences; determining the match of the index read with a specific index sequence based on the alignment score; and determining that the target read associated with the specific index read originates from the sample uniquely associated with the specific index sequence.

[0258] Index sequence design

[0259] In multiple embodiments, various factors are considered, including but not limited to means for detecting errors within the index sequences, transformation efficiency, assay compatibility, GC content, homopolymers, and manufacturing considerations to identify index sequences or oligonucleotides.

[0260] For example, the index sequences can be designed to provide a mechanism for facilitating error detection. Figure 3 Schematically illustrated is an index oligonucleotide design that provides a mechanism for detecting errors that occur in the index sequences during the sequencing process. According to this design, each of the index sequences has six nucleotides and differs from all other index sequences by at least two nucleotides. As Figure 3 illustrated, index sequence 344 differs from index sequence 342 in the first two nucleotides from the left, as shown by the underlined nucleotides T and G in index sequence 344 and nucleotides A and C in index sequence 342. Index sequence 346 is a sequence identified as part of a read, and it differs from all other index sequences of the adapters provided in the process. Since the index sequence in the read presumably originates from the index sequence in the adapter, errors may have occurred during the sequencing process such as during amplification or sequencing. Index sequences 342 and 344 are illustrated as the two index sequences most similar to index sequence 346 in the read. It can be seen that index sequence 346 differs from index sequence 342 by one nucleotide in the first nucleotide from the left, which is T instead of A. In addition, index sequence 346 also differs from index sequence 344 by one nucleotide (although in the second nucleotide from the left), which is C instead of G. Because index sequence 346 in the read differs from both index sequence 342 and index sequence 344 by one nucleotide, from the information illustrated, it is not possible to determine whether index sequence 346 originated from index sequence 342 or index sequence 344. However, in many other cases, the index sequence error in the read will not have an equally large difference from the two most similar index sequences. As shown in the example for index sequence 348, index sequences 342 and 344 are also the two index sequences most similar to index sequence 348. It can be seen that index sequence 348 differs from index sequence 342 by one nucleotide in the third nucleotide from the left, which is A instead of T. In contrast, index sequence 348 differs from index sequence 344 by three nucleotides. Thus, it can be determined that index sequence 348 originated from index sequence 342 rather than index sequence 344, and the error may have occurred in the third nucleotide from the left. By controlling the level of difference between the index sequences (e.g., as measured by Hamming distance or edit distance), some embodiments provide index oligonucleotides for identifying the origin of multiple samples, where sequencing errors, sample handling errors, and other errors can be corrected by assigning the index sequence reads to closely matching index sequences and the samples associated with the closely matching index sequences.

[0261] Some embodiments apply i5-i7 index pairs to multiple samples, where each ordered index pair is unique. The Hamming distance between any two index sequences in the complete index sequence set is controlled to be above a threshold. In some embodiments, the Hamming distance between the ordered index pairs is also controlled to be above a threshold. Additionally, in some embodiments, the edit distance between the index sequences is also controlled. These and other elements of the index oligonucleotides allow detection and correction of index hopping by identifying errors that would otherwise be ambiguous and uncorrectable.

[0262] A typical calculation of the edit distance is the Levenshtein distance, where each insertion, deletion, or substitution is considered a single edit operation and is scored equally. Consider the case of "ACTGACTA" and "ACTACTAA". In this case, the Levenshtein edit distance would be 2, as shown in the alignment below.

[0263] ACTGACTA-

[0264] ACT-ACTAA

[0265] However, in the case of index sequences, this may underestimate the true distance between the two sequences. In reality, the index sequences will be extended by bases from the surrounding adapters. If the bases from the surrounding adapters happen to match another index sequence, this will effectively only require a single deletion event to convert one index sequence into the other. Additionally, the index can be read in the opposite direction, in which case additional adapter sequences can occur at the 5′ end of the index. While it is possible to look at the expected adapter sequences to see the likelihood of this happening, this would make the index sequences only valid in the context of a specific adapter. More precisely, a custom edit distance is generated that always assumes that adjacent adapter sequences will match the adapter. In this custom edit distance, only a single insertion / deletion event is allowed. An edit distance threshold of 3 means that index pairs are not allowed where a single deletion + substitution can convert one index sequence into the other.

[0266] In some embodiments, the edit distance is a modified Levenshtein distance where terminal gaps are not assigned a penalty. U.S. Provisional Patent Application No. 62 / 447,851, which is incorporated herein by reference in its entirety, describes various methods for determining the modified Levenshtein distance of nucleic acid sequences.

[0267] Index oligonucleotides, adapters, and primers

[0268] In addition to referring to the above Figures 1A - 1CIn addition to the adapter designs described in exemplary workflow 100, other designs of indexing oligonucleotides can be used in various embodiments of the methods and systems disclosed herein.

[0269] Figures 2A - 2D Multiple embodiments of indexing oligonucleotides are shown. Although the adapters are labeled with various components, they can include additional components that are not labeled, such as additional primer binding sites or cleavage or digestion sites. Figure 2A Shows a standard Illumina dual-index adapter. The adapter is partially double-stranded and is formed by annealing two oligonucleotides corresponding to the two strands. The two strands have a number of complementary base pairs (e.g., 12 - 17 bp or 6 - 34 bp), which allow the two oligonucleotides to anneal at the ends to be ligated to the dsDNA fragment. The dsDNA fragment to be ligated at both ends to obtain paired-end reads is also referred to as an insert. The other base pairs on the two strands are mismatched (non-complementary), resulting in a fork-shaped or Y-shaped adapter with two single-stranded overhangs.

[0270] On the strand with a 5′ single-stranded overhang (top strand), in the 5′ to 3′ direction, the adapter has a P5 sequence, an i5 index sequence, and a sequencing primer binding sequence SP1 (e.g., SBS3). On the strand with a 3′ single-stranded overhang, in the 3′ to 5′ direction, the adapter has a P7′ sequence, an i7 index sequence, and an SP2 sequencing primer binding sequence (e.g., SBS12′). The P5 oligonucleotide and the P7′ oligonucleotide are complementary to the amplification primers that bind to the solid phase of the flow cell attached to the sequencing platform. They are also referred to as amplification primer binding sites, regions, or sequences. In some embodiments, the index sequences provide a means to track the origin of samples, thus allowing multiplexing of multiple samples on the sequencing platform.

[0271] The complementary base pairs are part of the sequencing primer binding sequences SP1 and SP2. Downstream of the SP1 primer sequence (e.g., SBS3) is a single nucleotide 3′-T overhang, which provides an overhang complementary to the single nucleotide 3′-A overhang of the dsDNA fragment to be sequenced, which can facilitate the hybridization of the two overhangs. The sequencing primer binding sequence SP2 (e.g., SBS12′) is located on the complementary strand, and a phosphate group is attached upstream of this complementary strand. The phosphate group helps to ligate the 5′ end of the SP2 sequence to the 3′-A overhang of the DNA fragment.

[0272] In some embodiments where the indexing sequences are selected from a set of indexing sequences, each strand of the adaptor comprises an indexing sequence selected from the set of indexing sequences such as those shown in Tables 1-3 and described elsewhere herein. In some embodiments, each double-stranded sequencing adaptor in the set of oligonucleotides comprises a first strand and a second strand, the first strand comprising an indexing sequence selected from a first subset of the set of indexing sequences, and the second strand comprising an indexing sequence selected from a second subset of the set of indexing sequences. The first subset does not overlap with the second subset. In some embodiments, the first subset of indexing sequences includes the indexing sequences listed in Table 1, and the second subset of indexing sequences includes the indexing sequences listed in Table 2.

[0273] Table 1. I7 index set including subsets (index groups 0 - 3)

[0274]

[0275]

[0276]

[0277] Table 2. I5 index set including subsets (index groups 0 - 3)

[0278]

[0279]

[0280] In some embodiments, the indexing sequences on the first strand of the adaptor and the indexing sequences on the second strand of the adaptor are both selected from the same subset of a plurality of subsets of the set of indexing sequences. In some embodiments, the subset of indexing sequences is one of the subsets of the indexing sequences in Table 3 (labeled by plate number).

[0281] Table 3. I5 and I7 index sets including subsets (plates 1 - 4)

[0282]

[0283]

[0284]

[0285] The set of index sequences included in the oligonucleotide set includes a plurality of unique index sequences. In some embodiments, the Hamming distance between any two index sequences in the index sequence set is not less than a first standard value, where the first standard value is 2 or greater. The index sequence set includes multiple pairs of color-balanced index sequences. Any two bases at the corresponding sequence positions of each pair of color-balanced index sequences include both of the following: (i) an A base or a C base, and (ii) a G base, a T base, or a U base. In some embodiments, the first standard value is 3. In some embodiments, the first standard value is 4.

[0286] In some embodiments, the index sequence set includes multiple non-overlapping subsets of index sequences, such as the subsets shown in Tables 1-3. In these subsets, the Hamming distance between any two index sequences is not less than a second standard value. In some embodiments, the second standard value is greater than the first standard value. In some embodiments, the first standard value is 4 and the second standard value is 5.

[0287] In some embodiments, the oligonucleotide includes an index sequence at its 3′ end and an index sequence at its 5′ end. In such embodiments, the oligonucleotide can be a single-stranded nucleic acid fragment with adaptors attached to both ends. It can be, for example, a denatured fragment obtained from the adaptor-target-adaptor construct 140 shown in Figure 1C In some embodiments, the oligonucleotide includes an index sequence at its 3′ end and an index sequence at its 5′ end. In such embodiments, the oligonucleotide can be a single-stranded nucleic acid fragment with adaptors attached to both ends. It can be, for example, a denatured fragment obtained from the adaptor-target-adaptor construct 140 shown in

[0288] In some embodiments, the edit distance between any two index sequences in the index sequence set is not less than a third standard value. In some embodiments, the third standard value is 3. In some embodiments, the edit distance is the modified Levenshtein distance, where end gaps are not assigned penalties. U.S. Patent Application No. 15 / 863,737, which is incorporated herein by reference in its entirety, describes various methods for determining the modified Levenshtein distance of nucleic acid sequences.

[0289] In some embodiments, each index sequence in the index sequence set has 8 bases; the first standard value is 3; and the third standard value is 2. In some embodiments, the index sequence set includes the sequences listed below in Example 2. In some embodiments, each index sequence in the index sequence set has 10 bases; the first standard value is 4; and the third standard value is 3. In some embodiments, the index sequence set includes the sequences listed below in Example 3.

[0290] From a bioinformatics perspective, longer oligonucleotides can provide more candidates that meet various constraints of interest such as edit distance or Hamming distance. However, longer oligonucleotides are more difficult to fabricate and will result in undesirable reactions (such as undesirable reactions through self-hybridization, cross-hybridization, folding) and other side effects. In contrast, while shorter oligonucleotides can avoid some of these side effects, they may not be able to meet bioinformatics constraints such as providing a large enough Hamming distance or edit distance to allow for error correction. A balance between bioinformatics robustness and biochemical functionality must be considered. In some embodiments, each index sequence in the oligonucleotide pool has 32 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 16 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 10 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 8 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 8 bases. In some embodiments, each index sequence in the oligonucleotide pool has 7 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 6 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 5 or fewer bases. In some embodiments, each index sequence in the oligonucleotide pool has 4 or fewer bases.

[0291] In some embodiments, the set of index sequences incorporated into the index oligonucleotides does not include index sequences that have been empirically determined to have poor performance in indexing the source of nucleic acid samples in multiplexed massively parallel sequencing. In some embodiments, the index sequences include the sequences in Table 4. Other sequences not listed in Table 4 may also not be included.

[0292] Table 4. Index sequences not included

[0293] Index tag Index sequence >N501 TAGATCGC >N504 AGAGTAGA >N513 TCGACTAG >N515 TTCTAGCT >N516 CCTAGAGT >N501 - rc GCGATCTA >N513 - rc CTAGTCGA >N515 - rc AGCTAGAA >N516 - rc ACTCTAGG >N504 - rc TCTACTCT >N704 TCCTGAGC >N715 ATCTCAGG >N710 CGAGGCTG >N705 GGACTCCT >N709 GCTACGCT >N709 - rc AGCGTAGC >N715 - rc CCTGAGAT >N705 - rc AGGAGTCC >N704 - rc GCTCAGGA >N710 - rc CAGCCTCG

[0294] In some embodiments, the set of index sequences comprises at least 12 different index sequences. In some embodiments, the set of index sequences comprises at least 20 different index sequences. In some embodiments, the set of index sequences comprises at least 24 different index sequences. In some embodiments, the set of index sequences comprises at least 28 different index sequences. In some embodiments, the set of index sequences comprises at least 48 different index sequences. In some embodiments, the set of index sequences comprises at least 80 or at least 96 different index sequences. In some embodiments, the set of index sequences comprises at least 112 or at least 384 different index sequences. In some embodiments, the set of index sequences comprises at least 734, at least 1,026 or at least 1,536 different index sequences.

[0295] In some embodiments, the set of index sequences comprises 4 subsets of 8 unique index sequences assigned as i5 sequences and 4 subsets of 12 unique index sequences assigned as i7 sequences. In some embodiments, the index sequences in a subset are color-balanced sequence pairs. In some embodiments, each subset comprises two or more pairs of index sequences to provide redundancy such that when any index needs to be replaced, the color-balanced index pairs can be replaced together with the redundant pairs in the subset. In some embodiments, the set of index sequences comprises 4 subsets of 12 unique index sequences assigned as i5 sequences and 4 subsets of 16 unique index sequences assigned as i7 sequences, for a total of 112 sequences.

[0296] In some embodiments, the set of index sequences comprises 4 subsets of index sequences, and each sequence in the subset can be applied as both an i5 index sequence and an i7 index sequence.

[0297] In some embodiments, the set of index sequences does not include any homopolymers having four or more consecutive identical bases. In some embodiments, the set of index sequences does not include index sequences that match or are reverse complementary to one or more sequencing primer sequences. In some embodiments, the sequencing primer sequence is included in the sequence of an oligonucleotide, such as Figure 2A the sequences shown in the dual-index adaptor (SP1 sequence or SP2 sequence). In some embodiments, the set of index sequences does not include index sequences that match or are reverse complementary to one or more flow cell amplification primer sequences, such as the P5 sequence or the P7 sequence (amplification primer sequences). In some embodiments, the flow cell amplification primer sequence is included in the oligonucleotide sequence, such as the P5 sequence and the P7' sequence at the 5' end and the 3' end of the bifurcated region of the Y-shaped adaptor.

[0298] In some embodiments, the set of index sequences does not include any subsequence of the sequence of an adapter or primer in the Illumina sequencing platform, or the reverse complement of such a subsequence. In some embodiments, the sequence of an adapter or primer in the Illumina sequencing platform comprises SEQ ID NO:1 (AGATGTGTATAAGAGACAG), SEQ ID NO:3 (TCGTCGGCAGCGTC), SEQ ID NO:5 (CCGAGCCCACGAGAC), SEQ ID NO:7 (CAAGCAGAAGACGGCATACGAGAT), and SEQ ID NO:8 (AATGATACGGCGACCACCGAGATCTACAC).

[0299] In some embodiments, the set of index sequences comprises index sequences having the same number of bases.

[0300] In some embodiments, each index sequence in the set of index sequences has a combined number of G and C bases between 2 and 6. In some embodiments, each index sequence has a guanine / cytosine (GC) content between 25% and 75%. In some embodiments, the set of oligonucleotides comprises DNA oligonucleotides or RNA oligonucleotides.

[0301] Figure 2B Different index oligonucleotide designs are shown, where only one strand of the Y-shaped adapter contains the index sequence. Figure 2B The sequencing adapter shown in Figure 2A is similar to the sequencing adapter in

[0302] Figure 2C except that the adapter contains the i7 index sequence only on the P7' arm of the Y-shaped adapter. The i7 index sequence is a member of the set of index sequences. The adapter does not contain an index sequence on its P5 arm.

[0303] The i5 index primer contains the i5 index sequence (210). The i7 index primer 206 contains the i7 index sequence. The i5 index primer 204 and the i7 index primer 206 can hybridize to the short universal adapter 214, which has a sequence complementary to Figure 2A and Figure 2BThe Y-shaped adaptor in [reference] is similar to the Y shape, except that the mismatched floppy ends of adaptor 202 are shorter and do not contain an index sequence or a flow cell amplification primer binding site. Instead, the index sequence and the flow cell amplification primer binding site are added to the adaptor via the i5 index primer 204 and the i7 index primer 206, through a nested PCR process such as that described in U.S. Patent No. 8,822,150, which is incorporated herein by reference in its entirety for all purposes.

[0304] The short universal adaptor 202 is common and shared among different samples, while Figure 2A the dual-index adaptor of [reference] and Figure 2B the single-index adaptor of [reference] are sample-specific. After attaching or ligating the short universal adaptor to the target nucleic acid fragment, the index-containing primers can be applied to the adaptor-target fragment in a sample-specific manner to allow identification of the origin of the sample. The i5 index primer 204 contains a P5 flow cell amplification primer binding site 208 at the 5′ end, an i5 index sequence 210 downstream of the P5 binding set, and a primer sequence 212 downstream of the i5 index sequence. The i7 index primer 206 contains a P7′ flow cell amplification primer binding site 216 at the 3′ end of the primer, an i7 index sequence upstream of the P7′ region, and a primer sequence 220 upstream of the i7 index sequence. When the i5 index primer 204 and the i7 index primer 206 are added to a reaction mixture containing the short universal adaptor 202 attached to the target fragment, the index sequence and the amplification primer binding site can be incorporated into the adaptor-target fragment through a PCR process (e.g., a nested PCR process) to provide a sequencing library containing the sample-specific index sequence.

[0305] Figure 2D Another index oligonucleotide design is shown, which involves index primers that can be used to bind to a double-stranded short universal adaptor. This design is similar to Figure 2C the design shown in [reference], but Figure 2D the short universal adaptor 232 in [reference] is double-stranded, rather than as Figure 2CThe Y shape shown by adapter 202. In addition, adapter 232 has blunt ends, rather than the T overhangs that adapter 202 has at 223. The i5 index primer 234 and the i7 index primer 236 can hybridize with the short universal adapter 232, thereby adding the relevant index sequences and amplification primer binding sites to the target sequence. The i5 index primer 234 includes a P5 flow cell amplification primer binding site 238 located at the 5′ end of the primer, an i5 index sequence 240 located downstream of the P5 binding site, and a primer sequence 242 located downstream of the i5 index sequence. The i5 index primer can be attached to the SP1 sequence primer binding site 244 of the double-stranded, short universal adapter 232. The i7 index primer 236 includes a P7′ flow cell amplification primer binding site 246 located at the 3′ end of the primer, an i7 index sequence 248 located upstream of the P7′ amplification primer binding site, and a primer sequence 250 located upstream of the i7 index sequence. Through a nested PCR reaction, the i5 index primer 234 and the i7 index primer 236 can be used to incorporate the index primers and amplification primer binding sites into the target sequence to provide a sequence library containing sample-specific index sequences.

[0306] In some embodiments, the set of index oligonucleotides is provided in a container comprising a plurality of separate compartments. In some embodiments, the container comprises a microtiter plate. Figures 4A - 4C Schematically shows a microtiter plate in which the index oligonucleotides can be provided. In some embodiments, each compartment contains a plurality of oligonucleotides, the plurality of oligonucleotides containing one index sequence from the set of index sequences. The index sequence in one compartment is different from the index sequences contained in other compartments. The oligonucleotides in each compartment can be applied to nucleic acid fragments from different sample sources to provide a mechanism for identifying the sample source.

[0307] In some embodiments, each compartment contains a first plurality of oligonucleotides, the first plurality of oligonucleotides containing a first index sequence from the set of index sequences. The compartment also contains a second plurality of oligonucleotides, the second plurality of oligonucleotides containing a second index sequence from the set of index sequences. The ordered combination of the first plurality of oligonucleotides and the second plurality of oligonucleotides is different from the ordered combination in any other compartment. The set of polynucleotides contains the first plurality of oligonucleotides and the second plurality of oligonucleotides.

[0308] Figure 4A The microtiter plate shown in contains an array of wells in the form of 8 rows and 12 columns, for a total of 96 compartments. In some embodiments, the array can have 16 rows and 24 columns, for a total of 384 compartments. In some embodiments, the set of oligonucleotides is provided in a Figure 4AIn the illustrated multi-well plate, each compartment in every 1 / 4 row contains oligonucleotides having at least a pair of color-balanced index sequences, and each compartment in every 1 / 4 column contains oligonucleotides having at least a pair of color-balanced indexes. In such a configuration, every quarter row and every quarter column can be used in a multiplex sequencing workflow. Thus, this configuration enables dual, triple, quadruple, sextuple, octuple, nonuple, and dodecaplex sequencing while fully utilizing the wells.

[0309] Figure 4B Shows the layout of the i5 index sequences in an 8×12 multi-well plate. The sequences are labeled such that 2n - 1 and 2n (where n is a positive integer) are color-balanced pairs. The i501 - i508 sequences can be selected from any subset of Table 1 or Table 3.

[0310] Figure 4C Shows the layout of the i7 index sequences. The i7 sequences are also organized in color-balanced pairs as described above. The i701 - i712 sequences can be selected from any subset of Table 2 or Table 3. For both the i5 sequences and the i7 sequences, when one sequence needs to be replaced due to various reasons such as poor performance or experimental considerations, its color-balanced pair should also be replaced. The removed color-balanced pair can be replaced with another color-balanced pair from the same subset in Tables 1 - 3. Such replacement will maintain the color balance of the plate. Figure 4B and Figure 4C The index sequence layouts shown in and are used for combinatorial dual-index applications. In other words, each well contains a first plurality of oligonucleotides and a second plurality of oligonucleotides, the first plurality of oligonucleotides containing a first index sequence of an index sequence set, and the second plurality of oligonucleotides containing a second index sequence of the index sequence set. The ordered combination of the first and second oligonucleotides in each compartment is different from the ordered combination of any other compartment.

[0311] In some embodiments, such as in the index sequence layouts illustrated in Figure 4B and Figure 4C the first plurality of oligonucleotides contains a P5 flow cell amplification primer binding site. The second plurality of oligonucleotides contains a P7′ flow cell amplification primer binding site. In some embodiments, such as in the embodiments illustrated in Figure 4B and Figure 4C the first plurality of oligonucleotides contains an i5 index sequence, and the second plurality of oligonucleotides contains an i7 index sequence.

[0312] In some embodiments, the set of oligonucleotides (including the first and second pluralities of oligonucleotides) is implemented as a Y-shaped adaptor containing index sequences, such as Figure 2A and Figure 2BThose Y-shaped adapters in. In some embodiments, the set of oligonucleotides provided in the plate includes double-stranded adapters containing index sequences. In some embodiments, the set of oligonucleotides includes primers containing index sequences, such as Figure 2C and Figure 2D the primers shown in.

[0313] In some embodiments, each index sequence in the first plurality of oligonucleotides is selected from a first subset of the set of index sequences, and each index sequence in the second plurality of oligonucleotides is selected from a second subset of the set of index sequences, and the first subset does not overlap with the second subset. In some embodiments, the Hamming distance between any two index sequences in the first subset or between any two index sequences in the second subset is not less than a second standard value. In some embodiments, the second standard value is greater than the first standard value. In some embodiments, the first standard value is 4 and the second standard value is 5. In other words, the Hamming distance between sequences within a subset is greater than the Hamming distance between sequences across subsets. In some applications, the larger Hamming distance within a subset can increase the probability of identifying index sequence reads containing errors (such as substitutions, insertions, or deletions). In some embodiments, the first subset is a subset selected from Table 1, and the second subset is a subset selected from Table 2. In some embodiments, the first subset contains i5 index sequences, and the second subset contains i7 index sequences.

[0314] In some embodiments, the index sequences are incorporated into sequencing adapters. In some embodiments, the sequencing adapter includes a Y-shaped sequencing adapter, where each sequencing adapter contains a first strand and a second strand, the first strand contains an index sequence selected from a first subset of the set of index sequences, and the second strand contains an index sequence selected from a second subset of the set of index sequences, and the first subset does not overlap with the second subset.

[0315] In some embodiments, the index sequences included in the first plurality of oligonucleotides and the second plurality of oligonucleotides are selected from the same subset of the set of index sequences. In some embodiments, the Hamming distance between any two index sequences in the same subset is not less than a second standard value. In some embodiments, the second standard value is greater than the first standard value. In some embodiments, the first standard value is 4 and the second standard value is 5. In some embodiments, the subset is selected from the subsets listed in Table 3. In some embodiments, the plurality of individual compartments of the multi-well plate are arranged in an array of one or more rows of compartments and one or more columns of compartments. In some embodiments, every 1 / n rows and / or every 1 / m columns of compartments contain oligonucleotides containing at least one pair of color-balanced index sequences, where n and m are each integers selected from the range of 1 to 24. In some embodiments, the plurality of individual compartments are arranged in an 8x12 array, as Figure 4A shown in.

[0316] Some embodiments provide oligonucleotides that consist primarily of multiple subsets of oligonucleotides. The set of oligonucleotides is configured to identify the source of a nucleic acid sample in multiplexed massively parallel sequencing, each of the nucleic acid samples comprising a plurality of nucleic acid molecules. Each subset of the multiple subsets of oligonucleotides comprises a unique index sequence, and the index sequences of the multiple subsets consist of a set of index sequences. The Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, where the first standard value is at least 2. The set of index sequences comprises multiple pairs of color-balanced index sequences, where any two bases at corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine (A) base or a cytosine (C) base, and (ii) a guanine (G) base, a thymine (T) base, or a uracil (U) base.

[0317] Construction of index oligonucleotides

[0318] Some embodiments provide methods for preparing multiple oligonucleotides for multiplexed massively parallel sequencing. The method includes selecting a set of index sequences from a pool of different index sequences. The set of index sequences includes at least six different sequences. The Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, where the first standard value is at least 2. The set of index sequences comprises multiple pairs of color-balanced index sequences. Any two bases at corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an A base or a C base, and (ii) a G base, a T base, or a U base.

[0319] In some embodiments, the process for selecting a set of index sequences from a pool of different index sequences proceeds according to Figure 5 steps 402 - 416 of process 400 therein. In some embodiments, selecting a set of index sequences from a pool of different index sequences includes selecting a candidate set of index sequences from the pool of index sequences; separating the selected candidate set into multiple groups of color-balanced index sequence pairs; and using a bipartite graph matching algorithm to partition each group into two subgroups of color-balanced pairs. Each color-balanced pair is a node in the bipartite graph.

[0320] Figure 5Process 400 for preparing indexed oligonucleotides such as indexed adaptors is shown. Process 400 involves providing a pool of all possible n-mer sequences. In some embodiments, the n-mer is an octamer. In some embodiments, the n-mer is a nonamer. In some embodiments, the n-mer is a decamer. Oligonucleotides of other sizes described herein can be generated similarly. See module 402. Process 400 also involves removing a subset of the indexed sequences from the pool of indexed sequences. See module 404. In some embodiments, the subset of indexed sequences removed includes indexed sequences having four or more consecutive identical bases. In some embodiments, the subset of indexed sequences removed includes indexed sequences having a combined number of G and C bases less than two and oligonucleotide sequences having a combined number of G and C bases greater than six. In some embodiments, the subset of indexed sequences removed includes indexed sequences having a sequence that matches or is reverse complementary to one or more sequencing primer sequences. In some embodiments, the sequencing primer sequence is included in the sequence of the indexed oligonucleotide, such as Figures 2A - 2D the sequences of the adaptor and primer shown in. In some embodiments, the subset of indexed sequences removed includes indexed sequences having a sequence that matches or is reverse complementary to one or more flow cell amplification primer sequences. In some embodiments, the flow cell amplification primer sequence is included in the sequence of the indexed oligonucleotide, such as Figures 2A - 2D the P5 sequence and the P7′ sequence in the adaptor and primer shown in. In some embodiments, the subset of indexed sequences removed includes indexed sequences empirically determined to have poor performance in indexing the source of nucleic acid samples in multiplex massively parallel sequencing. In some embodiments, the subset of indexed sequences removed includes the sequences in Table 4.

[0321] Process 400 continues by randomly selecting a pair of color-balanced sequences from the sequence pool. See module 406. Process 400 also involves adding the pair of color-balanced sequences to the candidate set and removing the pair from the pool. See module 408. Process 400 involves sorting the remaining indexed sequences in the indexed sequence pool based on the minimum Hamming distance from the members in the candidate set. See module 410. Process 400 also involves removing any remaining indexed sequences having a minimum Hamming distance from the members in the candidate set less than a first standard value or a minimum edit distance from the members in the candidate set less than a third standard value. In some embodiments, the first standard value is 4 and the third standard value is 3. See module 412.

[0322] Process 400 also involves determining whether any sequences remain in the pool. See module 414. If so, the process loops back to module 406 to randomly select a pair of color-balanced sequences from the sequence pool. See the "Yes" branch of decision module 414. If no sequences remain in the pool, process 400 proceeds to separate the candidate set into multiple groups of color-balanced pairs. See module 416. In some embodiments, the separation is performed by randomly selecting a seed for each of the multiple groups and greedily expanding each of the multiple groups. The greedy approach involves having each group take turns acquiring the farthest color-balanced pair remaining in the pool.

[0323] Process 400 also involves using a bipartite graph matching algorithm to partition each group into two subgroups of color-balanced pairs, where each color-balanced index sequence pair is a node in the bipartite graph. In the bipartite graph matching algorithm, two nodes are connected if the Hamming distance between them is less than a second criterion value, where the second criterion value is greater than the first criterion value. The matching algorithm produces two sets of index sequences. In some embodiments, the first criterion value is 4, and the second criterion value is 5. In some embodiments, one set can be used as the i5 index sequences, and the other set can be used as the i7 index sequences.

[0324] Then, process 400 involves synthesizing a plurality of oligonucleotides, where each oligonucleotide has at least one index sequence from the candidate set. In some embodiments, the plurality of oligonucleotides includes double-stranded sequencing adapters, where each strand of each double-stranded sequencing adapter includes an index sequence from the index sequence set. In some embodiments, the double-stranded sequencing adapter includes a first strand and a second strand, the first strand includes an index sequence selected from a first subset of the index sequence set, the second strand includes an index sequence selected from a second subset of the index sequence set, and the first subset does not overlap with the second subset. In some embodiments, the first strand of each double-stranded sequencing adapter includes a P5 flow cell amplification primer binding site, and the second strand of each double-stranded sequencing adapter includes a P7′ flow cell amplification primer binding site. Other forms of oligonucleotides described herein can be synthesized.

[0325] Sample

[0326] Samples that are used to determine the sequence of a DNA fragment can include samples obtained from any cell, fluid, tissue, or organ, the sample containing a nucleic acid in which the sequence of interest is to be determined. In some embodiments involving the diagnosis of cancer, circulating tumor DNA can be obtained from a subject's body fluid such as blood or plasma. In some embodiments involving the diagnosis of a fetus, it is advantageous to obtain cell-free nucleic acids, such as cell-free DNA (cfDNA), from maternal body fluid. Cell-free nucleic acids (including cell-free DNA) can be obtained from biological samples by a variety of methods known in the art, the biological samples including but not limited to plasma, serum, and urine (see, e.g., Fan et al., Proc Natl Acad Sci 105:16266-16271

[2008] ; Koide et al., Prenatal Diagnosis 25:604-607

[2005] ; Chen et al., Nature Med. 2:1033-1035

[1996] ; Lo et al., Lancet 350:485-487

[1997] ; Botezatu et al., Clin Chem. 46:1078-1084, 2000; and Su et al., J Mol. Diagn. 6:101-107

[2004] ).

[0327] In various embodiments, the nucleic acid (e.g., DNA or RNA) present in the sample can be specifically or non-specifically enriched prior to use (e.g., prior to preparing a sequencing library). Non-specific enrichment of sample DNA refers to whole genome amplification of genomic DNA fragments of the sample, which can be used to increase the level of sample DNA prior to preparing a cfDNA sequencing library. Methods for whole genome amplification are known in the art. Degenerate oligonucleotide-primed PCR (DOP), primer extension PCR technology (PEP), and multiple displacement amplification (MDA) are examples of whole genome amplification methods. In some embodiments, the sample is not enriched for DNA.

[0328] Samples that contain the nucleic acid to which the methods described herein are applied typically include biological samples (the "test samples") as described above. In some embodiments, the nucleic acid to be sequenced is purified or isolated by any of a number of well-known methods.

[0329] Accordingly, in certain embodiments, the sample comprises a purified or isolated polynucleotide or consists essentially of a purified or isolated polynucleotide, or it can be a sample such as a tissue sample, a biological fluid sample, a cell sample, etc. Suitable biological fluid samples include, but are not limited to, blood, plasma, serum, sweat, tears, sputum, urine, sputum, earflow, lymph fluid, saliva, cerebrospinal fluid, lavage fluid, bone marrow suspension, vaginal flow, transcervical lavage fluid, cerebral fluid, ascites, milk, secretions of the respiratory, intestinal, and urogenital tracts, amniotic fluid, milk, and leukapheresis samples. In some embodiments, the sample is a sample that can be readily obtained by a non-invasive procedure, such as blood, plasma, serum, sweat, tears, sputum, urine, feces, sputum, earflow, saliva, or feces. In certain embodiments, the sample is a peripheral blood sample, or the plasma and / or serum fraction of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy sample, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples. For example, the biological sample can include two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms "blood", "plasma", and "serum" expressly encompass their fractions or processed parts. Similarly, in the case where the sample is obtained from a biopsy, swab, smear, etc., the "sample" expressly encompasses the processed fractions or parts derived from the biopsy, swab, smear, etc.

[0330] In certain embodiments, the sample can be obtained from sources including, but not limited to: samples from different individuals, samples from different developmental stages of the same or different individuals, samples from different diseased individuals (e.g., individuals suspected of having a genetic disorder), normal individuals, samples obtained at different stages of an individual's disease, samples obtained from individuals undergoing different disease treatments, samples from individuals exposed to different environmental factors, samples from individuals with a pathological predisposition, samples from individuals exposed to an infectious disease agent, etc.

[0331] In an illustrative but non-limiting embodiment, the sample is a maternal sample obtained from a pregnant female, such as a pregnant woman. In such cases, the sample can be analyzed using the methods described herein to provide a prenatal diagnosis of potential chromosomal abnormalities in the fetus. The maternal sample can be a tissue sample, a biological fluid sample, or a cell sample. As non-limiting examples, the biological fluids include blood, plasma, serum, sweat, tears, sputum, urine, sputum, earflow, lymph fluid, saliva, cerebrospinal fluid, lavage fluid, bone marrow suspension, vaginal flow, transcervical lavage fluid, cerebral fluid, ascites, milk, secretions of the respiratory, intestinal, and urogenital tracts, and leukapheresis samples.

[0332] In certain embodiments, samples can also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. Cultured samples can be obtained from sources including, but not limited to: cultures (e.g., tissues or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissues or cells) maintained for different lengths of time, cultures (e.g., tissues or cells) treated with different factors or reagents (e.g., candidate drugs or modulators), or cultures of different types of tissues and / or cells.

[0333] Methods for isolating nucleic acids from biological sources are well known and will vary depending on the nature of the source. Those skilled in the art can readily isolate nucleic acids from the source as needed for the methods described herein. In some cases, it may be advantageous to fragment the nucleic acid molecules in the nucleic acid sample. The fragmentation can be random, or it can be specific, such as, for example, achieved using restriction endonuclease digestion. Methods for random fragmentation are well known in the art and include, for example, limited DNase digestion, alkali treatment, and physical shearing.

[0334] Sequencing library preparation

[0335] In various embodiments, sequencing can be performed on various sequencing platforms that require the preparation of a sequencing library. This preparation typically involves fragmenting the DNA (by sonication, nebulization, or shearing), followed by DNA repair and end polishing (to blunt ends or A overhangs), and the ligation of platform-specific adapters. In one embodiment, the methods described herein can utilize next-generation sequencing technology (NGS), which allows multiple samples to be sequenced individually as genomic molecules (i.e., single-end sequencing) or as a pooled sample containing indexed genomic molecules in a single sequencing run (e.g., multiplex sequencing). These methods can generate up to billions of DNA sequence reads. In various embodiments, the sequences of genomic nucleic acids and / or indexed genomic nucleic acids can be determined using, for example, the next-generation sequencing technology (NGS) described herein. In various embodiments, the analysis of the large amount of sequence data obtained using NGS can be performed using one or more processors as described herein.

[0336] In various embodiments, the use of such sequencing technologies does not involve the preparation of a sequencing library.

[0337] However, in certain embodiments, the sequencing methods contemplated herein involve the preparation of a sequencing library. In an illustrative method, sequencing library preparation involves the generation of a random collection of adapter-modified DNA fragments (e.g., polynucleotides) that are ready for sequencing. A sequencing library of polynucleotides can be prepared from DNA or RNA, including equivalents, analogs of DNA or cDNA, such as complementary DNA or cDNA, or copy DNA generated from an RNA template by the action of reverse transcriptase. The polynucleotides can originate in double-stranded form (e.g., dsDNA such as genomic DNA fragments, cDNA, PCR amplification products, etc.), or in certain embodiments, the polynucleotides can originate in single-stranded form (e.g., ssDNA, RNA, etc.) and have been converted to the dsDNA form. As an example, in certain embodiments, single-stranded mRNA molecules can be copied into double-stranded cDNA suitable for use in preparing a sequencing library. The exact sequence of the primary polynucleotide molecule is typically not important for the method of library preparation and can be known or unknown. In one embodiment, the polynucleotide molecule is a DNA molecule. More specifically, in certain embodiments, the polynucleotide molecule represents the entire gene complement or substantially the entire gene complement of an organism and is a genomic DNA molecule (e.g., cellular DNA, cell-free DNA (cfDNA), etc.), which typically contains both intron sequences and exon sequences (coding sequences), as well as non-coding regulatory sequences such as promoter sequences and enhancer sequences. In certain embodiments, the primary polynucleotide molecule comprises a human genomic DNA molecule, such as a cfDNA molecule present in the peripheral blood of a pregnant subject.

[0338] The preparation of a sequencing library for some NGS sequencing platforms is facilitated by using polynucleotides that comprise a specific range of fragment sizes. The preparation of such libraries typically involves the fragmentation of large polynucleotides (e.g., cellular genomic DNA) to obtain polynucleotides in the desired size range.

[0339] Paired-end reads can be used in the sequencing methods and systems disclosed herein. The fragment or insert length is longer than the read length and sometimes longer than the sum of the lengths of two reads.

[0340] In some illustrative embodiments, the sample nucleic acid is obtained as genomic DNA, which is fragmented into fragments longer than about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 base pairs, to which the NGS method can be readily applied. In some embodiments, paired-end reads are obtained from an insert of about 100 - 5000 bp. In some embodiments, the length of the insert is about 100 - 1000 bp. These are sometimes implemented as conventional short-insert paired-end reads. In some embodiments, the length of the insert is about 1000 - 5000 bp. These are sometimes implemented as long-insert mate-pair reads as described above.

[0341] In some embodiments, long inserts are designed for assessing very long sequences. In some embodiments, mate-pair reads can be applied to obtain reads separated by thousands of base pairs. In these embodiments, the insert or fragment is in the range of hundreds to thousands of base pairs and has two biotinylated linkers at both ends of the insert. The biotinylated linkers then ligate the two ends of the insert to form a circular molecule, which is then further fragmented. Sub-fragments containing the biotinylated linkers and both ends of the original insert are selected for sequencing on a platform designed to sequence shorter fragments.

[0342] Fragmentation can be achieved by any of a number of methods known to those of skill in the art. For example, fragmentation can be achieved by mechanical means, including but not limited to nebulization, sonication, and hydroshear. However, mechanical fragmentation typically cleaves the DNA backbone at C - O, P - O, and C - C bonds, producing a heterogeneous mixture of blunt ends and 3′- and 5′-overhangs with broken C - O, P - O, and C - C bonds (see, e.g., Alnemri and Liwack, J Biol.Chem 265:17323 - 17333

[1990] ; Richards and Boyer, J Mol Biol 11:327 - 240

[1965] ), and the broken C - O, P - O, and C - C bonds may need to be repaired because they may lack the 5′-phosphate necessary for subsequent enzymatic reactions, such as the ligation of sequencing adapters, which are required to prepare DNA for sequencing.

[0343] In contrast, cfDNA typically exists as fragments less than about 300 base pairs, and thus fragmentation is generally not required for generating a sequencing library using cfDNA samples.

[0344] Typically, whether polynucleotides are forcibly fragmented (e.g., in vitro fragmentation) or are naturally present as fragments, they are converted into blunt-ended DNA with 5′-phosphate and 3′-hydroxyl. Standard protocols, for example, using a protocol for sequencing on an Illumina platform as described in the exemplary workflow above referenced Figure 1A and Figure 1B direct the user to perform end repair on the sample DNA, purify the end-repaired product prior to adenylation or dA-tailing of the 3′ end, and purify the dA-tailed product prior to the adapter-ligation step of library preparation.

[0345] The various embodiments of the methods of sequence library preparation described herein eliminate the need to perform one or more steps typically enforced by standard protocols to obtain modified DNA products that can be sequenced by NGS. The Simplified Method (ABB Method), the 1-Step Method, and the 2-Step Method are examples of methods for preparing sequencing libraries, which are described in U.S. Patent Publication No. 2013 / 0029852A1, which is incorporated herein by reference in its entirety.

[0346] Sequencing method

[0347] The methods and devices described herein can employ next-generation sequencing technology (NGS), which allows for massively parallel sequencing. In certain embodiments, clonally amplified DNA templates or individual DNA molecules are sequenced in a massively parallel manner in a flow cell (e.g., as described in Volkerding et al Clin Chem 55:641-658

[2009] ; Metzker M Nature Rev11:31-46

[2010] ). Sequencing technologies for NGS include, but are not limited to, pyrosequencing, sequencing by synthesis using reversible dye terminators, sequencing by oligonucleotide probe ligation, and ion semiconductor sequencing. DNA from a single sample can be sequenced individually (i.e., singleplex sequencing), or DNA from multiple samples can be pooled and sequenced as barcoded genomic molecules in a single sequencing run (i.e., multiplex sequencing) to generate up to hundreds of millions of DNA sequence reads. Examples of sequencing technologies that can be used to obtain sequence information in accordance with the methods of the invention are further described herein.

[0348] Some sequencing technologies are commercially available, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, CA) and the sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, CT), Illumina / Solexa (Hayward, CA), and Helicos Biosciences (Cambridge, MA), as well as the sequencing-by-ligation platform from Applied Biosystems (Foster City, CA), as described below. In addition to single molecule sequencing using the sequencing-by-synthesis of Helicos Biosciences, other single molecule sequencing technologies include, but are not limited to, the SMRT TM technology of Pacific Biosciences, ION TORRENT TM technology, and nanopore sequencing such as that developed by Oxford Nanopore Technologies.

[0349] While the automated Sanger method is considered a "first generation" technology, Sanger sequencing, including automated Sanger sequencing, can also be used in the methods described herein. Additional suitable sequencing methods include, but are not limited to, nucleic acid imaging techniques such as atomic force microscopy (AFM) or transmission electron microscopy (TEM). Illustrative sequencing technologies are described in more detail below.

[0350] In some embodiments, the disclosed methods involve obtaining sequence information of nucleic acids in a test sample by performing massively parallel sequencing of millions of DNA fragments by using Illumina's synthesis sequencing and reversible terminator-based sequencing chemistry (e.g., as described in Bentley et al., Nature 6:53-59

[2009] ). The template DNA can be genomic DNA, such as cellular DNA or cfDNA. In some embodiments, genomic DNA from isolated cells is used as a template and is fragmented into lengths of several hundred base pairs. In other embodiments, cfDNA or circulating tumor DNA (ctDNA) is used as a template and fragmentation is not required because cfDNA or ctDNA exists as short fragments. For example, fetal cfDNA circulates in the bloodstream as fragments of approximately 170 base pairs (bp) in length (Fan et al., Clin Chem 56:1279-1286

[2010] ) and does not require fragmentation of the DNA prior to sequencing. Illumina sequencing technology relies on attaching fragmented genomic DNA to a planar, optically transparent surface to which oligonucleotide anchors are bound. The template DNA is end-repaired to generate 5'-phosphorylated blunt ends, and the polymerase activity of Klenow fragment is used to add a single A base to the 3' end of the phosphorylated DNA fragment with a blunt end. This addition prepares the DNA fragment for ligation to an oligonucleotide adaptor that has a single T-base overhang at its 3' end to increase ligation efficiency. The adaptor oligonucleotide is complementary to the flow cell anchor oligonucleotide. Under limiting dilution conditions, the adaptor-modified single-stranded template DNA is added to the flow cell and is immobilized to the anchor oligonucleotide by hybridization. The attached DNA fragments are extended and bridge amplified to produce an ultra-high density sequencing flow cell with hundreds of millions of clusters, each cluster containing approximately 1,000 copies of the same template. In one embodiment, randomly fragmented genomic DNA is amplified using PCR prior to its undergoing cluster amplification. Alternatively, a non-amplified genomic library preparation is used and cluster amplification alone is used to enrich randomly fragmented genomic DNA (Kozarewa et al., Nature Methods 6:291-295

[2009] ). In some applications, the template is sequenced using a robust four-color DNA synthesis sequencing technology that employs reversible terminators with removable fluorescent dyes. High-sensitivity fluorescence detection is achieved using laser excitation and total internal reflection optics. Short sequence reads of approximately dozens to hundreds of base pairs are aligned to a reference genome, and unique mapping of the short sequence reads to the reference genome is identified using specially developed data analysis pipeline software. After the first read is completed, the template can be regenerated in situ to enable a second read from the opposite end of the fragment.Thus, single-end or paired-end sequencing of DNA fragments can be used.

[0352] Various embodiments of the present disclosure can use synthetic sequencing that allows paired-end sequencing. In some embodiments, synthetic sequencing by Illumina involves clustering of fragments. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, as in the examples described herein, the fragment has two different adapters attached to the two ends of the fragment, and the adapters allow the fragment to hybridize to two different oligonucleotides on the surface of the flow cell lane. The fragment also contains two index sequences located at the two ends of the fragment or ligated to two index sequences at the two ends of the fragment, and the index sequences provide tags to identify different samples in multiplexed sequencing. In some sequencing platforms, the fragment to be sequenced from both ends is also referred to as an insert.

[0353] In some embodiments, the flow cell used for clustering in the Illumina platform is a slide with lanes. Each lane is a glass channel coated with a lawn of two types of oligonucleotides (e.g., P5 oligonucleotides and P7′ oligonucleotides). Hybridization is achieved by the first type of the two types of oligonucleotides on the surface. This oligonucleotide is complementary to the first adapter on one end of the fragment. The polymerase generates a complementary strand of the hybridized fragment. The double-stranded molecule is denatured, and the original template strand is washed away. The remaining strand, in parallel with many other remaining strands, is clonally amplified by bridge amplification.

[0354] In bridge amplification and other sequencing methods involving clustering, the strand folds, and the second adapter region on the second end of the strand hybridizes to the second type of oligonucleotide on the flow cell surface. The polymerase generates a complementary strand, forming a double-stranded bridge molecule. The double-stranded molecule is denatured, resulting in two single-stranded molecules that are tethered to the flow cell by the two different oligonucleotides. Then, the process is repeated over and over again and occurs simultaneously for millions of clusters, resulting in clonal amplification of all fragments. After bridge amplification, the reverse strand is cleaved and washed away, leaving only the forward strand. The 3′ end is blocked to prevent unwanted priming.

[0355] After clustering, sequencing begins by extending a first sequencing primer to generate a first read. In each cycle, fluorescently labeled nucleotides compete for addition to the growing strand. Based on the template sequence, only one nucleotide is incorporated. After adding each nucleotide, the cluster is excited by a light source and emits a characteristic fluorescent signal. The number of cycles determines the length of the read. The emission wavelength and signal intensity determine the base call. For a given cluster, all identical strands are read simultaneously. Hundreds of millions of clusters are sequenced in a massively parallel manner. After completion of the first read, the read product is washed away.

[0356] In the next step of the protocol involving two indexing primers, the index 1 primer is introduced and hybridizes to the index 1 region on the template. The index region provides identification of the fragment, which is useful for demultiplexing samples during multiplexed sequencing. The index 1 read is generated similar to the first read. After completion of the index 1 read, the read product is washed away, and the 3′ end of the strand is deprotected. The template strand then folds and binds to a second oligonucleotide on the flow cell. The index 2 sequence is read in the same manner as index 1. The index 2 read product is then washed away upon completion of this step.

[0357] After reading both indexes, read 2 begins by using polymerase to extend the second flow cell oligonucleotide, forming a double-stranded bridge. The double-stranded DNA is denatured, and the 3′ end is blocked. The original forward strand is cleaved off and washed away, leaving the reverse strand. Read 2 starts with the introduction of the read 2 sequencing primer. As with read 1, the sequencing steps are repeated until the desired length is reached. The read 2 product is washed away. This entire process generates millions of reads, representing all the fragments. The sequences from the pooled sample libraries are separated based on the unique indexes introduced during sample preparation. For each sample, reads of base calls for similar segments are locally clustered. The forward and reverse reads are paired, generating contiguous sequences. These contiguous sequences are aligned to a reference genome for variant identification.

[0358] The example of synthetic sequencing described above involves paired-end reads, which are used in many embodiments of the disclosed methods. Paired-end sequencing involves two reads from both ends of a fragment. Paired-end reads are used to resolve ambiguous alignments. Paired-end sequencing allows the user to select the length of the insert (or fragment to be sequenced) and sequence either end of the insert, generating high-quality, alignable sequence data. Since the distance between each paired read is known, alignment algorithms can use this information to more precisely map the reads over repetitive regions. This results in better alignment of the reads, especially those spanning difficult-to-sequence repetitive regions in the genome. Paired-end sequencing can detect rearrangements, including insertions and deletions (indels) as well as inversions.

[0359] Paired-end reads can use inserts of different lengths (i.e., different sizes of the fragments to be sequenced). By default in the present disclosure, paired-end reads are used to refer to reads obtained from various insert lengths. In some cases, to distinguish short-insert paired-end reads from long-insert paired-end reads, the latter are specifically referred to as mate-pair reads. In some embodiments involving mate-pair reads, two biotinylated adapters are first attached to the two ends of a relatively long insert (e.g., several kb). Then, the biotinylated adapters ligate the two ends of the insert to form a circular molecule. Then, sub-fragments encompassing the biotinylated adapters can be obtained by further fragmenting the circular molecule. Then, sub-fragments containing the two ends of the original fragment in opposite sequence order can be sequenced by the same procedures as described above for short-insert paired-end sequencing. Additional details of mate-pair sequencing using the Illumina platform are shown in the online publication at the following address, which is incorporated by reference in its entirety: wwwdot illumina dot com / documents / products / technotes / technote_nextera_matepair_data_processin g.pdf.

[0360] After sequencing DNA fragments, sequence reads of a predetermined length (e.g., 100 bp) are localized by mapping to a known reference genome (aligning with a known reference genome). The mapped reads and their corresponding positions on the reference sequence are also referred to as tags. In another embodiment of the procedure, localization is achieved by k-mer sharing and read-to-read alignment. The analysis of many embodiments disclosed herein utilizes poorly matching or non-matching reads, as well as matching reads (tags). In one embodiment, the reference genome sequence is the NCBI36 / hg18 sequence, which is available on the World Wide Web at genome dot ucsc dot edu / cgi-bin / hgGateway?org=Human&db=hg18&hgsid=166260105. Optionally, the reference genome sequence is GRCh37 / hg19 or GRCh38, which is available on the World Wide Web at genome dot ucsc dot edu / cgi-bin / hgGateway. Other sources of public sequence information include GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (DNA Databank of Japan). Many computer algorithms can be used to align sequences, including but not limited to BLAST (Altschul et al., 1990), BLITZ (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988), BOWTIE (Langmead et al., Genome Biology 10:R25.1-R25.10

[2009] ), or ELAND (Illumina, Inc., San Diego, CA, USA). In one embodiment, one end of a copy of a plasma cfDNA molecule that has been clonally amplified is sequenced and processed by bioinformatics alignment analysis on an Illumina Genome Analyzer that uses Efficient Large-Scale Alignment of Nucleotide Databases (ELAND) software.

[0361] Other sequencing methods can also be used to obtain sequence reads and their alignments. Additional suitable methods are described in U.S. Patent Publication No. 2016 / 0319345A1, which is incorporated herein by reference in its entirety.

[0362] In some embodiments of the methods described herein, the sequence reads are about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. It is expected that technological advancements will enable single-end reads greater than 500 bp to be achieved, which when generating paired-end reads enables reads greater than about 1000 bp to be achieved. In some embodiments, paired-end reads are used to determine a sequence of interest, which comprises sequence reads of about 20 bp to 1000 bp, about 50 bp to 500 bp, or 80 bp to 150 bp. In various embodiments, paired-end reads are used to evaluate a sequence of interest. The sequence of interest is longer than the read. In some embodiments, the sequence of interest is longer than about 100 bp, 500 bp, 1000 bp, or 4000 bp. Mapping of the sequence reads is accomplished by comparing the sequence of the read to a reference sequence to determine the chromosomal origin of the nucleic acid molecule being sequenced and does not require specific gene sequence information. A small degree of mismatch (0 - 2 mismatches / read) may be allowed to account for minor polymorphisms that may exist between the reference genome and the genomes in the mixed sample. In some embodiments, reads that match the reference sequence are used as anchor reads, and reads that are paired with the anchor reads but do not match or poorly match the reference are used as anchored reads. In some embodiments, poorly matching reads may have a relatively large percentage of mismatches / read, e.g., at least about 5%, at least about 10%, at least about 15%, or at least about 20% mismatches / read.

[0363] Typically, multiple sequence tags (i.e., reads that match the reference sequence) are obtained for each sample. In some embodiments, for each sample, at least about 3 x 10 6 sequence tags of 100 bp, at least about 5 x 10 6 sequence tags, at least about 8 x 10 6 sequence tags, at least about 10 x 10 6 sequence tags, at least about 15 x 10 6 sequence tags, at least about 20 x 10 6 sequence tags, at least about 30 x 10 6 sequence tags, at least about 40 x 10 6 sequence tags, or at least about 50 x 10 6sequence tags. In some embodiments, all sequence reads are mapped to all regions of a reference genome to provide genome-wide reads. In other embodiments, the reads are mapped to sequences of interest.

[0364] Devices and systems for preparing index oligonucleotides

[0365] As should be apparent, certain embodiments of the present invention employ processes that operate under the control of instructions and / or data stored in or transmitted through one or more computer systems. Certain embodiments also relate to apparatus for performing these operations. The apparatus may be specially designed and / or constructed for the required purposes, or it may be a general purpose computer selectively configured by one or more computer programs and / or data structures stored in the computer or otherwise available to the computer. In particular, various general purpose machines may be used with programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized apparatus to perform the required method steps. Specific structures of various such machines are shown and described below.

[0366] Certain embodiments also provide functionality (e.g., code and processes) for storing any of the results (e.g., query results) or data structures generated as described herein. Such results or data structures are typically stored at least temporarily on a computer-readable medium. The results or data structures may also be output in any of a variety of ways such as display, printing, etc.

[0367] Examples of tangible computer-readable media suitable for use with computer program products and computing devices of the present invention include, but are not limited to, magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD-ROM disks; magneto-optical media; semiconductor memory devices (e.g., flash memory), and hardware devices specially configured to store and execute program instructions such as read-only memory devices (ASIC) and random access memory (ROM), and sometimes application specific integrated circuits (ASIC), programmable logic devices (PLD), and signal transmission media for conveying computer-readable instructions such as local area networks, wide area networks, and the Internet. The data and program instructions provided herein may also be embodied in a carrier wave or other transmission medium (including electronic or optical conduction paths). The data and program instructions of the present invention may also be embodied in a carrier wave or other transmission medium (e.g., optical lines, electrical lines, and / or radio waves).

[0368] Examples of program instructions include low-level code such as that produced by a compiler, and high-level code that may be executed by a computer using an interpreter. In addition, the program instructions may be machine code, source code, and / or any other code that directly or indirectly controls the operation of a computing machine. The code may specify inputs, outputs, computations, conditions, branches, iterative loops, etc.

[0369] The analysis of sequencing data and the diagnostics derived therefrom typically use a variety of computer-executed algorithms and programs. Thus, certain embodiments employ processes involving data stored in or transmitted through one or more computer systems or other processing systems. The embodiments disclosed herein also relate to apparatuses for performing these operations. The apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer (or group of computers) selectively activated or reconfigured by a computer program and / or data structures stored in the computer. In some embodiments, groups of processors cooperate (e.g., via a network or cloud computing) and / or perform some or all of the recited analysis operations in parallel. The processor or group of processors for performing the methods described herein may be various types of processors or groups of processors, including microcontrollers and microprocessors, such as programmable devices (e.g., CPLDs and FPGAs) and non-programmable devices such as gate array ASICs or general-purpose microprocessors.

[0370] One embodiment provides a system for determining sequences in a plurality of test samples containing nucleic acids, the system including a sequencer configured to receive nucleic acid samples and provide nucleic acid sequence information from the samples; a processor; and a machine-readable storage medium having stored thereon instructions for execution on the processor to determine a sequence of interest in a test sample by the methods described above.

[0371] In some embodiments of any of the systems provided herein, the sequencer is configured to perform next-generation sequencing (NGS). In some embodiments, the sequencer is configured to use reversible dye terminators and perform massively parallel sequencing using sequencing by synthesis. In other embodiments, the sequencer is configured to perform ligation sequencing. In still other embodiments, the sequencer is configured to perform single-molecule sequencing.

[0372] Another embodiment provides a system that includes a nucleic acid synthesizer, a processor, and a machine-readable storage medium having instructions stored thereon for execution on the processor to prepare sequencing adapters. The instructions include: (a) code for adding to a candidate set of index sequences a color-balanced pair of index sequences randomly selected from a pool of different index sequences, wherein any two bases at corresponding sequence positions of each pair of color-balanced index sequences include both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base; (b) code for sorting the remaining index sequences in the pool of index sequences based on the minimum Hamming distance from members of the candidate set; (c) code for removing any remaining index sequences having a minimum Hamming distance from members of the candidate set less than a first standard value or a minimum edit distance from members of the candidate set less than a second standard value; (d) code for repeating (a)-(c) to maximize the size of the candidate set; and (e) code for selecting from the candidate set a set of index sequences to be incorporated into a set of oligonucleotides configured for multiplexed massively parallel sequencing.

[0373] In some embodiments, the instructions include: (a) code for receiving a plurality of index reads and a plurality of target reads of target sequences obtained from target nucleic acids from a plurality of samples, wherein each target read contains a target sequence obtained from a target nucleic acid of a sample from the plurality of samples, each index read contains an index sequence obtained from a target nucleic acid of a sample from the plurality of samples, the index sequence being selected from a set of index sequences, each target read is associated with at least one index read, each sample in the plurality of samples is uniquely associated with one or more index sequences in the set of index sequences, and the Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, wherein the first standard value is at least 2; (b) code for identifying, in the plurality of target reads, a subset of target reads associated with an index read that matches at least one index sequence uniquely associated with a particular sample in the plurality of samples; and (c) code for determining the target sequence of the particular sample based on the identified subset of target reads.

[0374] In addition, certain embodiments relate to tangible and / or non-transitory computer-readable media or computer program products that include program instructions and / or data (including data structures) for performing various computer-implemented operations. Examples of computer-readable media include, but are not limited to, semiconductor memory devices, magnetic media such as disk drives, magnetic tapes, optical media such as CDs, magneto-optical media, and hardware devices such as read-only memory devices (ROM) and random access memory (RAM) that are specifically configured to store and execute program instructions. The computer-readable media may be directly controlled by an end user, or the media may be indirectly controlled by the end user. Examples of directly controlled media include media located at a user facility and / or media not shared with other entities. Examples of indirectly controlled media include media indirectly accessible by a user via an external network and / or via a service that provides shared resources such as a "cloud". Examples of program instructions include both machine code such as that produced by a compiler and files containing high-level code that can be executed by a computer using an interpreter.

[0375] In various embodiments, data or information employed in the disclosed methods and devices is provided in electronic format. Such data or information may include reads and tags derived from nucleic acid samples, reference sequences (including reference sequences that provide polymorphisms only or primarily), determinations such as cancer diagnosis determinations, counseling advice, diagnoses, etc. As used herein, data or other information provided in electronic format can be used to be stored on a machine and transmitted between machines. Conventionally, data in electronic format is provided digitally and can be stored as bits and / or bytes in various data structures, lists, databases, etc. The data can be embodied electronically, optically, etc.

[0376] One embodiment provides a computer program product for generating an output indicative of the sequence of a DNA fragment of interest in a test sample. The computer product may contain instructions for performing any one or more of the methods described above for determining the sequence of interest. As explained, the computer product may include a non-transitory and / or tangible computer-readable medium having recorded thereon computer-executable or compilable logic (e.g., instructions) for causing a processor to be able to determine the sequence of interest. In one example, the computer product includes a computer-readable medium having recorded thereon computer-executable or compilable logic (e.g., instructions) for causing a processor to be able to diagnose a condition or determine a nucleic acid sequence of interest.

[0377] It should be understood that for an unassisted human, performing the computational operations of the methods disclosed herein is not practical, or even impossible in most cases. For example, mapping a single 30bp read from a sample to any one of the human chromosomes may require years of effort without the aid of a computing device. Of course, the problem is complex because reliable determination of low allele frequency mutations typically requires mapping thousands (e.g., at least about 10,000) or even millions of reads to one or more chromosomes.

[0378] The methods disclosed herein can be performed using a system for determining sequences of interest in a plurality of test samples. The system can include: (a) a sequencer configured to receive nucleic acids from a test sample and provide nucleic acid sequence information from the sample; (b) a processor; and (c) one or more computer-readable storage media having instructions stored thereon that are executable on the processor to determine sequences of interest in the test sample. In some embodiments, the method is directed by a computer-readable medium having computer-readable instructions stored thereon for performing a method for determining sequences of interest. Accordingly, one embodiment provides a computer program product that includes a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for determining the sequences of nucleic acid fragments in a plurality of test samples.

[0379] In some embodiments, the program code or instructions may also include automatically recording information related to the method. Patient medical records may be maintained by, for example, a laboratory, a physician's office, a hospital, a health maintenance organization, an insurance company, or a personal medical record website. Additionally, based on the results of the processor-implemented analysis, the method may also involve prescribing, initiating, and / or changing the treatment of a human subject from whom the test sample was obtained. This may involve performing one or more additional tests or analyses on additional samples obtained from the subject.

[0380] The disclosed methods can also be performed using a computer processing system adapted or configured to perform a method for determining sequences of interest. One embodiment provides a computer processing system adapted or configured to perform the methods as described herein. In one embodiment, the device includes a sequencing device adapted or configured to sequence at least a portion of the nucleic acid molecules in a sample to obtain the type of sequence information described elsewhere herein. The device may also include components for processing the sample. Such components are described elsewhere herein.

[0381] Sequences or other data can be input directly or indirectly into a computer or stored on a computer-readable medium. In one embodiment, a computer system is directly connected to a sequencing device that reads and / or analyzes the sequence of nucleic acids from a sample. The sequence or other information from such a tool is provided via an interface in the computer system. Alternatively, the sequences processed by the system are provided from sequence storage sources such as databases or other repositories. After being made available to the processing device, a memory device or mass storage device at least temporarily buffers or stores the nucleic acid sequences. In addition, the memory device can store tag counts for various chromosomes or genomes, etc. The memory can also store various routines and / or programs for analyzing the presented sequences or mapping data. Such programs / routines can include programs for performing statistical analysis, etc.

[0382] In one example, a user provides a sample to a sequencing device. Data is collected and / or analyzed by the sequencing device connected to a computer. Software on the computer allows for data collection and / or analysis. The data can be stored, presented (via a monitor or other similar device), and / or sent to another location. The computer can be connected to the Internet, which is used to transmit the data to a handheld device used by a remote user (e.g., a physician, scientist, or analyst). It should be understood that the data can be stored and / or analyzed before transmission. In some embodiments, the raw data is collected and sent to a remote user or to a device that analyzes and / or stores the data. The transmission can occur via the Internet, but can also occur via satellite or other connections. Alternatively, the data can be stored on a computer-readable medium, and the medium can be shipped to the end user (e.g., via mail). The remote user can be in the same or different geographical locations, including but not limited to buildings, cities, states, countries, or continents.

[0383] In some embodiments, the method further includes collecting data regarding a plurality of polynucleotide sequences (e.g., reads, tags, and / or reference chromosomal sequences), and sending the data to a computer or other computing system. For example, a computer can be connected to laboratory equipment such as sample collection equipment, nucleotide amplification equipment, nucleotide sequencing equipment, or hybridization equipment. The computer can then collect the applicable data collected by the laboratory devices. The data can be stored on the computer at any step, e.g., at the time of real-time collection, before sending, during sending, or in conjunction with or after sending. The data can be stored on a computer-readable medium that can be retrieved from the computer. The collected or stored data can be transmitted from the computer to a remote location, e.g., via a local area network or a wide area network such as the Internet. At the remote location, various operations can be performed on the transmitted data, as described below.

[0384] The types of electronic format data that can be stored, transmitted, analyzed, and / or manipulated in the systems, devices, and methods disclosed herein are the following types:

[0385] a) Reads obtained by sequencing nucleic acids in a test sample

[0386] b) Tags obtained by aligning the reads with a reference genome or one or more other reference sequences

[0387] c) Reference genome or sequence

[0388] d) Thresholds for determining a test sample as affected, unaffected, or no determination

[0389] e) Actual determination of a medical condition related to a sequence of interest

[0390] f) Diagnosis (clinical condition associated with the determination)

[0391] g) Recommendations for further testing resulting from the determination and / or diagnosis

[0392] h) Treatment and / or monitoring plan resulting from the determination and / or diagnosis

[0393] These different types of data can be obtained, stored, transmitted, analyzed, and / or manipulated at one or more locations using different devices. The processing options span a broad spectrum. At one end of the spectrum, all or much of the information is stored and used at the location where the test sample is processed (e.g., a doctor's office or other clinical setting). At the other extreme, a sample is obtained at one location, the sample is processed and optionally sequenced at a different location, the reads are aligned and determined at one or more different locations, and a diagnosis, recommendation, and / or plan are prepared at still another location (which may be the location where the sample was obtained).

[0394] In various embodiments, reads are generated by a sequencing device and then transmitted to a remote site where the reads are processed to determine a sequence of interest. As an example, at the remote location, the reads are aligned with a reference sequence to produce anchored reads and anchored read segments. Processing operations that can be employed at different locations are the following operations:

[0395] a) Sample collection

[0396] b) Sample processing prior to sequencing

[0397] c) Sequencing

[0398] d) Analyze sequence data and obtain a medical determination

[0399] e) Diagnosis

[0400] f) Reporting the diagnosis and / or determination to the patient or health care provider

[0401] g) Developing a plan for further treatment, testing, and / or monitoring

[0402] h) Executing the plan

[0403] i) Consulting

[0404] Any one or more of these operations may be automated, as described elsewhere herein. Typically, the sequencing and analysis of sequence data and obtaining a medical determination will be performed computationally. Other operations may be performed manually or automatically.

[0405] Figure 6 An embodiment of a decentralized system for generating a determination or diagnosis from multiple test samples is shown. Sample collection location 01 is used to obtain test samples. The samples are then provided to processing and sequencing location 03, where the test samples may be processed and sequenced as described above. Location 03 includes equipment for processing the samples and equipment for sequencing the processed samples. As described elsewhere herein, the result of the sequencing is a collection of reads, which typically are provided in electronic format and are provided to a network such as the Internet, as indicated by reference numeral 05 in Figure 6 the drawings.

[0406] The sequence data is provided to a remote location 07, where analysis and determination generation are performed. This location may include one or more powerful computing devices such as a computer or a processor. After the computing resources at location 07 have completed their analysis and generated a determination from the received sequence information, the determination is transmitted back over network 05. In some embodiments, not only a determination but also an associated diagnosis are generated at location 07. The determination and / or diagnosis are then transmitted over the network and transmitted back to sample collection location 01, as Figure 6 illustrated in the drawings. As explained, this is merely one of many variations on how the various operations associated with generating a determination or diagnosis are divided among the various locations. A common variation involves providing sample collection, as well as processing and sequencing, at a single location. Another variation involves providing processing and sequencing at the same location as analysis and determination generation.

[0407] Figure 7A typical computer system is illustrated in a simple modular format that, when appropriately configured or designed, can be used as a computing device in accordance with certain embodiments. Computer system 2000 includes any number of processors 2002 (also referred to as central processing units or CPUs), which are connected to a memory device that includes main memory 2006 (typically random access memory or RAM) and main memory 2004 (typically read only memory or ROM). The CPU 2002 can be of various types, including microcontrollers and microprocessors, such as programmable devices (e.g., CPLDs and FPGAs) and non-programmable devices such as gate array ASICs or general purpose microprocessors. In the depicted embodiment, main memory 2004 is used to transfer data and instructions unidirectionally to the CPU, and main memory 2006 is typically used to transfer data and instructions in a bidirectional manner. These two main memory devices can include any suitable computer-readable medium, such as the computer-readable media described above. Mass storage device 2008 is also connected to main memory 2006 bidirectionally and provides additional data storage capacity and can include any of the computer-readable media described above. The mass storage device 2008 can be used to store programs, data, etc., and is typically an auxiliary storage medium such as a hard disk. Such programs, data, etc. are often temporarily copied to main memory 2006 for execution on the CPU 2002. It will be understood that, where appropriate, information retained in the mass storage device 2008 can be incorporated in a standard manner as part of main memory 2004. A particular mass storage device such as a CD-ROM 2014 can also transfer data unidirectionally to the CPU or main memory.

[0408] The CPU 2002 is also connected to an interface 2010 that is connected to one or more input / output devices, such as a nucleic acid sequencer (2020), a nucleic acid synthesizer (2022), a video monitor, a trackball, a mouse, a keyboard, a microphone, a touch-sensitive display, a transducer reader, a magnetic or paper tape reader, a tablet computer, a stylus, a voice or handwriting recognition peripheral, a USB port, or other well-known input devices, and of course, such as other computers. Finally, the CPU 2002 can optionally be connected to an external device, such as a database or a computer or telecommunications network, using an external connection generally shown at 2012. Through such connections, it is contemplated that, during the course of performing the method steps described herein, the CPU can receive information from the network or can output information to the network. In some embodiments, a nucleic acid sequencer or a nucleic acid synthesizer can be communicatively connected to the CPU 2002 via the network connection 2012 without going through the interface 2010 or in addition to going through the interface 2010 via the network connection 2012.

[0409] In one embodiment, a system such as computer system 2000 is used as a data import, data association, and query system capable of performing some or all of the tasks described herein. Information and programs, including data files, may be provided via network connection 2012 for access or download by researchers. Alternatively, such information, programs, and files may be provided to researchers on a memory device.

[0410] In a specific embodiment, computer system 2000 is directly connected to a data acquisition system such as a microarray, a high-throughput screening system, or a nucleic acid sequencer (2020) that captures data from samples. Data from such systems is provided via interface 2010 for analysis by system 2000. Alternatively, data processed by system 2000 is provided from a data storage source such as a database or other repository of related data. Once in device 2000, a memory device such as main memory 2006 or mass storage 2008 at least temporarily buffers or stores the relevant data. The memory may also store various routines and / or programs for importing, analyzing, and presenting data, including code for selecting and / or validating index sequences, determining sequence reads, and correcting errors in the reads, etc.

[0411] In certain embodiments, the computers used herein may include user terminals, which may be any type of computer (e.g., desktop computer, laptop computer, tablet computer, etc.), media computing platform (e.g., cable, satellite TV set-top box, digital video recorder, etc.), handheld computing device (e.g., PDA, email client, etc.), mobile phone, or any other type of computing or communication platform.

[0412] In certain embodiments, the computers used herein may also include a server system in communication with the user terminal. The server system may include server devices or decentralized server devices and may include mainframe computers, minicomputers, supercomputers, personal computers, or combinations thereof. Multiple server systems may also be used without departing from the scope of the invention. The user terminal and the server system may communicate with each other via a network. The network may include, for example, wired networks such as LAN (local area network), WAN (wide area network), MAN (metropolitan area network), ISDN (integrated services digital network), etc., and wireless networks such as wireless LAN, CDMA, Bluetooth, and satellite communication networks, etc., without limiting the scope of the invention.

[0413] The present disclosure may be embodied in other specific forms without departing from the spirit or essential characteristics of the present disclosure. The described embodiments are to be considered in all respects only illustrative and not restrictive. Thus, the scope of the present disclosure is indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within the scope of the claims.

[0414] Experiments

[0415] Example 1

[0416] Index sequence verification

[0417] Computer simulation experiments were conducted to verify the effectiveness and efficacy of the index sequences according to some embodiments. The index sequences satisfy the following conditions.

[0418] ● The index sequence has no direct match with an 8-mer subsequence (or reverse complement) of the sequencing adapter or primer of the sequencing platform

[0419] ○SBS491

[0420] ○P7

[0421] ○P5

[0422] ○SBS3

[0423] ● No four-nucleotide homopolymer appears

[0424] ● The G / C content is 2 to 6

[0425] ● The sequences listed in Table 4 (known poor performers) are not used

[0426] ● Design 1: The index sequences are selected from those index sequences in Table 1 and Table 2

[0427] ○ No two sequences have a Hamming distance < 4

[0428] ● Design 2: The index sequences are selected from those index sequences in Table 3

[0429] ○ Hamming distance

[0430] ■ Among any pairs in the I5 sequences or I7 sequences of a given plate, the Hamming distance > = 5

[0431] ■ Among all sequences (I5 and I7) across the entire design, the Hamming distance between any pair > = 4

[0432] ■ Among all sequences (I5 and I7) across the entire design, the Hamming distance between any pair > = 4

[0433] ○ Edit distance

[0434] ■ For all sequences (I5 and I7) across the entire design, the "edit" distance between any pair >= 3

[0435] Computer simulation experiments generated 1 million randomly varied index sequences through 3 edit operations, and no direct matches were found among the resulting sequences. The experiment generated 1 million randomly varied index sequences with 1 deletion and 1 substitution operation, and no direct matches were found among the resulting sequences.

[0436] The resulting sequences were assigned to a microplate layout as shown in Figure 4. It was found that no i5 / i7 pair occurred more than once. The designed pairs are color-balanced. The triples (quarter rows) are color-balanced. The quadruples (half columns) are color-balanced. The hextuples (half rows) are color-balanced. The octuples (full columns) are color-balanced. The designed pools (doubles, triples, quadruples, hextuples, octuples) do not have duplicate pairs (i.e., they mitigate index hopping).

[0437] The experimental results demonstrate that the index sequences and index oligonucleotides provided according to some embodiments can detect and correct errors in index sequence reads, thereby providing more accurate sample indexing in multiplexed massively parallel sequencing.

[0438] Example 2

[0439] 8 - mer index set

[0440] This example describes the design considerations for an 8-mer index sequence set according to some embodiments and lists the sequences in the index set.

[0441] This index set supports a larger number of indexes and maintains the index length at 8 bp. Compared with the index design of Example 1, the edit distance threshold is reduced to zero, and the Hamming distance threshold is reduced to 3.

[0442] An overview of the design strategy is given below.

[0443] ● The index sequence has no direct match with the 8-mer subsequence (or reverse complement) of the sequencing platform adapter or primer sequence, such as

[0444] ○ SBS491

[0445] ○ P7

[0446] ○ P5

[0447] ○ SBS3

[0448] ● No four-nucleotide homopolymers are present

[0449] ● The GC content is between 25% and 75% (inclusive of the endpoints)

[0450] ● The sequences listed in Table 4 are not used (known poor performers)

[0451] ● The minimum Hamming distance is 3

[0452] ● The minimum modified edit distance is 2

[0453] ● Indexers are provided as color balance pairs. Each pair is labeled with the numbers 2n - 1 and 2n, where n is a positive integer.

[0454] A total of 734 sequences were obtained, as listed below. They include all the sequences tested in Example 1 and shown in Table 3.

[0455] scytale-001: GAACCGCG; scytale-002: AGGTTATA; scytale-003: TCATCCTT; scytale-004: CTGCTTCC; scytale-005: GGTCACGA; scytale-006: AACTGTAG; scytale-007: GTGAATAT; scytale-008: ACAGGCGC; scytale-009: CATAGAGT; scytale-010: TGCGAGAC; scytale-011: GACGTCTT; scytale-012: AGTACTCC; scytale-013: TGGCCGGT; scytale-014: CAATTAAC; scytale-015: ATAATGTG; scytale-016: GCGGCACA; scytale-017: CTAGCGCT; scytale-018: TCGATATC; scytale-019: CGTCTGCG; scytale-020: TACTCATA; scytale-021: ACGCACCT; scytale-022: GTATGTTC; scytale-023: CGCTATGT; scytale-024: TATCGCAC; scytale-025: TCTGTTGG; scytale-026: CTCACCAA; scytale-027: TATTAGCT; scytale-028: CGCCGATC; scytale-029: TCTCTACT; scytale-030: CTCTCGTC; scytale-031: CCAAGTCT; scytale-032: TTGGACTC; scytale-033: GGCTTAAG; scytale-034: AATCCGGA; scytale-035: TAATACAG; scytale-036: CGGCGTGA; scytale-037: ATGTAAGT; scytale-038: GCACGGAC; scytale-039: GGTACCTT; scytale-040: AACGTTCC; scytale-041: GCAGAATT; scytale-042: ATGAGGCC; scytale-043: ACTAAGAT; scytale-044: GTCGGAGC; scytale-045: CCGCGGTT; scytale-046: TTATAACC; scytale-047: GGACTTGG;scytale-048: AAGTCCAA; scytale-049: ATCCACTG; scytale-050: GCTTGTCA; scytale-051: CAAGCTAG; scytale-052: TGGATCGA; scytale-053: AGTTCAGG; scytale-054: GACCTGAA; scytale-055: TGACGAAT; scytale-056: CAGTAGGC; scytale-057: AGCCTCAT; scytale-058: GATTCTGC; scytale-059: TCGTAGTG; scytale-060: CTACGACA; scytale-061: TAAGTGGT; scytale-062: CGGACAAC; scytale-063: ATATGGAT; scytale-064: GCGCAAGC; scytale-065: AAGATACT; scytale-066: GGAGCGTC; scytale-067: ATGGCATG; scytale-068: GCAATGCA; scytale-069: GTTCCAAT; scytale-070: ACCTTGGC; scytale-071: CTTATCGG; scytale-072: TCCGCTAA; scytale-073: GCTCATTG; scytale-074: ATCTGCCA; scytale-075: CTTGGTAT; scytale-076: TCCAACGC; scytale-077: CCGTGAAG; scytale-078: TTACAGGA; scytale-079: GGCATTCT; scytale-080: AATGCCTC; scytale-081: TACCGAGG; scytale-082: CGTTAGAA; scytale-083: CACGAGCG; scytale-084: TGTAGATA; scytale-085: GATCTATC; scytale-086: AGCTCGCT; scytale-087: CGGAACTG; scytale-088: TAAGGTCA; scytale-089: TTGCCTAG; scytale-090: CCATTCGA; scytale-091: ACACTAAG; scytale-092: GTGTCGGA; scytale-093: TTCCTGTT; scytale-094: CCTTCACC;scytale-095: GCCACAGG; scytale-096: ATTGTGAA; scytale-097: ACTCGTGT; scytale-098: GTCTACAC; scytale-099: GTTCGCCG; scytale-100: ACCTATTA; scytale-101: ATATCTCG; scytale-102: GCGCTCTA; scytale-103: AACAGGTT; scytale-104: GGTGAACC; scytale-105: CAACAATG; scytale-106: TGGTGGCA; scytale-107: AGGCAGAG; scytale-108: GAATGAGA; scytale-109: TGCGGCGT; scytale-ll0: CATAATAC; scytale-l1l: GAGGATGG; scytale-ll2: AGAAGCAA; scytale-113: TAGAGCCG; scytale-114: CGAGATTA; scytale-l15: CCGGTCAT; scytale-ll6: TTAACTGC; scytale-117: TTCAGTTG; scytale-118: CCTGACCA; scytale-ll9: TGAGTACG; scytale-120: CAGACGTA; scytale-121: AACCAACA; scytale-122: GGTTGGTG; scytale-123: TCTACCAG; scytale-124: CTCGTTGA; scytale-125: AGAGCTGT; scytale-126: GAGATCAC; scytale-127: TAGGCAAT; scytale-128: CGAATGGC; scytale-129: ACGGTGCG; scytale-130: GTAACATA; scytale-131: TTGTTCCT; scytale-132: CCACCTTC; scytale-133: ACTAGACG; scytale-134: GTCGAGTA; scytale-135: TGCCATCG; scytale-136: CATTGCTA; scytale-137: GAGCGCGT; scytale-138: AGATATAC; scytale-139: TTCGTCAG; scytale-140: CCTACTGA; scytale-141: CGTGTATT;scytale-142: TACACGCC; scytale-143: GTAGCCGG; scytale-144: ACGATTAA; scytale-145: ACCGGAAT; scytale-146: GTTAAGGC; scytale-147: CACTCCGG; scytale-148: TGTCTTAA; scytale-149: CCTGCGTG; scytale-150: TTCATACA; scytale-151: AAGCATTC; scytale-152: GGATGCCT; scytale-153: CGCAGGAG; scytale-154: TATGAAGA; scytale-155: CAACTCCT; scytale-156: TGGTCTTC; scytale-157: GATGGAAG; scytale-158: AGCAAGGA; scytale-159: CAGTCTCT; scytale-160: TGACTCTC; scytale-161: GCCTATAT; scytale-162: ATTCGCGC; scytale-163: GTGTGTAG; scytale-164: ACACACGA; scytale-165: GCCACGCT; scytale-166: ATTGTATC; scytale-167: CTGATTGT; scytale-168: TCAGCCAC; scytale-169: CTACACCG; scytale-170: TCGTGTTA; scytale-171: AATGTGGC; scytale-172: GGCACAAT; scytale-173: AGCGGAGG; scytale-174: GATAAGAA; scytale-175: TGAAGGTT; scytale-176: CAGGAACC; scytale-177: GTATTCCG; scytale-178: ACGCCTTA; scytale-179: CCTTACTT; scytale-180: TTCCGTCC; scytale-181: ATTGTTGT; scytale-182: GCCACCAC; scytale-183: CAGTGGTG; scytale-184: TGACAACA; scytale-185: TGTGCCTG; scytale-186: CACATTCA; scytale-187: AAGCCAGT; scytale-188: GGATTGAC;scytale-189: TTGGAGAT; scytale-190: CCAAGAGC; scytale-191: TCCTATGG; scytale-192: CTTCGCAA; scytale-193: TACATCTG; scytale-194: CGTGCTCA; scytale-195: AGGTGCCG; scytale-196: GAACATTA; scytale-197: CTGACATT; scytale-198: TCAGTGCC; scytale-199: ATTAATGG; scytale-200: GCCGGCAA; scytale-201: GTGGTGGT; scytale-202: ACAACAAC; scytale-203: ATATACCT; scytale-204: GCGCGTTC; scytale-205: GACGCAGT; scytale-206: AGTATGAC; scytale-207: TGGACGCG; scytale-208: CAAGTATA; scytale-209: GATTATAG; scytale-210: AGCCGCGA; scytale-211: CCTTGCGG; scytale-212: TTCCATAA; scytale-213: ACACCGTT; scytale-214: GTGTTACC; scytale-215: CGTCACAT; scytale-216: TACTGTGC; scytale-217: AACGTGTG; scytale-218: GGTACACA; scytale-219: CAACGGAT; scytale-220: TGGTAAGC; scytale-221: TCGGCGCT; scytale-222: CTAATATC; scytale-223: GACAACCG; scytale-224: AGTGGTTA; scytale-225: TATATCGT; scytale-226: CGCGCTAC; scytale-227: TTATCAGT; scytale-228: CCGCTGAC; scytale-229: TCACGTCG; scytale-230: CTGTACTA; scytale-231: GTTACTTG; scytale-232: ACCGTCCA; scytale-233: ACCAAGTG; scytale-234: GTTGGACA; scytale-235: CGGTTAGG;scytale-236: TAACCGAA; scytale-237: AATTATCC; scytale-238: GGCCGCTT; scytale-239: GCATTGTT; scytale-240: ATGCCACC; scytale-241: TAGAGATT; scytale-242: CGAGAGCC; scytale-243: TCACAAGG; scytale-244: CTGTGGAA; scytale-245: GCTATTAT; scytale-246: ATCGCCGC; scytale-247: TCTGGCCG; scytale-248: CTCAATTA; scytale-249: AATTACGA; scytale-250: GGCCGTAG; scytale-251: AGATCAAT; scytale-252: GAGCTGGC; scytale-253: CACACACT; scytale-254: TGTGTGTC; scytale-255: TATGCGAG; scytale-256: CGCATAGA; scytale-257: ACAACTGG; scytale-258: GTGGTCAA; scytale-259: GAATACGT; scytale-260: AGGCGTAC; scytale-261: AAGGAATA; scytale-262: GGAAGGCG; scytale-263: CTCCATCT; scytale-264: TCTTGCTC; scytale-265: GCGCCTGT; scytale-266: ATATTCAC; scytale-267: CCTGTACG; scytale-268: TTCACGTA; scytale-269: ATTCGGTT; scytale-270: GCCTAACC; scytale-271: TGGATAAT; scytale-272: CAAGCGGC; scytale-273: GTTAACAG; scytale-274: ACCGGTGA; scytale-275: CAATTCTG; scytale-276: TGGCCTCA; scytale-277: CGGCTCGT; scytale-278: TAATCTAC; scytale-279: TGGTGATG; scytale-280: CAACAGCA; scytale-281: GTAGATGT; scytale-282: ACGAGCAC;scytale-283: ATCCTTAG; scytale-284: GCTTCCGA; scytale-285: AACAGACC; scytale-286: GGTGAGTT; scytale-287: ACGGCCAG; scytale-288: GTAATTGA; scytale-289: TGCTAACT; scytale-290: CATCGGTC; scytale-291: AGATAGTT; scytale-292: GAGCGACC; scytale-293: CTAGTAAG; scytale-294: TCGACGGA; scytale-295: CACTGGCT; scytale-296: TGTCAATC; scytale-297: GCCTGTTG; scytale-298: ATTCACCA; scytale-299: GTGGCTCT; scytale-300: ACAATCTC; scytale-301: AATGTTAG; scytale-302: GGCACCGA; scytale-303: TACCAATT; scytale-304: CGTTGGCC; scytale-305: CTAACAGG; scytale-306: TCGGTGAA; scytale-307: TCTAATCG; scytale-308: CTCGGCTA; scytale-309: GTGATGCG; scytale-310: ACAGCATA; scytale-311: GAATGTAT; scytale-312: AGGCACGC; scytale-313: CGCGACAG; scytale-314: TATAGTGA; scytale-315: CCGCTTGG; scytale-316: TTATCCAA; scytale-317: CCTACAAT; scytale-318: TTCGTGGC; scytale-319: AGCAATAT; scytale-320: GATGGCGC; scytale-321: GAGCCATG; scytale-322: AGATTGCA; scytale-323: TTGTGGTT; scytale-324: CCACAACC; scytale-325: TCAGGAGT; scytale-326: CTGAAGAC; scytale-327: AAGTCGCG; scytale-328: GGACTATA; scytale-329: ATTACACT;scytale-330: GCCGTGTC; scytale-331: AACATTGG; scytale-332: GGTGCCAA; scytale-333: CGGCCTTG; scytale-334: TAATTCCA; scytale-335: CCGTACCG; scytale-336: TTACGTTA; scytale-337: AATCGAAT; scytale-338: GGCTAGGC; scytale-339: TCTCAGGT; scytale-340: CTCTGAAC; scytale-341: TTAGGTGG; scytale-342: CCGAACAA; scytale-343: CACGTTAT; scytale-344: TGTACCGC; scytale-345: CTACCGTG; scytale-346: TCGTTACA; scytale-347: GTTCAGCT; scytale-348: ACCTGATC; scytale-349: AATTAGTG; scytale-350: GGCCGACA; scytale-351: ATGAGCGT; scytale-352: GCAGATAC; scytale-353: GCTGTCGT; scytale-354: ATCACTAC; scytale-355: AATGTACA; scytale-356: GGCACGTG; scytale-357: TTCGAAGG; scytale-358: CCTAGGAA; scytale-359: TGATTCAT; scytale-360: CAGCCTGC; scytale-361: AACCGCCT; scytale-362: GGTTATTC; scytale-363: AGAATTCG; scytale-364: GAGGCCTA; scytale-365: GAGATAGG; scytale-366: AGAGCGAA; scytale-367: CCATGATT; scytale-368: TTGCAGCC; scytale-369: AGGTGTGT; scytale-370: GAACACAC; scytale-371: CTCGCACG; scytale-372: TCTATGTA; scytale-373: TATTGTCG; scytale-374: CGCCACTA; scytale-375: TACACCAT; scytale-376: CGTGTTGC;scytale-377: ACCGATAG; scytale-378: GTTAGCGA; scytale-379: CCTCCGCT; scytale-380: TTCTTATC; scytale-381: AAGAATGA; scytale-382: GGAGGCAG; scytale-383: GAGTTGAT; scytale-384: AGACCAGC; scytale-385: CCAACCTG; scytale-386: TTGGTTCA; scytale-387: AGGTACAT; scytale-388: GAACGTGC; scytale-389: CACCTGGT; scytale-390: TGTTCAAC; scytale-391: TTGCGCTG; scytale-392: CCATATCA; scytale-393: CATATAAG; scytale-394: TGCGCGGA; scytale-395: ATTGACTT; scytale-396: GCCAGTCC; scytale-397: ACCAGCGG; scytale-398: GTTGATAA; scytale-399: TGTCCACG; scytale-400: CACTTGTA; scytale-401: AAGCTAAC; scytale-402: GGATCGGT; scytale-403: TCAATTGT; scytale-404: CTGGCCAC; scytale-405: GCGGTATG; scytale-406: ATAACGCA; scytale-407: AATCAAGC; scytale-408: GGCTGGAT; scytale-409: AGGTATTG; scytale-410: GAACGCCA; scytale-411: TCACATAT; scytale-412: CTGTGCGC; scytale-413: CGAATCAG; scytale-414: TAGGCTGA; scytale-415: ACTCGGAG; scytale-416: GTCTAAGA; scytale-417: AGGCGATT; scytale-418: GAATAGCC; scytale-419: CTTAAGTT; scytale-420: TCCGGACC; scytale-421: TCCGACAT; scytale-422: CTTAGTGC; scytale-423: CAAGGACT;scytale-424: TGGAAGTC; scytale-425: GCTTCATG; scytale-426: ATCCTGCA; scytale-427: TTCGCCTT; scytale-428: CCTATTCC; scytale-429: TGACTGAG; scytale-430: CAGTCAGA; scytale-431: ATATTAGG; scytale-432: GCGCCGAA; scytale-433: AATTGGAC; scytale-434: GGCCAAGT; scytale-435: CACACTTG; scytale-436: TGTGTCCA; scytale-437: GTGATCTT; scytale-438: ACAGCTCC; scytale-439: CTGGATAG; scytale-440: TCAAGCGA; scytale-441: TGTCGTTG; scytale-442: CACTACCA; scytale-443: TACCTGCG; scytale-444: CGTTCATA; scytale-445: CGACAGGT; scytale-446: TAGTGAAC; scytale-447: GTTAGTCT; scytale-448: ACCGACTC; scytale-449: AGTAACCG; scytale-450: GACGGTTA; scytale-451: ACGGTAGT; scytale-452: GTAACGAC; scytale-453: TTAGCTAT; scytale-454: CCGATCGC; scytale-455: GTGCAACG; scytale-456: ACATGGTA; scytale-457: CGACCATT; scytale-458: TAGTTGCC; scytale-459: AGCTACGG; scytale-460: GATCGTAA; scytale-461: AATAGCTG; scytale-462: GGCGATCA; scytale-463: CCGTCGAT; scytale-464: TTACTAGC; scytale-465: GTTGAATG; scytale-466: ACCAGGCA; scytale-467: TACAATAG; scytale-468: CGTGGCGA; scytale-469: TCTTCTGT; scytale-470: CTCCTCAC;scytale-471: CGATGACG; scytale-472: TAGCAGTA; scytale-473: GTCTTGTG; scytale-474: ACTCCACA; scytale-475: GCGAGCTG; scytale-476: ATAGATCA; scytale-477: TGTATGCT; scytale-478: CACGCATC; scytale-479: GAACTAAT; scytale-480: AGGTCGGC; scytale-481: GCCGAGGT; scytale-482: ATTAGAAC; scytale-483: CTGAACCT; scytale-484: TCAGGTTC; scytale-485: GATACCGG; scytale-486: AGCGTTAA; scytale-487: CTATGTGT; scytale-488: TCGCACAC; scytale-489: TACGCTCG; scytale-490: CGTATCTA; scytale-491: TAACGCTT; scytale-492: CGGTATCC; scytale-493: GCCTTAGT; scytale-494: ATTCCGAC; scytale-495: ACATAATG; scytale-496: GTGCGGCA; scytale-497: ATTGGCAG; scytale-498: GCCAATGA; scytale-499: AACCTCGC; scytale-500: GGTTCTAT; scytale-501: CATCAACT; scytale-502: TGCTGGTC; scytale-503: GCAGACCG; scytale-504: ATGAGTTA; scytale-505: TAGTCCGT; scytale-506: CGACTTAC; scytale-507: TCAATCCG; scytale-508: CTGGCTTA; scytale-509: AGCGTGGT; scytale-510: GATACAAC; scytale-511: GAAGAGTG; scytale-512: AGGAGACA; scytale-513: CTGTTCAG; scytale-514: TCACCTGA; scytale-515: TTCTGACG; scytale-516: CCTCAGTA; scytale-517: CGAACGAT;scytale-518: TAGGTAGC; scytale-519: AGTGGACT; scytale-520: GACAAGTC; scytale-521: TTCACCGG; scytale-522: CCTGTTAA; scytale-523: GCTAGTGG; scytale-524: ATCGACAA; scytale-525: TACCGTAT; scytale-526: CGTTACGC; scytale-527: CAATCGTT; scytale-528: TGGCTACC; scytale-529: AAGATGTC; scytale-530: GGAGCACT; scytale-531: TCCTCAAG; scytale-532: CTTCTGGA; scytale-533: ATACGATG; scytale-534: GCGTAGCA; scytale-535: CTCATAAT; scytale-536: TCTGCGGC; scytale-537: TGGCACTT; scytale-538: CAATGTCC; scytale-539: AATATGCG; scytale-540: GGCGCATA; scytale-541: CCGGATCT; scytale-542: TTAAGCTC; scytale-543: GCTGTGAG; scytale-544: ATCACAGA; scytale-545: AACTACTT; scytale-546: GGTCGTCC; scytale-547: TTAGGCCT; scytale-548: CCGAATTC; scytale-549: ACTTCTAG; scytale-550: GTCCTCGA; scytale-551: AAGCGGAA; scytale-552: GGATAAGG; scytale-553: TATCCTGG; scytale-554: CGCTTCAA; scytale-555: GTGCCGTT; scytale-556: ACATTACC; scytale-557: CGGAGTCG; scytale-558: TAAGACTA; scytale-559: CCAGCCGT; scytale-560: TTGATTAC; scytale-561: ATAGAGAG; scytale-562: GCGAGAGA; scytale-563: CACCGCTG; scytale-564: TGTTATCA;scytale-565: GAGTAACT; scytale-566: AGACGGTC; scytale-567: CTTGGAGG; scytale-568: TCCAAGAA; scytale-569: TTACTTCT; scytale-570: CCGTCCTC; scytale-571: AATCTCTA; scytale-572: GGCTCTCG; scytale-573: TCGTGCAT; scytale-574: CTACATGC; scytale-575: AAGACAAG; scytale-576: GGAGTGGA; scytale-577: ACTTAACT; scytale-578: GTCCGGTC; scytale-579: AAGGAGGT; scytale-580: GGAAGAAC; scytale-581: CGTCAAGG; scytale-582: TACTGGAA; scytale-583: AGAGGCTT; scytale-584: GAGAATCC; scytale-585: CAGCTATT; scytale-586: TGATCGCC; scytale-587: CTATATTG; scytale-588: TCGCGCCA; scytale-589: TGCACTGT; scytale-590: CATGTCAC; scytale-591: GTCCAGAG; scytale-592: ACTTGAGA; scytale-593: TAACCACT; scytale-594: CGGTTGTC; scytale-595: AACATCAA; scytale-596: GGTGCTGG; scytale-597: ATGTCTAT; scytale-598: GCACTCGC; scytale-599: TCGACTTG; scytale-600: CTAGTCCA; scytale-601: AACGCGAC; scytale-602: GGTATAGT; scytale-603: TATCTGAT; scytale-604: CGCTCAGC; scytale-605: TCCGTTCT; scytale-606: CTTACCTC; scytale-607: TCAAGATG; scytale-608: CTGGAGCA; scytale-609: GTATCAAG; scytale-610: ACGCTGGA; scytale-611: GATCACTT;scytale-612: AGCTGTCC; scytale-613: GAGCTTAG; scytale-614: AGATCCGA; scytale-615: CTCCGAGT; scytale-616: TCTTAGAC; scytale-617: GCCAACTT; scytale-618: ATTGGTCC; scytale-619: CGAGTGTG; scytale-620: TAGACACA; scytale-621: TCGGCAGG; scytale-622: CTAATGAA; scytale-623: TGCCGGCT; scytale-624: CATTAATC; scytale-625: ATGCTCCG; scytale-626: GCATCTTA; scytale-627: GATAGCAT; scytale-628: AGCGATGC; scytale-629: AGGACGTT; scytale-630: GAAGTACC; scytale-631: ACCTTCAG; scytale-632: GTTCCTGA; scytale-633: TGTCTCGG; scytale-634: CACTCTAA; scytale-635: CGGTGCTT; scytale-636: TAACATCC; scytale-637: CTAAGTAG; scytale-638: TCGGACGA; scytale-639: GCAGCGAT; scytale-640: ATGATAGC; scytale-641: GTCGTACT; scytale-642: ACTACGTC; scytale-643: CCACTAGT; scytale-644: TTGTCGAC; scytale-645: CATGAGAT; scytale-646: TGCAGAGC; scytale-647: TAAGCATG; scytale-648: CGGATGCA; scytale-649: AACCAGAT; scytale-650: GGTTGAGC; scytale-651: AGTATATG; scytale-652: GACGCGCA; scytale-653: CCAGATGG; scytale-654: TTGAGCAA; scytale-655: TCGTTCGG; scytale-656: CTACCTAA; scytale-657: AAGCGTCG; scytale-658: GGATACTA;scytale-659: CGTAGCCT; scytale-660: TACGATTC; scytale-661: GTCTGATT; scytale-662: ACTCAGCC; scytale-663: CATGCTGT; scytale-664: TGCATCAC; scytale-665: TGAACTAG; scytale-666: CAGGTCGA; scytale-667: CAGCCGAG; scytale-668: TGATTAGA; scytale-669: ACTGTCTG; scytale-670: GTCACTCA; scytale-671: ACAGAGCT; scytale-672: GTGAGATC; scytale-673: CGTTCCAG; scytale-674: TACCTTGA; scytale-675: CGAGTTCT; scytale-676: TAGACCTC; scytale-677: CTCTAGAT; scytale-678: TCTCGAGC; scytale-679: GCTAGGTT; scytale-680: ATCGAACC; scytale-681: CCACGCAG; scytale-682: TTGTATGA; scytale-683: CCGACACG; scytale-684: TTAGTGTA; scytale-685: CACAACGT; scytale-686: TGTGGTAC; scytale-687: AGCGCTTG; scytale-688: GATATCCA; scytale-689: ATACTGGT; scytale-690: GCGTCAAC; scytale-691: TCTTCGCG; scytale-692: CTCCTATA; scytale-693: GTCTCCGT; scytale-694: ACTCTTAC; scytale-695: TGTAAGAG; scytale-696: CACGGAGA; scytale-697: AAGGACCG; scytale-698: GGAAGTTA; scytale-699: AAGAGTAT; scytale-700: GGAGACGC; scytale-701: CATGGTTG; scytale-702: TGCAACCA; scytale-703: GCTCTAGG; scytale-704: ATCTCGAA; scytale-705: AGGCTGCT;scytale-706: GAATCATC; scytale-707: TGGAATGG; scytale-708: CAAGGCAA; scytale-709: AGACATCT; scytale-710: GAGTGCTC; scytale-711: ATCTAGCG; scytale-712: GCTCGATA; scytale-713: GTCAGAAG; scytale-714: ACTGAGGA; scytale-715: ACTATCCT; scytale-716: GTCGCTTC; scytale-717: GTACGCAT; scytale-718: ACGTATGC; scytale-719: GATCCTCT; scytale-720: AGCTTCTC; scytale-721: TGTGCAGT; scytale-722: CACATGAC; scytale-723: CTTACGAG; scytale-724: TCCGTAGA; scytale-725: GCACCTAG; scytale-726: ATGTTCGA; scytale-727: CTTCTCTT; scytale-728: TCCTCTCC; scytale-729: GAGAGGAG; scytale-730: AGAGAAGA; scytale-731: CGATCTGG; scytale-732: TAGCTCAA; scytale-733: AGTAGTAG; scytale-734: GACGACGA;

[0456] Example 3

[0457] 10 - mer index set

[0458] This example describes the design considerations for a set of 10-mer index sequences according to some embodiments and lists the sequences in the index set. The index set supports a larger number of indices and keeps the index length at 10 bp. An overview of the design strategy is given below.

[0459] ● The index sequences do not have a direct match to the 10-mer subsequence (or reverse complement) of the sequencing platform adapter or primer sequences, such as

[0460] ○ SBS491

[0461] ○ P7

[0462] ○ P5

[0463] ○SBS3

[0464] ●No four-nucleotide homopolymers appear

[0465] ●The GC content is between 25% and 75% (including the endpoints)

[0466] ●The sequences listed in Table 4 (known poor performers) are not used

[0467] ●The minimum Hamming distance is 4

[0468] ●The minimum modified edit distance is 3

[0469] ●The indexers are provided as color balance pairs

[0470] A total of 1026 decamer sequences were obtained, and they were combined and shown in the combined sequence of SEQ ID NO:9. The nth decamer sequence includes nucleotides 10(n - 1)+1, 10(n - 1)+2, 10(n - 1)+3, ……, 10(n - 1)+10 in SEQ ID NO:9. The (2m - 1)th decamer and the 2mth decamer in SEQ ID NO:9 are color balanced, where m is an integer between 1 and 513. In some embodiments, an index sequence set is obtained. The index sequence set includes 1026 decamer sequences of SEQ ID NO:9. In some embodiments, oligonucleotides are generated using the index sequences. The oligonucleotides include double-stranded or Y-shaped sequencing adapters. Each strand of each double-stranded or Y-shaped sequencing adapter contains an index sequence corresponding to a decamer in SEQ ID NO:9. In other embodiments, the oligonucleotide set includes single-stranded oligonucleotide pairs, such as primer pairs. Each pair of single-stranded oligonucleotides is provided together in a reagent. Each oligonucleotide in a pair contains an index sequence corresponding to a decamer in SEQ ID NO:9.

[0471] Ten subsets are selected from the 1026 decamers in SEQ ID NO:9 as subsets of the index sequences. To select the subsets of decamers, sequencing is performed using double-stranded sequencing adapters. Each strand of the double-stranded sequencing adapter contains an index sequence from the 1026 sequences in SEQ ID NO:9. Different index sequence pairs are used to generate different sequencing adapters. The sequencing performance of the adapters is measured. Based on the measured sequencing performance, the index sequences are sorted. Using the sorting of the index sequences as a criterion, one or more subsets of decamer sequences can be selected from the 1026 sequences.

[0472] In some embodiments, a subset of 96 different pairs of index sequences with the highest sequencing performance is selected. In one embodiment, each pair of the 96 index sequence pairs comprises the nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11.

[0473] The 10-mers in each of the ten subsets are combined and shown as combined sequences. The ten subsets respectively correspond to ten combined sequences: SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, and SEQ ID NO:19. The nth 10-mer sequence in the combined sequence comprises nucleotides 10(n - 1)+1, 10(n - 1)+2, 10(n - 1)+3, ……, 10(n - 1)+10. The (2m - 1)th 10-mer and the 2mth 10-mer in SEQ ID NO:10 to SEQ ID NO:19 are color balanced, where m is a positive integer. Each 10-mer is different from any other 10-mer in the subset. Each 10-mer in any one of the five subsets (corresponding to SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:16, and SEQ ID NO:18) is different from all other 10-mers in any one of the five subsets. Each 10-mer in any one of the other five subsets (corresponding to SEQ ID NO:11, SEQ ID NO:13, SEQ ID NO:15, SEQ ID NO:17, and SEQ ID NO:19) is different from all other 10-mers in any one of the other five subsets.

[0474] In some embodiments, each Y-shaped or double-stranded sequencing adaptor comprises a first strand and a second strand, the first strand comprising a first index sequence selected from a first subset of an index sequence set, and the second strand comprising a second index sequence selected from a second subset (or the reverse complement of the second subset) of the index sequence set. In some embodiments, each pair of oligonucleotides (e.g., primers) comprises a first oligonucleotide and a second oligonucleotide, the first oligonucleotide comprising a first index sequence selected from a first subset of an index sequence set, and the second oligonucleotide comprising a second index sequence selected from a second subset (or the reverse complement of the second subset) of the index sequence set.

[0475] In some embodiments, the first index sequence and the second index sequence are, respectively: the nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 (or the reverse complement of SEQ ID NO:11); the nth 10-mer in SEQ ID NO:12 and the nth 10-mer in SEQ ID NO:13 (or its reverse complement); the nth 10-mer in SEQ ID NO:14 and the nth 10-mer in SEQ ID NO:15 (or its reverse complement); the nth 10-mer in SEQ ID NO:16 and the nth 10-mer in SEQ ID NO:17 (or its reverse complement); the nth 10-mer in SEQ ID NO:18 and the nth 10-mer in SEQ ID NO:19 (or its reverse complement).

[0476] In some embodiments, the first index sequence and the second index sequence are included in an oligonucleotide, which is provided in a reaction compartment of a container comprising a plurality of separate compartments. Each compartment contains (a) a first plurality of oligonucleotides comprising the first index sequence and (b) a second plurality of oligonucleotides comprising the second index sequence, and the ordered combination of (a) and (b) in the compartment is different from the ordered combination of (a) and (b) in any other compartment.

[0477] In some embodiments, the container comprises a microplate. In some embodiments, the container contains 8x12 compartments. When the compartments are labeled as rows A-H and columns 1-12, they can be listed as A1, A2, A3, ……, A12, B1, B2, ……, B12, ……, H1, H2, ……, H12. In some embodiments, in the nth compartment in the list, the first index sequence and the second index sequence are, respectively: the nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 (or the reverse complement of the nth 10-mer in SEQ ID NO:11); the nth 10-mer in SEQ ID NO:12 and the nth 10-mer in SEQ ID NO:13 (or the reverse complement of the nth 10-mer in SEQ ID NO:13); the nth 10-mer in SEQ ID NO:14 and the nth 10-mer in SEQ ID NO:15 (or its reverse complement); or the nth 10-mer in SEQ ID NO:16 and the nth 10-mer in SEQ ID NO:17 (or its reverse complement).

[0478] The present application also provides the following:

[0479] 1. A method for sequencing target nucleic acids derived from multiple samples, the method comprising

[0480] (a) contacting a plurality of indexed polynucleotides with target nucleic acids derived from the multiple samples to generate a plurality of indexed-target polynucleotides, wherein

[0481] the indexed polynucleotides contacting the target nucleic acids from each sample comprise an index sequence or a combination of index sequences uniquely associated with this sample,

[0482] the index sequence or the combination of index sequences is selected from a set of index sequences, and

[0483] the Hamming distance between any two index sequences in the set of index sequences is not less than a first standard value, wherein the first standard value is at least 2;

[0484] (b) pooling the plurality of indexed-target polynucleotides;

[0485] (c) sequencing the pooled indexed-target polynucleotides to obtain a plurality of index reads of the index sequences and a plurality of target reads of the target sequences, each target read being associated with at least one index read; and

[0486] (d) using the index reads to determine the sample origin of the target reads.

[0487] 2. The method according to item 1, wherein the set of index sequences comprises multiple pairs of color-balanced index sequences, wherein any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine (A) base or a cytosine (C) base, and (ii) a guanine (G) base, a thymine (T) base or a uracil (U) base.

[0488] 3. The method according to any one of the preceding items, wherein the set of index sequences comprises at least 6 different index sequences.

[0489] 4. The method according to any one of the preceding items, wherein using the index reads to determine the sample origin of the target reads comprises:

[0490] for each index read, obtaining an alignment score with respect to the set of index sequences, each alignment score indicating the similarity between the sequence of the index read and the index sequences of the set of index sequences;

[0491] determining that a specific index read matches a specific index sequence based on the alignment score; and

[0492] determining that the target read associated with the specific index read is derived from the sample uniquely associated with the specific index sequence.

[0493] 5. The method according to any one of the preceding items, wherein the plurality of indexed polynucleotides comprise a plurality of indexing primers, the indexing primers comprising an indexing sequence from the set of indexing sequences.

[0494] 6. The method according to item 5, wherein each indexing primer further comprises a flow cell amplification primer binding sequence.

[0495] 7. The method according to item 6, wherein the flow cell amplification primer binding sequence comprises a P5 sequence or a P7′ sequence.

[0496] 8. The method according to item 5, wherein the target nucleic acids from the plurality of samples comprise nucleic acids having universal adapters covalently attached to one or both ends.

[0497] 9. The method according to item 8, wherein contacting the plurality of indexed polynucleotides with the target nucleic acids from the plurality of samples comprises:

[0498] hybridizing the plurality of indexing primers to the universal adapter covalently attached to one or both ends of the nucleic acid; and

[0499] extending the plurality of indexing primers to obtain a plurality of indexed - adapter - target polynucleotides.

[0500] 10. The method according to item 9, wherein the universal adapter and the target nucleic acid are double - stranded, and hybridizing the plurality of indexing primers to the universal adapter comprises hybridizing the plurality of indexing primers to only one strand of the universal adapter.

[0501] 11. The method according to item 9, wherein the universal adapter and the target nucleic acid are double - stranded, and hybridizing the plurality of indexing primers to the universal adapter comprises hybridizing the plurality of indexing primers to both strands of the universal adapter.

[0502] 12. The method according to item 11, wherein a first indexing primer hybridizing to the first strand of a specific universal adapter comprises a first indexing sequence selected from a first subset of the set of indexing sequences, and a second indexing primer hybridizing to the second strand of the specific universal adapter comprises a second indexing sequence selected from a second subset of the set of indexing sequences.

[0503] 13. The method according to item 12, wherein the first indexing sequence and the second indexing sequence are respectively:

[0504] the nth 10 - mer in SEQ ID NO:10 and the nth 10 - mer in SEQ ID NO:11 or the reverse complement of SEQ ID NO:11;

[0505] The nth decamer in SEQ ID NO:12 and the nth decamer in SEQ ID NO:13 or the reverse complement of SEQ ID NO:13;

[0506] The nth decamer in SEQ ID NO:14 and the nth decamer in SEQ ID NO:15 or the reverse complement of SEQ ID NO:15;

[0507] The nth decamer in SEQ ID NO:16 and the nth decamer in SEQ ID NO:17 or the reverse complement of SEQ ID NO:17; or

[0508] The nth decamer in SEQ ID NO:18 and the nth decamer in SEQ ID NO:19 or the reverse complement of SEQ ID NO:19.

[0509] 14. The method according to item 12, wherein the first subset comprises the index sequences listed in Table 1, and the second subset comprises the index sequences listed in Table 2.

[0510] 15. The method according to item 11, wherein the index primer hybridizing to both strands of the universal adaptor comprises an index sequence selected from the same subset of the set of index sequences.

[0511] 16. The method according to item 15, wherein the subset of index sequences is selected from one of the subsets of index sequences in Table 3.

[0512] 17. The method according to item 8, further comprising attaching the universal adaptor to one or both ends of the nucleic acid before step (a).

[0513] 18. The method according to item 17, wherein the attachment comprises attaching the universal adaptor by transposome-mediated fragmentation.

[0514] 19. The method according to item 18, wherein the transposome-mediated fragmentation comprises:

[0515] Providing nucleic acid molecules obtained from the plurality of samples and a plurality of transposome complexes, wherein each transposome complex comprises a transposase and two transposon end compositions, and the transposon end compositions comprise the sequence of the universal adaptor; and

[0516] Obtaining the target nucleic acid, wherein the target nucleic acid comprises at one or both ends the sequence of the universal adaptor transposed from the transposon end compositions.

[0517] 20. The method according to item 17, wherein said attachment comprises connecting said universal adaptor to one or both ends of said nucleic acid.

[0518] 21. The method according to item 20, wherein said ligation comprises enzymatic ligation or chemical ligation.

[0519] 22. The method according to item 21, wherein said chemical ligation comprises click chemical reaction ligation.

[0520] 23. The method according to item 17, wherein said attachment is carried out by amplification with a target-specific primer comprising a sequence of the universal adaptor.

[0521] 24. The method according to item 8, wherein said universal adaptor comprises a double-stranded adaptor.

[0522] 25. The method according to item 8, wherein said universal adaptor comprises a Y-shaped adaptor.

[0523] 26. The method according to item 8, wherein said universal adaptor comprises a single-stranded adaptor.

[0524] 27. The method according to item 8, wherein said universal adaptor comprises a hairpin adaptor.

[0525] 28. The method according to item 8, wherein each of said universal adaptors comprises an overhang at one end to be attached to the nucleic acid before being attached to the nucleic acid.

[0526] 29. The method according to item 8, wherein each of said universal adaptors comprises a blunt end to be attached to the nucleic acid before being attached to the nucleic acid.

[0527] 30. The method according to any one of the preceding items, wherein said plurality of indexed polynucleotides comprises a sample-specific adaptor, said sample-specific adaptor comprising an index sequence from said set of index sequences.

[0528] 31. The method according to item 30, wherein said sample-specific adaptor comprises an adaptor having two strands.

[0529] 32. The method according to item 31, wherein only one strand of said sample-specific adaptor comprises an index sequence.

[0530] 33. The method according to item 31, wherein each strand of said sample-specific adaptor comprises an index sequence.

[0531] 34. The method according to item 33, wherein the first strand of each sample-specific adaptor comprises a first index sequence selected from a first subset of the set of index sequences, and the second strand of the sample-specific adaptor comprises a second index sequence selected from a second subset of the set of index sequences.

[0532] 35. The method according to item 34, wherein the first index sequence and the second index sequence are respectively:

[0533] The nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 or the reverse complement of SEQ ID NO:11;

[0534] The nth 10-mer in SEQ ID NO:12 and the nth 10-mer in SEQ ID NO:13 or the reverse complement of SEQ ID NO:13;

[0535] The nth 10-mer in SEQ ID NO:14 and the nth 10-mer in SEQ ID NO:15 or the reverse complement of SEQ ID NO:15;

[0536] The nth 10-mer in SEQ ID NO:16 and the nth 10-mer in SEQ ID NO:17 or the reverse complement of SEQ ID NO:17; or

[0537] The nth 10-mer in SEQ ID NO:18 and the nth 10-mer in SEQ ID NO:19 or the reverse complement of SEQ ID NO:19.

[0538] 36. The method according to item 34, wherein the first subset comprises the index sequences listed in Table 1, and the second subset comprises the index sequences listed in Table 2.

[0539] 37. The method according to item 34, wherein the first subset and the second subset are the same.

[0540] 38. The method according to item 37, wherein the subset of index sequences is selected from one of the subsets of index sequences in Table 3.

[0541] 39. The method according to item 30, wherein each sample-specific adaptor comprises a flow cell amplification primer binding sequence.

[0542] 40. The method according to item 39, wherein the flow cell amplification primer binding sequence comprises a P5 sequence, a P5′ sequence, a P7 sequence or a P7′ sequence.

[0543] 41. The method according to item 30, wherein contacting the plurality of indexed polynucleotides with the target nucleic acid comprises attaching a sample-specific adaptor to the target nucleic acid by transposome-mediated fragmentation.

[0544] 42. The method according to item 41, wherein the transposome-mediated fragmentation comprises:

[0545] providing nucleic acid molecules obtained from the plurality of samples;

[0546] providing a plurality of transposome complexes, wherein each transposome complex comprises a transposase and two transposon end compositions, the transposon end compositions comprising the sequence of the sample-specific adaptor; and

[0547] obtaining the target nucleic acid, wherein the target nucleic acid comprises at one or both ends the sequence of the sample-specific adaptor transposed from the transposon end compositions.

[0548] 43. The method according to item 30, wherein contacting the plurality of indexed polynucleotides with the target nucleic acid comprises ligating the sample-specific adaptor to the target nucleic acid.

[0549] 44. The method according to item 43, wherein the ligation comprises enzymatic ligation or chemical ligation.

[0550] 45. The method according to item 44, wherein the chemical ligation comprises click chemical reaction ligation.

[0551] 46. The method according to item 30, wherein the sample-specific adaptor comprises a Y-shaped adaptor having complementary double-stranded regions and mismatched single-stranded regions.

[0552] 47. The method according to item 46, wherein each strand of the sample-specific adaptor comprises an index sequence at the mismatched single-stranded region.

[0553] 48. The method according to item 46, wherein only one strand of the sample-specific adaptor comprises an index sequence at the mismatched single-stranded region.

[0554] 49. The method according to item 30, wherein the sample-specific adaptor comprises a single-stranded adaptor.

[0555] 50. The method according to item 30, wherein the sample-specific adaptor comprises a hairpin adaptor.

[0556] 51. The method according to any one of the preceding items, wherein contacting the plurality of indexed polynucleotides with the target nucleic acid comprises attaching the plurality of indexed polynucleotides to both ends of the target nucleic acid.

[0557] 52. The method according to any one of items 1-50, wherein contacting the plurality of indexed polynucleotides with the target nucleic acid comprises attaching the plurality of indexed polynucleotides to only one end of the target nucleic acid.

[0558] 53. The method according to any one of the preceding items, wherein the combination of index sequences is an ordered combination of index sequences.

[0559] 54. The method according to any one of the preceding items, the method further comprising amplifying the pooled indexed-target polynucleotides before sequencing the pooled indexed-target polynucleotides.

[0560] 55. The method according to any one of the preceding items, the method further comprising fragmenting nucleic acid molecules obtained from the plurality of samples to obtain the target nucleic acid before step (a).

[0561] 56. The method according to item 55, wherein the fragmentation comprises transposase-mediated fragmentation.

[0562] 57. The method according to item 56, wherein the transposase-mediated fragmentation comprises:

[0563] providing the nucleic acid molecule and a plurality of transpososome complexes, wherein each transpososome complex comprises a transposase and two transposon end compositions; and

[0564] obtaining the target nucleic acid, the target nucleic acid comprising sequences transposed from the transposon end compositions at one or both ends.

[0565] 58. The method according to item 55, wherein the fragmentation comprises contacting with a plurality of PCR primers targeting a sequence of interest to obtain the target nucleic acid comprising the sequence of interest.

[0566] 59. The method according to any one of the preceding items, wherein the set of index sequences comprises a plurality of non-overlapping subsets of index sequences, and the Hamming distance between any two index sequences in any subset is not less than a second standard value.

[0567] 60. The method according to item 59, wherein the second standard value is greater than the first standard value.

[0568] 61. The method according to item 60, wherein the first standard value is 4, and the second standard value is 5.

[0569] 62. The method according to any one of items 1-60, wherein the first standard value is 3.

[0570] 63. The method according to any one of items 1-60, wherein the first standard value is 4.

[0571] 64. The method according to any one of items 1-61, wherein the edit distance between any two index sequences in the index sequence set is not less than a third standard value.

[0572] 65. The method according to item 64, wherein the edit distance is a modified Levenshtein distance, and end gaps are not assigned penalty scores.

[0573] 66. The method according to item 65, wherein the third standard value is 3.

[0574] 67. The method according to item 65, wherein:

[0575] each index sequence in the index sequence set has 10 bases;

[0576] the first standard value is 4; and

[0577] the third standard value is 3.

[0578] 68. The method according to any one of the foregoing items, wherein the index sequence set comprises 10-mers in SEQ ID NO:9.

[0579] 69. The method according to any one of the foregoing items, wherein each index sequence has 32 bases or fewer.

[0580] 70. The method according to item 69, wherein each index sequence has 10 bases or fewer.

[0581] 71. The method according to item 69, wherein each index sequence has 6 to 8 bases.

[0582] 72. The method according to any one of the foregoing items, wherein the index sequence set does not include index sequences that are empirically determined to have poor performance in indexing the source of nucleic acid samples in massively parallel sequencing.

[0583] 73. The method according to item 72, wherein the index sequences not included comprise the sequences in Table 4.

[0584] 74. The method according to any one of the foregoing items, wherein the index sequence set does not include any subsequences of the sequences of adapters or primers in the sequencing platform, or the reverse complements of such subsequences.

[0585] 75. The method according to item 74, wherein the sequence of the adaptor or primer in the sequencing platform comprises SEQ ID NO: 1 (AGATGTGTATAAGAGACAG), SEQ ID NO: 3 (TCGTCGGCAGCGTC), SEQ ID NO: 5 (CCGAGCCCACGAGAC), SEQ ID NO: 7 (CAAGCAGAAGACGGCATACGAGAT), and SEQ ID NO: 8 (AATGATACGGCGACCACCGAGATCTACAC).

[0586] 76. The method according to any one of the preceding items, wherein each of the index sequences in the index sequence set has a guanine / cytosine (GC) content between 25% and 75%.

[0587] 77. The method according to any one of the preceding items, wherein the index sequence set comprises at least 12 different index sequences.

[0588] 78. The method according to any one of the preceding items, wherein the index sequence set comprises at least 20 different index sequences.

[0589] 79. The method according to any one of the preceding items, wherein the index sequence set comprises at least 24 different index sequences.

[0590] 80. The method according to any one of the preceding items, wherein the index sequence set comprises at least 28 different index sequences.

[0591] 81. The method according to any one of the preceding items, wherein the index sequence set comprises at least 48 different index sequences.

[0592] 82. The method according to any one of the preceding items, wherein the index sequence set comprises at least 80 different index sequences.

[0593] 83. The method according to any one of the preceding items, wherein the index sequence set comprises at least 96 different index sequences.

[0594] 84. The method according to any one of the preceding items, wherein the index sequence set comprises at least 112 different index sequences.

[0595] 85. The method according to any one of the preceding items, wherein the index sequence set does not include any homopolymer having four or more consecutive identical bases.

[0596] 86. The method according to any one of the preceding items, wherein the index sequence set does not include an index sequence that matches or is reverse complementary to one or more sequencing primer sequences.

[0597] 87. The method according to item 86, wherein the sequencing primer sequence is included in the sequence of the plurality of indexed polynucleotides.

[0598] 88. The method according to any one of the preceding items, wherein the set of index sequences does not include an index sequence that matches or is reverse complementary to one or more flow cell amplification primer sequences.

[0599] 89. The method according to item 88, wherein the flow cell amplification primer sequence is included in the sequence of the plurality of indexed polynucleotides.

[0600] 90. The method according to any one of the preceding items, wherein the set of index sequences comprises index sequences having the same number of bases.

[0601] 91. The method according to any one of the preceding items, wherein each index sequence in the set of index sequences has a combined number of guanine and cytosine bases that is not less than 2 and not greater than 6.

[0602] 92. The method according to any one of the preceding items, wherein the plurality of indexed polynucleotides comprises DNA or RNA.

[0603] 93. A method for sequencing target nucleic acids derived from a plurality of samples, the method comprising:

[0604] (a) providing a plurality of double-stranded nucleic acid molecules derived from the plurality of samples;

[0605] (b) providing a plurality of transpososome complexes, wherein each transpososome complex comprises a transposase and two transposon end compositions;

[0606] (c) incubating the double-stranded nucleic acid molecules with the transpososome complexes to obtain double-stranded nucleic acid fragments, wherein the double-stranded nucleic acid fragments comprise sequences transposed from the transposon end compositions at one or both ends;

[0607] (d) contacting a plurality of indexing primers with the double-stranded nucleic acid fragments to generate a plurality of indexed-fragment polynucleotides, wherein

[0608] the indexing primers contacting the double-stranded nucleic acid fragments derived from each sample comprise an index sequence or a combination of index sequences uniquely associated with this sample, and

[0609] the index sequence or the combination of index sequences is selected from a set of index sequences;

[0610] (e) pooling the plurality of indexed-fragment polynucleotides;

[0611] (f) Sequencing the pooled index-fragment polynucleotides to obtain index reads of the index sequences and multiple target reads of the target sequences, each target read being associated with at least one index read; and

[0612] (g) Using the index reads to determine the sample origin of the target reads.

[0613] 94. The method according to item 93, wherein the Hamming distance between any two index sequences in the index sequence set is not less than a first standard value, and the first standard value is at least 2.

[0614] 95. The method according to any one of items 93-94, wherein the index sequence set contains multiple pairs of color-balanced index sequences, and any two bases at the corresponding sequence positions of each pair of color-balanced index sequences include both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base.

[0615] 96. The method according to any one of items 93-95, wherein contacting the multiple index primers with the double-stranded nucleic acid fragment includes:

[0616] Hybridizing the multiple index primers with the sequences transposed from the transposon end composition at one or both ends of the double-stranded nucleic acid fragment; and

[0617] Extending the multiple index primers to obtain the index-fragment polynucleotides.

[0618] 97. The method according to item 96, wherein the hybridization includes hybridizing the multiple index primers with only one strand of the double-stranded nucleic acid fragment.

[0619] 98. The method according to item 96, wherein the hybridization includes hybridizing the multiple index primers with both strands of the double-stranded nucleic acid fragment.

[0620] 99. The method according to item 98, wherein the first index primer hybridizing with the first strand of a specific double-stranded nucleic acid fragment contains a first index sequence selected from a first subset of the index sequence set, and the second index primer hybridizing with the second strand of the specific double-stranded nucleic acid fragment contains a second index sequence selected from a second subset of the index sequence set.

[0621] 100. The method according to item 99, wherein the first index sequence and the second index sequence are respectively:

[0622] The nth 10-mer in SEQ ID NO:10 and the nth 10-mer in SEQ ID NO:11 or the reverse complement of SEQ ID NO:11;

[0623] The nth decamer in SEQ ID NO:12 and the nth decamer in SEQ ID NO:13 or the reverse complement of SEQ ID NO:13;

[0624] The nth decamer in SEQ ID NO:14 and the nth decamer in SEQ ID NO:15 or the reverse complement of SEQ ID NO:15;

[0625] The nth decamer in SEQ ID NO:16 and the nth decamer in SEQ ID NO:17 or the reverse complement of SEQ ID NO:17; or

[0626] The nth decamer in SEQ ID NO:18 and the nth decamer in SEQ ID NO:19 or the reverse complement of SEQ ID NO:19.

[0627] 101. The method according to item 99, wherein the first subset comprises the index sequences listed in Table 1, and the second subset comprises the index sequences listed in Table 2.

[0628] 102. The method according to item 99, wherein the first subset and the second subset are the same.

[0629] 103. The method according to item 102, wherein the first subset or the second subset of the index sequences is selected from one of the subsets of the index sequences in Table 3.

[0630] 104. The method according to any one of items 93 - 103, wherein the set of index sequences comprises at least 6 different index sequences.

[0631] 105. The method according to any one of items 93 - 104, wherein each index primer comprises a flow cell amplification primer binding sequence.

[0632] 106. The method according to item 105, wherein the flow cell amplification primer binding sequence comprises a P5 sequence or a P7′ sequence.

[0633] 107. The method according to any one of items 93 - 106, wherein at least one of the transposome complexes comprises Tn5 transposase and a Tn5 transposon end composition.

[0634] 108. The method according to any one of items 93 - 107, wherein at least one of the transposome complexes comprises Mu transposase and a Mu transposon end composition.

[0635] 109. A computer program product comprising a non-transitory machine-readable medium storing program code that, when executed by one or more processors of a computer system, causes the computer system to implement a method for sequencing target nucleic acids derived from a plurality of samples, the program code comprising:

[0636] (a) code for receiving a plurality of indexed reads and a plurality of target reads of target sequences obtained from target nucleic acids derived from the plurality of samples, wherein

[0637] each target read comprises a target sequence obtained from a target nucleic acid of a sample among the plurality of samples,

[0638] each indexed read comprises an index sequence obtained from a target nucleic acid of a sample among the plurality of samples, the index sequence being selected from a set of index sequences,

[0639] each target read is associated with at least one indexed read,

[0640] each sample among the plurality of samples is uniquely associated with one or more index sequences from the set of index sequences, and

[0641] the Hamming distance between any two index sequences from the set of index sequences is not less than a first standard value, wherein the first standard value is at least 2;

[0642] (b) code for identifying, among the plurality of target reads, a subset of target reads associated with indexed reads that match at least one index sequence uniquely associated with a particular sample among the plurality of samples; and

[0643] (c) code for determining the target sequence of the particular sample based on the identified subset of target reads.

[0644] 110. The computer program product according to item 109, wherein the set of index sequences comprises a plurality of pairs of color-balanced index sequences, wherein any two bases at corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base.

[0645] 111. A computer system comprising:

[0646] one or more processors;

[0647] system memory; and

[0648] One or more computer-readable storage media having computer-executable instructions stored thereon that cause a computer system to implement a method for sequencing nucleic acids in a plurality of samples, the instructions including:

[0649] (a) receiving a plurality of indexed reads and a plurality of target reads of target sequences obtained from target nucleic acids derived from the plurality of samples, wherein

[0650] each target read comprises a target sequence obtained from a target nucleic acid derived from a sample of the plurality of samples,

[0651] each indexed read comprises an index sequence obtained from a target nucleic acid derived from a sample of the plurality of samples, the index sequence being selected from a set of index sequences,

[0652] each target read is associated with at least one indexed read,

[0653] each sample of the plurality of samples is uniquely associated with one or more index sequences of the set of index sequences, and

[0654] the Hamming distance between any two index sequences of the set of index sequences is not less than a first standard value, wherein the first standard value is at least 2;

[0655] (b) identifying a subset of target reads in the plurality of target reads that are associated with indexed reads that match at least one index sequence that is uniquely associated with a particular sample of the plurality of samples; and

[0656] (c) determining the target sequence of the particular sample based on the identified subset of target reads.

[0657] 112. The computer system of item 111, wherein the set of index sequences comprises a plurality of pairs of color-balanced index sequences, wherein any two bases at corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base, or a uracil base.

Claims

1. A method for processing target nucleic acids from multiple samples, the method comprises: contacting a plurality of indexed polynucleotides with the target nucleic acids from the multiple samples to generate a plurality of indexed-target polynucleotides, wherein: the indexed polynucleotides contacting the target nucleic acids from each sample comprise an index sequence or a combination of index sequences uniquely associated with this sample, the index sequence or the combination of index sequences is selected from an index sequence set, and the index sequence set comprises multiple pairs of color-balanced index sequences, wherein any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both of the following: (i) an adenine (A) base or a cytosine (C) base, and (ii) a guanine (G) base, a thymine (T) base or a uracil (U) base.

2. The method according to claim 1, further comprises: pooling the plurality of indexed-target polynucleotides.

3. The method according to claim 2, further comprises: sequencing the pooled indexed-target polynucleotides to obtain a plurality of index reads of the index sequences and a plurality of target reads of the target sequences, each target read being associated with at least one index read.

4. The method according to claim 3, further comprises: using the index reads to determine the sample origin of the target reads.

5. The method according to claim 4, wherein using the index reads to determine the sample origin of the target reads comprises: for each index read, obtaining an alignment score with respect to the index sequence set, each alignment score indicating the similarity between the sequence of the index read and the index sequences of the index sequence set; determining that a specific index read matches a specific index sequence based on the alignment score; and determining that the target read associated with the specific index read is derived from the sample uniquely associated with the specific index sequence.

6. The method according to claim 1, wherein the index sequence set comprises at least 6 different index sequences.

7. The method according to claim 1, wherein the plurality of indexed polynucleotides comprise a plurality of index primers, and the index primers comprise the index sequences in the index sequence set.

8. The method according to claim 7, wherein each index primer further comprises a flow cell amplification primer binding sequence.

9. The method according to claim 7, wherein the target nucleic acids from the multiple samples comprise nucleic acids having universal adapters covalently attached to one or both ends.

10. The method according to claim 9, wherein the universal adapter comprises a double-stranded adapter.

11. The method according to claim 9, wherein the universal adapter comprises a Y-shaped adapter.

12. The method according to claim 9, wherein the universal adapter comprises a single-stranded adapter.

13. The method according to claim 9, wherein the universal adapter comprises a hairpin adapter.

14. A method for processing multiple nucleic acid samples, the method comprises: applying different oligonucleotides of an oligonucleotide set to the target nucleic acids of each nucleic acid sample of the multiple nucleic acid samples, thereby generating a plurality of indexed-target polynucleotides, wherein The oligonucleotide set is configured to identify the source of a nucleic acid sample in multiplexed massively parallel sequencing; Each oligonucleotide of the oligonucleotide set comprises an index sequence; The oligonucleotide set comprises an index sequence set, the index sequence set comprising at least 6 different index sequences; and The index sequence set comprises multiple pairs of color-balanced index sequences, wherein any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine (A) base or a cytosine (C) base, and (ii) a guanine (G) base, a thymine (T) base or a uracil (U) base; Pool the multiple index-target polynucleotides; Sequence the multiple index-target polynucleotides; And Use the index sequence set to determine the source of the nucleic acid samples from which the multiple index-target polynucleotides are derived.

15. The method according to claim 14, wherein the Hamming distance between any two index sequences in the index sequence set is not less than a first standard value, wherein the first standard value is at least 2.

16. A method for processing target nucleic acids derived from multiple samples, the method comprising: (a) providing multiple double-stranded nucleic acid molecules derived from the multiple samples; (b) providing multiple transpososome complexes, wherein each transpososome complex comprises a transposase and two transposon end compositions; (c) incubating the double-stranded nucleic acid molecules with the transpososome complexes to obtain double-stranded nucleic acid fragments, wherein the double-stranded nucleic acid fragments comprise sequences transposed from the transposon end compositions at one or both ends; (d) contacting multiple index primers with the double-stranded nucleic acid fragments to generate multiple index-fragment polynucleotides, wherein: The index primers contacting the double-stranded nucleic acid fragments derived from each sample comprise an index sequence or a combination of index sequences uniquely associated with this sample, The index sequence or the combination of index sequences is selected from the index sequence set; and The index sequence set comprises multiple pairs of color-balanced index sequences, wherein any two bases at the corresponding sequence positions of each pair of color-balanced index sequences comprise both: (i) an adenine base or a cytosine base, and (ii) a guanine base, a thymine base or a uracil base.

17. The method according to claim 16, further comprising: (e) pooling the multiple index-fragment polynucleotides.

18. The method according to claim 17, further comprising: (f) sequencing the pooled index-fragment polynucleotides to obtain index reads of the index sequences and multiple target reads of the target sequences, each target read being associated with at least one index read.

19. The method according to claim 18, further comprising: (g) using the index reads to determine the sample source of the target reads.

20. The method according to claim 16, wherein the Hamming distance between any two index sequences in the index sequence set is not less than a first standard value, wherein the first standard value is at least 2.

Citation Information

Patent Citations

  • Polymerase enzymes and reagents for enhanced nucleic acid sequencing

    US20080108082A1

  • Detecting and classifying copy number variation

    US20130029852A1

  • Error suppression in sequenced DNA fragments using redundant reads with unique molecular indices (UMIS)

    US20160319345A1

  • Methods and systems for generation and error-correction of unique molecular index sets with heterogeneous molecular lengths

    US20180201992A1

  • Process for amplifying, detecting, and / or-cloning nucleic acid sequences

    US4683195A