Comparison of polynucleotide copies with different characteristics

JP7902116B2Active Publication Date: 2026-08-07ILLUMINA INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ILLUMINA INC
Filing Date
2021-05-24
Publication Date
2026-08-07

Smart Images

  • Figure 0007902116000002
    Figure 0007902116000002
  • Figure 0007902116000003
    Figure 0007902116000003
  • Figure 0007902116000004
    Figure 0007902116000004
Patent Text Reader

Abstract

Methods are provided that include generating copies of two or more populations of polynucleotides comprising a discriminating sequence, the copies being bound to a substrate; hybridizing an oligonucleotide to the discriminating sequence; and comparing the amount of oligonucleotide hybridized to the copies of the two or more populations of polynucleotides, wherein at least one characteristic differs between the two or more populations of polynucleotides or between the copies of the two or more populations of substrate-bound polynucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 031,230, filed May 28, 2020, the entire contents of which are incorporated herein by reference.

[0002] Sequence Listing This application includes a sequence listing created on May 18, 2021. This file in ASCII format is named H2055903.txt and is 1 KB in size. This file is incorporated herein by reference in its entirety.

Background Art

[0003] Many current sequencing platforms use "sequencing by synthesis" (SBS) technology and methods using fluorescence for detection. In some examples, multiple polynucleotides isolated from one or more populations of nucleotides to be sequenced are bound to the surface of a substrate and replicated. Then, SBS can be performed on the copies bound to the surface. By creating copies of the polynucleotide or amplifying the polynucleotide and sequencing its copies, the fluorescence signal emitted during sequencing increases, thereby improving the sequencing process.

[0004] Copies of a substrate-bound polynucleotide can be synthesized by solid-phase nucleic acid amplification, a method that allows the amplified product to be immobilized on a solid support to form an array containing clusters of immobilized nucleic acid molecules. Each cluster or colony on such an array is a plurality of copies of the target polynucleotide chain and a plurality of immobilized polynucleotide chains complementary to it. Cluster amplification or cluster formation is an example of a method by which surface-bound copies of the target polynucleotide and complements of the target polynucleotide are synthesized for SBS. Some examples of similarly suitable methods that can be used to generate surface-bound copies, etc., include bridge amplification, kinetic exclusion amplification ("ExAmp"), or others.

[0005] Cluster formation involves the use of polymerase to synthesize clusters bound to a surface. However, a known problem associated with certain polymerases and polymerization methods is quantitative synthesis bias related to the various characteristics of the target polynucleotide. For example, in some cases, cluster formation methods may be biased to amplify more copies of target polynucleotides with a lower percentage of guanine (G)-cytosine (C) base pairs compared to polynucleotides with a relatively higher GC content. In other cases, cluster formation methods may be biased to amplify more copies of relatively shorter target polynucleotides compared to relatively longer polynucleotides. In several other cases, other theoretical causes of bias, such as the method of polynucleotide sample preparation or other differences, may affect the relative amplification levels of the polynucleotides. [Overview of the Initiative]

[0006] Therefore, taking into consideration at least the foregoing, sequencing techniques would benefit from methods for determining the presence of such biases in cluster formation and other amplification processes, as well as methods for identifying, isolating, and modifying such techniques that can minimize such biases and yield more accurate sequencing results.

[0007] In one embodiment, a method is provided for producing copies of two or more populations of polynucleotides containing an identification sequence, wherein the copies are bound to a substrate, hybridizing an oligonucleotide to the identification sequence, and comparing the amount of oligonucleotide hybridized to the copies of the two or more populations of polynucleotides, wherein at least one characteristic differs between the two or more populations of polynucleotides or between the production of copies of two or more populations of polynucleotides bound to a substrate.

[0008] In one example, at least one feature is selected from length, guanine-cytosine content, and preparation method. In another example, at least one feature includes guanine-cytosine content. In yet another example, at least one feature includes length. In yet another example, at least one feature includes preparation method. In a further example, at least one feature differs between the preparation of copies of two or more populations of substrate-bound polynucleotides. In yet another example, the oligonucleotide includes a fluorophore.

[0009] In one example, the method further includes detecting differences between the amounts of oligonucleotides hybridized to copies of two or more populations of a substrate-bound polynucleotide, where the difference is at least about 10%. In another example, the difference is at least about 20%. In yet another example, the difference is at least about 30%.

[0010] In one example, at least one feature includes a combination, the combination includes two or more of guanine-cytosine content, length, preparation method, and the creation of copies of two or more populations of polynucleotides bound to a substrate, the two or more populations of polynucleotides include three or more populations of polynucleotides, and each combination of the three or more populations of polynucleotides is different from another combination of polynucleotides.

[0011] Another example further involves detecting a difference between the amounts of oligonucleotides hybridized to two or more copies of three or more populations of a substrate-bound polynucleotide, where the difference is at least about 10%. In one example, the difference is at least about 20%. In another example, the difference is at least about 30%.

[0012] In another embodiment, a method is provided for producing copies of two or more populations of polynucleotides containing a recognition sequence, wherein the copies are bound to a substrate, hybridizing an oligonucleotide containing a fluorophore to the recognition sequence, and detecting the amount of oligonucleotide hybridized to the copies of the two or more populations of polynucleotides, wherein at least one feature differs between the two or more populations of polynucleotides or between the production of copies of two or more populations of polynucleotides bound to a substrate, the at least one feature being selected from length, guanine-cytosine content, preparation method, and production of copies of two or more populations of polynucleotides bound to a substrate.

[0013] In one example, at least one feature includes guanine-cytosine content. In another example, at least one feature includes length. In yet another example, at least one feature includes preparation method. In yet another example, at least one feature differs between the preparation of copies of two or more populations of substrate-bound polynucleotides.

[0014] Another example further involves detecting a difference between the amounts of oligonucleotides hybridized to copies of two or more populations of polynucleotides, where the difference is at least about 10%. In one example, the difference is at least about 20%. In yet another example, the difference is at least about 30%. [Brief explanation of the drawing]

[0015] These and other features, aspects, and advantages of this disclosure will be better understood when the modes for carrying out the invention described below are read with reference to the accompanying drawings. [Figure l] This is a flowchart illustrating one embodiment of the method disclosed herein. [Figure 2] This is an illustrative diagram of elements of one embodiment of the method according to the embodiments of the present disclosure. [Figure 3] This graph shows the difference in average intensity detected from fluorescently labeled oligonucleotides hybridized to polynucleotide copies from loaded DNA in different ratios, according to an embodiment of the present disclosure. [Figure 4] This graph compares the fluorescence detection intensity after cluster formation of polynucleotides loaded with 40%, 50%, or 60% of the total DNA amount in one example of the cluster formation procedure. [Figure 5] This is a flowchart illustrating one embodiment of the method disclosed herein. [Modes for carrying out the invention]

[0016] This disclosure relates to a method for evaluating bias in the replication of polynucleotides, such as as part of the SBS process. In particular, it includes a process for identifying the presence of a bias toward producing more or fewer copies of a given population of polynucleotides compared to copies of different populations. Polynucleotides from different populations may be distinguished from one another by different characteristics. These characteristics may be any characteristics of the polynucleotides of a population, including the physical properties of the polynucleotide chain or the process by which the population of polynucleotides is supplied as a sample preparation.

[0017] For example, polynucleotides from one population may have a lower or higher ratio of C and / or G bases to A and / or T bases compared to polynucleotides from another population. In another example, polynucleotides from one population may have a certain number of nucleotide lengths, and the length of polynucleotides from one population may differ from the length of polynucleotides from another population. In yet another example, different populations of polynucleotides may be subjected to different preparation methods. For example, they may be subjected to different methods of fragmenting target molecules into shorter polynucleotides for replication and sequencing, or different methods of adding oligonucleotide sequences or identifiers to polynucleotides (methods of tagging or indexing polynucleotides for identification of copies subsequently made, processes sometimes referred to as indexing, indexing, or barcoding), or different methods of separating polynucleotide sequences from an initial sample (e.g., isolating and selecting polynucleotides of a certain size or within a certain size range). In yet another example, a method of clustering from one population of polynucleotides may differ from a method of clustering from different populations of polynucleotides.

[0018] In some examples, any feature can distinguish two or more populations, whether it directly relates to the physical properties of the polynucleotides of different populations, or indirectly exhibits their properties as a result of the preparation, storage, processing, handling, or preparation or clustering processes of the polynucleotides, or relates to other properties, such as other components that may be present with the polynucleotides. The methods disclosed herein may be used to determine whether differences in features result in a bias that leads to disproportionately large or faster replication of polynucleotides from one population compared to another population in a replication process, such as a clustering process.

[0019] In some examples, populations may differ with respect to two or more characteristics (including GC content, length, preparation method, or the process of making copies, such as during cluster formation). For example, populations may differ with respect to length (e.g., the number of nucleotides in the polynucleotide of the population) and GC content (e.g., the relative amount of G residues and / or C residues in the polynucleotide population compared to A residues and / or T residues in the polynucleotide population). Or, they may differ with respect to these and any of the preparation methods, any of the replication methods during cluster formation, or any combination of two or more of the aforementioned. In some examples, populations may differ with respect to one or more characteristics, or any combination of any two or more characteristics, or any combination of any three or more characteristics (e.g., length, GC content, or the method by which the polynucleotides are prepared for replication or cluster formation, and / or replication methods, e.g., cluster formation methods).

[0020] Differences in the preparation methods of different polynucleotide populations, which characterize the population, may confer different structural characteristics to the population, such as differences in the efficiency of obtaining polynucleotides of the intended size, consistency of polynucleotide size within the population, and how many polynucleotides within the population correctly bound adapters or other sequences. All of these can lead to bias or replication differences that provide evidence following cluster formation. The methods disclosed herein may be used to confirm such effects of differences in preparation methods.

[0021] In some examples, one or more features of one or more of the two or more populations of polynucleotides can be preselected, including any of the foregoing features or any two or more of the foregoing features combined with each other. For example, other aspects of the clustering process or the replication process may cause a bias, increase, decrease, remove, or otherwise affect a bias with respect to polynucleotide length, GC content, preparation process, or other features, or the method of making copies, or any combination of two or more thereof. Thus, the characteristics of the polynucleotide population can be preselected, can be configured to indicate such a possibility, or can be assumed to be the cause or source of the bias, the clustering or other replication process can be performed, and the amount of copies of two or more populations of polynucleotides can be compared. A greater amount of copies of one population compared to another population, normalized by the respective starting amount at the start of replication, can indicate a bias in the direction of the polynucleotide having the preselected feature under the replication conditions used or a bias in the direction opposite to the polynucleotide.

[0022] An example of such a method is shown in the flowchart of Figure 1. Two or more populations of polynucleotides are prepared for replication, for example, by a clustering process. The preparation process involves adding an oligonucleotide sequence to a polynucleotide of one population, and adding another oligonucleotide sequence to a polynucleotide of another population. In examples where three or more populations of polynucleotides are used, an oligonucleotide sequence different from the oligonucleotide sequence added to each polynucleotide of another population may be added to a polynucleotide of one population, such that the polynucleotide of each population contains an oligonucleotide sequence that is specific to the polynucleotide of that population and is different from the oligonucleotides added to any other polynucleotide of any other population. The sequence of such oligonucleotide, called a discriminant sequence, may be different so that the discriminant sequence hybridizes to an oligonucleotide that has sequence complementarity to the discriminant sequence. For example, each discriminant sequence added to polynucleotides of two or more populations may hybridize to an oligonucleotide sequence that is not hybridizable to the discriminant sequence of any other polynucleotide of those two or more populations. As will be further explained below, the presence of sequence identifiers that are distinguishable between polynucleotides of different populations may, according to the methods disclosed herein, enable the identification of a copy of a polynucleotide of a given population rather than a copy of any other population.

[0023] Next, the single-stranded polynucleotides of two or more populations can be replicated along with the copies bound to the substrate. For example, replication can be carried out as described above by a cluster formation process by solid-state exclusion amplification, a cluster formation by bridge amplification, or other processes. In a non-limiting example, the 3' end of the polynucleotide can hybridize to a primer sequence bound to the substrate, and a polymerization process can be carried out starting with a primer bound to the surface to create a complement to the polynucleotide and extending the complement to the 5' end of each polynucleotide. Next, the polynucleotides of two or more populations can be amplified from their complements bound to the surface. According to a non-limiting example of the bridge PCR process, the free 3' end of the complement bound to the surface to the polynucleotides of two or more populations can then hybridize to another primer sequence bound to the substrate. The complement can then be replicated by a polymerase reaction, thereby resulting in copies of the polynucleotides of two or more populations of polynucleotides extending from the surface and their complements. Next, the complements and copies bound to the surface can be denatured from each other, and a further polymerization reaction can be carried out (after hybridization of their free 3' ends to a primer bound to the surface as the starting site of the polymerase reaction) in which the copies of the polynucleotides of two or more populations of polynucleotides bound to the surface and their complements are replicated, and then the complementary pairs of polynucleotides bound to the surface can be denatured from each other. By repeating this process, clusters of copies of the polynucleotides of two or more populations and their complements bound to the substrate can be formed. Other similar methods for making copies of a population of polynucleotides, regardless of PCR, rolling circle amplification, multiple displacement amplification, random primer amplification, isothermal amplification, etc., can also be used in other examples.

[0024] Next, the amount of a substrate-bound copy of one polynucleotide from two or more populations of polynucleotides can be determined. For example, an oligonucleotide capable of hybridizing to a distinctive sequence of one polynucleotide from two or more populations can be added to hybridize to the distinctive sequence present in the polynucleotide. The hybridizable oligonucleotide may include a detectable marker, such as a fluorescent marker that can emit detectable fluorescence upon stimulation with electromagnetic radiation of a given wavelength. By generating fluorescence in such a hybridized oligonucleotide and detecting the amount of emitted fluorescence, the amount of a copy of one polynucleotide from two or more populations can be evaluated.

[0025] The oligonucleotide can then be dehybridized and subsequently incubated with another oligonucleotide, the other oligonucleotide being hybridizable to the distinctive sequence of a polynucleotide from another population of two or more populations of polynucleotides. The other hybridizable oligonucleotide may include a detectable marker, such as a fluorescent marker that can emit detectable fluorescence upon stimulation with electromagnetic radiation of a given wavelength. By generating fluorescence in such other hybridized oligonucleotides and detecting the amount of emitted fluorescence, the amount of copies of the polynucleotide from another population of two or more populations can be assessed. In cases where potential replication bias is caused by or arises from the characteristics or combinations of characteristics of three or more populations of polynucleotides, the process of hybridizing a hybridizable oligonucleotide to each distinctive sequence of a polynucleotide from an individual population of nucleotides, measuring the amount of the hybridized oligonucleotide, and then dehybridizing them (if followed by hybridization with another oligonucleotide) can be repeated to obtain a measure of the amount of each type of oligonucleotide as a measure of the amount of copies of each polynucleotide from two or more populations of polynucleotides.

[0026] For example, differences between the amounts of oligonucleotides hybridized to each copy of two or more populations of polynucleotides can be detected, for example, by comparing the relative amounts of fluorescence emitted from oligonucleotides that can hybridize to each identification sequence. For example, a sample may contain two populations of polynucleotides characterized by containing different identification sequences and having different features. Different samples may contain polynucleotides from each of the two populations in different relative proportions. For example, one population may constitute about 0%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% of the total nucleotide content of the sample, and the other population constitutes the remainder of the sample. Copies and complements of the populations can then be prepared, for example, in a clustering process, according to the disclosure herein.

[0027] Next, oligonucleotides can hybridize to the identification sequences of copies of a population of polynucleotides. The amount of hybridized oligonucleotides can be measured, for example, if the oligonucleotides contain fluorophores, and fluorescence emission can be detected and quantified as a measure of the total amount of oligonucleotides hybridized to the identification sequences of a given population. In this way, the amount of oligonucleotides hybridized to each population can be measured and compared to obtain an indicator of the relative abundance of polynucleotide copies in each population after replication. For example, if a sample contains polynucleotides from each population in different relative proportions, a difference may be detectable. For instance, if, before replication by cluster formation, one population constitutes approximately 0%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, or 45% of the nucleotide content of the sample, and the other population constitutes the remainder, a difference may be detectable.

[0028] In one example, fluorescence emission is measured from oligonucleotides containing fluorophores hybridized to the identification sequences of each population, and differences in fluorescence are confirmed. For example, oligonucleotides capable of hybridizing to the identification sequences of one population of a polynucleotide may contain different fluorophores that are detectable from oligonucleotides capable of hybridizing to the identification sequences of the other population of the oligonucleotide (e.g., Alexa 647, Alexa 532, etc.), so that fluorescence emission from one population can be detected independently of fluorescence emission from another population, and vice versa. For example, the fluorescence emitted from an oligonucleotide hybridized to one identification sequence may be at least about 10% stronger or weaker, at least about 15% stronger or weaker, at least about 20% stronger or weaker, at least about 25% stronger or weaker, at least about 30% stronger or weaker, at least about 35% stronger or weaker, at least about 40% stronger or weaker, at least about 45% stronger or weaker, or at least about 50% stronger or weaker than the fluorescence emitted from an oligonucleotide hybridized to the other identification sequence. In another example, the fluorescence emitted from an oligonucleotide hybridized to one identification sequence may be about 10% stronger or weaker, or about 15% stronger or weaker, or about 20% stronger or weaker, or about 25% stronger or weaker, or about 30% stronger or weaker, or about 35% stronger or weaker, or about 40% stronger or weaker, or about 45% stronger or weaker, or about 50% stronger or weaker, than the fluorescence emitted from an oligonucleotide hybridized to the other identification sequence.

[0029] An example is shown in Figure 2. The leftmost panel shows two polynucleotides, each from two populations of polynucleotides. Each contains an index or identification sequence. In a real sample, multiple polynucleotides from two or more populations may be used. The surface where solid-phase replication occurs is shown. In this example, the surface is the surface of a flow cell. Complementary and hybridizable primers (e.g., primers P5 and P7) are bound to the surface. The polynucleotide is then hybridized to the surface-bound primers, which are then extended by polymerase to form a polynucleotide complement. The polynucleotide is then dehybridized, separating from the surface-bound complement that was extended from the surface-bound primer. The result is the formation of a surface-bound copy of the population's polynucleotide, polymerized using the surface-bound complements of two or more populations' polynucleotides as templates. Such chains are then linearized, dehybridized from each other, and the process is repeated thereon. By repeatedly performing this process, the number of copies and their complements bound to the surface of two or more polynucleotide populations is amplified, thereby creating surface-bound clusters. Subsequently, a pair of chains corresponding to either copies or complements of the polynucleotide population can be removed from the substrate by enzymatic cleavage or other means. See the arrow in Figure 2 indicating "clustering and linearization".

[0030] If a feature or combination thereof that distinguishes one of two or more populations from another population causes bias in the replication process, or if differences in the replication process have an effect, such bias may be reflected in the difference in the amount of surface-bound copies of polynucleotides from one population compared to another population. Such differences can be confirmed by hybridizing an oligonucleotide probe having a detectable appendage, such as a fluorescent marker, that can hybridize to a identifiable sequence of one population but not to any other identifiable sequence. As shown in the third panel of Figure 2, such a probe can hybridize to a copy of one population, excess unbound probe can be washed away, and then the amount of hybridized probe can be detected by measuring the extent to which fluorescence is emitted after stimulating the surface with electromagnetic radiation of a wavelength known to induce emission from the fluorescent marker bound to the oligonucleotide probe.

[0031] Subsequently, the first probe may be dehybridized, washed away, and then hybridized with another probe. This other oligonucleotide is capable of hybridizing to a recognition sequence of another population (but not to any other recognition sequence) and possesses a detectable appendage, such as a fluorescent marker. As shown in the last panel of Figure 2, such another probe may hybridize to a copy of such another population, excess unbound probe may be washed away, and the amount of hybridized probe can then be detected by measuring how much fluorescence is emitted after stimulating the surface with electromagnetic radiation of a wavelength known to induce emission from the fluorescent marker bound to the other oligonucleotide probe. A comparison of the amount of fluorescence detected from the first hybridized probe and the second hybridized probe indicates the difference in how many copies of polynucleotides from two or more populations are bound to the surface.

[0032] By comparing this difference with the difference in the amount of polynucleotides from each of the two or more populations of polynucleotides used to initiate the clustering process, the effect of one or more features distinguishing the two or more populations of polynucleotides on the replication bias, or bias caused by the replication method, can be identified. That is, the existence and magnitude of a given bias in replication can be confirmed by comparing with each other the amount of each such oligonucleotide hybridized into copies of each of the two or more populations, normalized by the relative amounts of polynucleotides from each population used for replication. For example, if a feature causes a bias in replication, the relative amount of copies of polynucleotides from one of the two or more populations of polynucleotides characterized by that feature (e.g., higher GC content, longer length, sample preparation process, any combination of two or more of the above) may exceed the relative amount of copies of polynucleotides from the other population of polynucleotides characterized by differences based on that feature (e.g., lower GC content, shorter length, different sample preparation process, or any other different combination of two or more of the above). Next, the detection of such differences may indicate a bias towards or against the replication direction of the polynucleotide characterized by the feature or combination of features.

[0033] The solid-phase amplification process results in the formation of copies and complements of the initial population of polynucleotides bound to the surface. The copies of the polynucleotides contain the identification sequence. The complements of these copies then contain the complements of the identification sequence, and these complements of the identification sequence may also be uniquely hybridizable to oligonucleotide probes that do not hybridize to other copies or complements of the polynucleotide bound to the surface. The extent to which polynucleotide replication occurred during the replication process is indicated by hybridizing oligonucleotides to the identification sequence on the copies bound to the surface of the polynucleotides and measuring the amount of such hybridized oligonucleotides. Similarly, the extent to which polynucleotide replication occurred during the replication process is also indicated by hybridizing oligonucleotides to the complements of the identification sequence on the complements bound to the surface of the copies of the polynucleotides and measuring the amount of such hybridized oligonucleotides. The detection of either a probe that is hybridizable to and hybridized with the identification sequence of a copy bound to the surface of a polynucleotide, or a probe that is hybridizable to and hybridized with the complement of the identification sequence of the complement bound to the surface, can be used as an indicator of the extent to which replication has occurred in a population of polynucleotides.

[0034] In some examples, polynucleotides from two populations may not have identical identification sequences, while in other examples, polynucleotides from two or more populations may contain identical identification sequences. For example, a polynucleotide may contain two or more identification sequences. Polynucleotides from all of two or more populations of a polynucleotide may have a first identification sequence that is different in each population. They may have a second identification sequence that is shared among two or more of the two or more populations, but different from any other population among the two or more populations. A population may have a third identification sequence that is also shared by some populations, but makes it distinguishable from other populations. A population may further have a fourth identification sequence that is shared by all populations. In this example, the difference between populations in a given identification sequence may be such that an oligo that is hybridizable by a hybridization sequence relating to one sequence is not hybridizable by an identification sequence relating to another sequence. Therefore, the population from which the copies and complements bound to the surface of a polynucleotide originate can be determined by hybridizing probes specific to a given identification sequence.

[0035] In a non-restrictive example, there may be four populations of polynucleotides. Two populations may have a higher GC content than the other two populations, and two populations may have longer polynucleotides than the other two populations. Length and GC content may be combined among the four populations; that is, the first population has long polynucleotides with high GC content, the second population has long polynucleotides with low GC content, the third population has short polynucleotides with high GC content, and the fourth population has short polynucleotides with low GC content. Each population may have one, two, three, four, or more identifying sequences. The first identifying sequence may be unique to each population. The second identification sequence can distinguish between groups of different lengths; that is, the first and second groups have the same second identification sequence, the third and fourth groups have the same identification sequence, and the second identification sequence of the first and second groups is different from the second identification sequence of the third and fourth groups. The third identification sequence can distinguish between groups of different GC content; that is, the first and third groups have the same second identification sequence, the second and fourth groups have the same identification sequence, and the third identification sequence of the first and third groups is different from the third identification sequence of the second and fourth groups. The fourth identification sequence can be shared by all groups.

[0036] After replication, the hybridization of probes specific to a given identification sequence and the measurement of such hybridization can indicate, for different populations or combinations of differences in polynucleotides, the amount of surface-bound copies, i.e., the degree of replication, depending on the features that distinguish them or that they share. For example, the amount of each population can be determined individually by measuring the hybridization of probes specific to each sequence of a first hybridization sequence. The replication amounts of long and short polynucleotides can be determined by measuring the hybridization of oligonucleotides to each sequence of a second identification sequence. The replication amounts of high-GC-content and low-GC-content polynucleotides can be determined by measuring the hybridization of nucleotides to each sequence of a third identification sequence. The total replication amount can then be determined as a whole by measuring the hybridization of nucleotides to a fourth identification sequence. In other examples, more or fewer identification sequences may be included in some or all populations of polynucleotides, and different populations may be combined in different ways. In other examples, for instance, when several examples of a given feature (e.g., low GC content, intermediate GC content, or high GC content, or short polynucleotide length, intermediate polynucleotide length, and long polynucleotide length) are compared, there may be three or more sequences that can be present in a given identification sequence.

[0037] Although cluster formation has occurred, copies and complements of two or more populations of polynucleotides are bound to the surface before the hybridization of oligonucleotides to the identification sequence as described above. It may be advantageous to remove the surface-bound complements before evaluating the hybridization of oligonucleotides to the identification sequence. Alternatively, in another example, it may be advantageous to remove the surface-bound copies before measuring the hybridization of oligonucleotides to the identification sequence complements of the surface-bound complements. Removal of surface-bound copies or complements of two or more populations of polynucleotides can be achieved by including a range of such copies and complements, a residue that can be selectively cleaved in the surface-bound primer, and removing the copies or complements extending from the primer after cluster formation. For example, the primer may contain a deoxyuridine (dU) moiety. Subsequent treatment with an enzyme preparation such as LMX1 can cleave the primer at the dU residue, releasing the polynucleotides extending from the primer. In another example, the surface-bound primer may contain an 8-oxoguanine (oxo-G) residue. Subsequent treatment with enzyme preparations such as LMX2 can cleave the primer at the oxo-G residue, releasing the polynucleotide extending from the primer.

[0038] Furthermore, aspects of replication processes, such as clustering processes, may be modified or compared to determine whether such aspects reduce, eliminate, partially eliminate, exacerbate, or otherwise affect bias arising from the characteristics of the polynucleotide population. For example, if a characteristic that distinguishes polynucleotides from two or more different populations is determined to give rise to, cause, or be associated with the characteristic by the methods disclosed herein, replication, such as a clustering process, may be performed under different conditions, and the effect of such differences in replication conditions on such bias can be determined. A aspect of the process by which copies of a polynucleotide population are made (e.g., an aspect of a clustering process) may be a characteristic, and such characteristic may differ between different populations. For example, polynucleotides from two different populations that differ in a first characteristic (e.g., GC content, length, and / or preparation process, in non-limiting examples) may be replicated under each of two different conditions (the second characteristic). Any difference in the amount of copies of polynucleotides from the two populations replicated under one condition, indicating that the first characteristic is associated with bias in replication, may then be compared to any difference in the amount of copies of polynucleotides from the two populations replicated under the other condition. If such differences are distinct from each other, it would indicate that differences in replication conditions, caused by biases associated with the features, can be influenced by the conditions (i.e., by the second feature). In another example, the distinguishing feature could be a manner of making copies of two or more populations of polynucleotides, for example, a manner of a clustering process, and no other features would be distinct.

[0039] For example, a bias related to a feature that is reflected in the difference in the degree to which two populations are replicated when replicated under certain conditions may be smaller or larger than a bias related to a feature that is reflected in the difference in the degree to which two populations are replicated when replicated under different conditions (indicated by a smaller or larger difference in the amount of copies between the two populations). Any component, circumstances, environment, or other aspect of replication may be modified or tested for its effect on the bias reflected in the difference in the degree to which two populations of polynucleotides, distinguished with respect to features, are replicated. For example, different polymerases, additives in the polymerization reaction (e.g., polyethylene glycol, salts, nucleotides, etc.), substrates, polymer coatings of substrates, flow cell properties, temperature, timing, or number of polymerization cycles, components used to rehybridize or linearize copies of polynucleotides and their complements (e.g., LMX1 or LMX2 used in some biochemical processes to release a subset of surface-bound polynucleotides after cluster formation but before resynthesis of surface-bound polynucleotides for subsequent sequencing reactions), or any other conditions may be modified and compared. Two or more examples of conditions may be compared with two or more other examples. Furthermore, multiple conditions may be modified, for example, to determine whether there is an interaction between multiple conditions for replication bias associated with a feature. In addition, numerous features may be compared as described above for their individual and combined effects on bias, and one or more conditions may be tested individually, in combination, or in combination for their effect on bias or multiple biases associated with any given feature.

[0040] The methods disclosed herein offer advantages over other methods for evaluating potential bias in clustering or other replication processes used in next-generation sequencing technologies. For example, as disclosed herein, the causes of bias can be evaluated without requiring sequencing of replicated polynucleotides. Potential causes of bias, and adjustments that may minimize, eliminate, or otherwise affect the bias, can be identified without the need to proceed with extra time, expense, and the computational burden of performing and analyzing polynucleotide sequencing. Furthermore, the examples disclosed herein, individually or in combination, provide high-throughput methods for evaluating multiple possible causes of bias, such as the morphology of polynucleotide features, and multiple variables that may be modified to eliminate or otherwise alter bias in replication, such as the conditions under which replication, e.g., clustering occurs, or the conditions under which replication, e.g., clustering occurs, or conditions that may otherwise affect any aspect of polynucleotide replication that occurs before sequencing in the SBS process.

[0041] Regarding the evaluation of bias arising from GC content as a characteristic, a population of polynucleotides can be characterized by its average relative GC content. For example, certain species of microorganisms are known to have a relatively higher or lower average percentage of GC content compared to, for example, humans, whose genomes have roughly equal ratios of GC and AT content on average. Some microorganisms, such as bacteria of the genus Rhodobacter, are known to have high GC content, such as over 60%. Other microorganisms, such as Bacillus cereus, are known to have lower GC content, such as less than 40%. In one example, GC content can be a characteristic. A population of polynucleotides may be polynucleotides prepared from the genus Rhodobacter, which exhibit a higher GC content as a characteristic compared to other populations, and another population may be prepared from humans, which exhibit an intermediate GC content as a characteristic compared to other populations, or from Bacillus cereus, which exhibits a lower GC content as a characteristic compared to other populations. "Higher" and "lower" are used relatively in this specification. Therefore, when humans are used as a population, that population may have a higher or lower GC compared to other populations, depending on the GC content of such other populations (e.g., Bacillus cereus and Rhodobacter, respectively).

[0042] In other examples, polynucleotides may be of synthetic or artificial origin with a predetermined GC content, determined, for example, by directly determining the sequence of the polynucleotide in a sample, or by stoichiometrically controlling the relative amount of incorporation of a given type of nucleotide into the chain, depending on the synthesis method (e.g., using a template-independent method for sequence synthesis). A population of polynucleotides may contain any intended or known GC content, where GC content refers to the total number of guanine and cytosine nucleic acid bases combined out of the total number of nucleic acid bases (sum of G, C, A, and T). A population may be defined by its GC content as a characteristic of the population as a whole, even if individual polynucleotides in that population may have different GC content from the overall population GC content.

[0043] The population may have a GC content of approximately 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or any intermediate GC content. All other possible comparisons are expressly included as aspects of this disclosure.

[0044] Regarding the evaluation of bias due to polynucleotide length, a population of polynucleotides can be characterized by its average relative polynucleotide length. For example, nucleic acid molecules can be isolated from a sample such as a cell or other biological source and can be fragmented during sample preparation by one of a variety of methods. By adjusting parameters used in the fragmentation method, such as sonication time, polynucleotides of various lengths can be produced. Polynucleotides of the desired length can then be isolated from the resulting fragments. A population can be defined by its polynucleotide length as a characteristic of the population as a whole, even if individual polynucleotides in that population may have lengths that are judged to differ from the polynucleotide length of the population. In another example, the polynucleotide length as a characteristic of a population may be predetermined by polymerizing polynucleotides of a predetermined length from a designed template.

[0045] Polynucleotide clusters are approximately 100 nucleotides long, 150 nucleotides long, 200 nucleotides long, 250 nucleotides long, 300 nucleotides long, 350 nucleotides long, 400 nucleotides long, 450 nucleotides long, 500 nucleotides long, 550 nucleotides long, 600 nucleotides long, 650 nucleotides long, 700 nucleotides long, 750 nucleotides long, 800 nucleotides long, 850 nucleotides long, 900 nucleotides long, 950 nucleotides long, 1,000 nucleotides long, and 1,050 nucleotides long. It may have lengths of approximately 1,200 nucleotides, 1,250 nucleotides, 1,300 nucleotides, 1,350 nucleotides, 1,400 nucleotides, 1,450 nucleotides, 1,500 nucleotides, 1,550 nucleotides, 1,600 nucleotides, 1,650 nucleotides, 1,700 nucleotides, 1,750 nucleotides, 1,800 nucleotides, 1,850 nucleotides, 1,900 nucleotides, 1,950 nucleotides, 2,000 nucleotides, or more.

[0046] Other characteristics may include nucleotide populations (DNA libraries) prepared by other aspects of population preparation, such as different library preparation methods or library preparation kits from different suppliers. The effects of aspects of the clustering process may also be compared to determine whether and to what extent such conditions affect bias or assumed bias related to the features. Each of two or more populations of polynucleotides distinguished from one or more by one or more features may be subjected to one, two or more replication processes, such as a clustering process, using different conditions between the two or more processes. Whether replication bias arises due to differences in features under certain and / or other conditions, or whether the amount or presence of bias differs depending on which replication conditions are applied, may indicate whether the bias related to the features can be altered by such modification of the conditions. For example, multiple conditions may be modified, or several different examples of conditions may be compared. Examples of conditions that may be modified include the components of the reaction solution used in solid-phase replication processes such as cluster formation (e.g., polymerase used, pH, concentration of polynucleotides or any other components, linearizing enzymes used such as LMX1 or LMX2 or others, performance-enhancing additives including GP32 or UvsX or other nucleotide-binding proteins, polyethylene glycol, creatine phosphate, or other additives, type of substrate surface, or type, presence, or thickness of polymer coating on the surface where solid-phase PCR or cluster formation occurs), the number or duration of replication reactions, and temperature. In some examples, any aspect of the replication method, such as the method for preparing a population of polynucleotides and / or the method for replication in cluster formation from a population of polynucleotides, may be modified and evaluated by the methods disclosed herein to determine the possibility of bias.

[0047] Non-limiting examples of parameters that may vary as characteristics of reagents used in methods for creating copies (e.g., ExAmp, bridging, or other clustering processes) include various enzyme concentrations (and ratios), additives (concentrations and ratios), solution pH, polynucleotide concentration in the replication solution, and nucleotide concentration in the replication solution.

[0048] Non-limiting examples such as the temperature at which replication occurs (e.g., in one example, around 20 degrees Celsius, above 20 degrees Celsius, or below 20 degrees Celsius), the duration of the polymerization or washing process, and the reagent replenishment method or duration, as well as the reagent flow rate to the flow cell or other substrate used, may be modified and investigated accordingly. Cluster formation times can differ in their characteristics (for example, less than approximately 30 minutes, or within the range of approximately 30 minutes to approximately 1 hour, or within the range of approximately 1 hour to approximately 2 hours, or within the range of approximately 2 hours to approximately 3 hours, or within the range of approximately 3 hours to approximately 4 hours, or within the range of approximately 4 hours to approximately 5 hours, or within the range of approximately 5 hours to approximately 6 hours, or within the range of approximately 6 hours to approximately 7 hours, or within the range of approximately 7 hours to approximately 8 hours, or within the range of approximately 8 hours to approximately 9 hours, or within the range of approximately 9 hours to approximately 10 hours, or within the range of approximately 10 hours to approximately 11 hours, or within the range of approximately 11 hours to approximately 12 hours, or within the range of approximately 12 hours to approximately 24 hours, or within the range of approximately 24 hours to approximately 36 hours, or within the range of approximately 36 hours to approximately 48 hours, or within the range of approximately 48 hours to approximately 72 hours, or longer). The duration of each aspect of the replication or clustering process, for example, the incubation period of the reagent in solution, may differ in characteristics (e.g., about 10 seconds, or 20 seconds, or 30 seconds, or 40 seconds, or 50 seconds, or 60 seconds, or 70 seconds, or 80 seconds, or 90 seconds, or 100 seconds, or 110 seconds, or 120 seconds, or longer).

[0049] The fluid velocity or the rate at which the reagent flows through the flow cell may be a distinguishing feature (for example, the flow velocity may be about 10 uL / min, or about 20 uL / min, or about 30 uL / min, or about 40 uL / min, or about 50 uL / min, or about 60 uL / min, or about 70 uL / min, or about 80 uL / min, or about 90 uL / min, or about 100 uL / min, or about 110 uL / min, or about 120 uL / min, or about 130 uL / min, or about 140 uL / min, or about 150 uL / min, or higher). Other embodiments that may be modified in terms of features include pH (e.g., greater than approximately pH 7.5, less than approximately pH 7.5, or approximately pH 7.5), type of buffer (e.g., Tris-based or other), and concentration of other or more components of the buffer or cluster-forming solution (e.g., approximately 100 nM, less than 100 nM, or greater than 100 nM).

[0050] In one example, polynucleotides from two or more populations may be combined in a solution, and this solution may be added to a substrate such as a flow cell for replication, for example, involving replication by a clustering process. Different identification sequences of polynucleotides from two or more populations allow for the identification of which population a copy bound to a surface belongs to. In one example, the relative proportion of polynucleotides from one population to the total amount of polynucleotides added may be controlled and may differ between several different solutions. For example, the total polynucleotides in a solution added to a flow cell or a lane of a flow cell may have approximately equal proportions of polynucleotides from each of two populations of polynucleotides. The total polynucleotides in another solution added to a flow cell or a lane of a flow cell may have about 25% polynucleotides from one population and about 75% from the other population, while the total polynucleotides in yet another solution added to a flow cell or a lane of a flow cell may have about 75% polynucleotides from the other population and about 25% from the first population. Any other division between the ratios of polynucleotides from one population and other populations may be used in different solutions (e.g., approximately 5% / 95%, approximately 10% / 90%, approximately 15% / 85%, approximately 20% / 80%, approximately 25% / 75%, approximately 30% / 70%, approximately 35% / 65%, approximately 40% / 60%, approximately 45% / 55%, approximately 50% / 50%, approximately 55% / 45%, approximately 60% / 40%, approximately 65% / 35%, approximately 70% / 30%, approximately 75% / 25%, approximately 80% / 20%, approximately 85% / 15%, approximately 90% / 10%, approximately 95% / 5%, or any intermediate relative ratio).

[0051] Library preparation A library containing polynucleotides can be prepared by any preferred method for attaching oligonucleotide adapters to target polynucleotides. As used herein, “library” is a collection of polynucleotides derived from a given source or sample. The library contains several target polynucleotides. As used herein, “target polynucleotide” is a polynucleotide that is desired to be included in a replication process, such as a clustering process. The target polynucleotide can be any polynucleotide of known or unknown sequence. The target polynucleotide may be, for example, a fragment of genomic DNA or cDNA. The target polynucleotide may be derived from a randomly fragmented primary polynucleotide sample. The target polynucleotide may be processed into a template suitable for amplification by placing primer sequences, such as a recognition sequence or a sequence complementary to the surface-bound primer, at the ends of each target fragment. The target polynucleotide may be obtained from a primary RNA sample by reverse transcription to cDNA.

[0052] As used herein, the terms “polynucleotide” and “oligonucleotide” may be used interchangeably and typically refer to molecules comprising two or more nucleotide monomers covalently linked to one another by phosphodiester bonds. Polynucleotides typically contain more nucleotides than oligonucleotides. For illustrative purposes only and not limiting purposes, polynucleotides may be considered to contain 15, 20, 30, 40, 50, 100, 200, 300, 400, 500 nucleotides or more, while oligonucleotides may be considered to contain 100, 50, 20, 15 nucleotides or fewer.

[0053] Examples of polynucleotides and oligonucleotides include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). These terms should be understood to encompass, as equivalents, analogues of either DNA or RNA derived from nucleotide analogues, and to be applicable to single-stranded (e.g., sense or antisense) polynucleotides and double-stranded polynucleotides. As used herein, these terms also include cDNA, which is complementary DNA or copy DNA produced from an RNA template by the action of reverse transcriptase.

[0054] Primary polynucleotide molecules may arise from double-stranded DNA (dsDNA) morphology (e.g., genomic DNA fragments, PCR products, and amplification products), or from single-stranded morphology as DNA or RNA, and may be converted to dsDNA morphology. For example, mRNA molecules can be replicated to double-stranded cDNA using standard techniques well known in the art. The exact sequence of primary polynucleotides is generally not important to the disclosures presented herein and may be known or unknown.

[0055] In some examples, the primary target polynucleotide is an RNA molecule. In embodiments of such examples, RNA isolated from a particular sample is first converted to double-stranded DNA using techniques known in the art. The double-stranded DNA can then be index-tagged with a library-specific tag. Different preparations of such double-stranded DNA containing library-specific index tags can be produced in parallel from RNA isolated from different sources or samples. Subsequently, these different preparations of double-stranded DNA containing different library-specific index tags can be mixed, replicated together, and the identity of each sequenced fragment to the population from which that fragment was isolated / taken can be determined by the presence of the library-specific index tag sequence.

[0056] In some cases, the primary target polynucleotide is a DNA molecule. For example, the primary polynucleotide may correspond to the entire genetic complement of an organism, and is a genomic DNA molecule such as a human DNA molecule, which includes intron and exon sequences (coding sequences) as well as non-coding regulatory sequences, such as promoter and enhancer sequences. However, it can also be assumed that a polynucleotide sequence or a specific subset of genomic DNA, such as a specific chromosome or a part thereof, may be used. In many cases, the sequence of the primary polynucleotide is unknown. The DNA target polynucleotide may be chemically or enzymatically treated before or after fragmentation treatments such as random fragmentation, and before, during, or after ligation of adapter oligonucleotides.

[0057] In one example, the primary target polynucleotide is fragmented into appropriate lengths suitable for sequencing. The target polynucleotide can be fragmented in any preferred way. Preferably, the target polynucleotide is fragmented randomly. Random fragmentation refers to the disorderly fragmentation of the polynucleotide, for example, by enzymatic, chemical, or mechanical means. Any preferred fragmentation method can be used. To be clear, generating smaller fragments of a larger portion of a polynucleotide by specific PCR amplification of such smaller fragments is not equivalent to fragmenting a larger portion of the polynucleotide because the larger portion of the polynucleotide remains intact (i.e., not fragmented by PCR amplification) (however, the methods disclosed herein can be performed on a population of polynucleotides produced by any of these techniques). Furthermore, random fragmentation is designed to generate fragments regardless of the sequence identity or position of the nucleotide containing the cleavage and / or the nucleotides surrounding the cleavage.

[0058] In some cases, random fragmentation is performed by mechanical means, such as atomization or ultrasonic treatment, to generate fragments of about 50 base pairs to about 1500 base pairs in length, for example, about 50 to 700 base pairs or about 50 to 500 base pairs in length.

[0059] Fragmentation of polynucleotide molecules by mechanical means (e.g., atomization, sonication, and hydroshearing) can result in fragments with a heterogeneous mixture of blunt ends and 3' and 5' overhangs. For example, fragment ends can be repaired using methods or kits known in the art (e.g., the Lucigen DNA Terminator End Repair Kit) to produce ends best suited for insertion into blunt regions of a cloning vector. In some examples, the fragment ends of a population of nucleic acids are blunt ends. Fragment ends may be blunt ends or phosphorylated. Phosphate moieties can be introduced, for example, by enzymatic treatment with polynucleotide kinases.

[0060] In some cases, a target polynucleotide sequence with a single overhanging nucleotide is prepared by the activity of a specific type of DNA polymerase, such as Taq polymerase or Klenow exo-minus polymerase, which has template-independent terminal transferase activity to add, for example, a single deoxynucleotide, such as deoxyadenosine (A), to the 3' end of a PCR product. Such enzymes can be used to add a single nucleotide "A" to the 3' blunt end of each strand of a double-stranded target polynucleotide. Thus, "A" can be added to the 3' end of each strand of the double-stranded target polynucleotide after reaction with Taq or Klenow exo-minus polymerase, while the adapter polynucleotide construct may be a T construct in which a compatible "T" overhang exists at the 3' end of each double-stranded region of the adapter construct. This terminal modification also prevents self-ligation of the target polynucleotide, resulting in a bias toward the formation of a combined, ligated adapter-target polynucleotide.

[0061] In some cases, fragmentation is achieved by tagmentation. In this method, a transposase is used to fragment a double-stranded polynucleotide and to ligate a universal primer sequence to one of the strands of the double-stranded polynucleotide. The gaps in the resulting molecule can be filled and subjected to extension by PCR amplification using a primer containing, for example, a 3' end with a sequence complementary to the ligated universal primer sequence and a 5' end containing other sequences of the adapter.

[0062] The adapter may be conjugated to the target polynucleotide in any other preferred manner. In some examples, the adapter may be introduced in a single-step process. In some examples, the adapter may be introduced in a multi-step process, such as a two-step process, which includes ligation of a portion of the adapter having a universal primer sequence to the target polynucleotide. The second step includes extension by PCR amplification using a primer having a 3' end with a sequence complementary to the conjugated universal primer sequence and a 5' end containing other sequences of the adapter. Additional extension may be performed to provide an additional sequence to the 5' end of the already extended polynucleotide.

[0063] In some examples, the entire adapter is ligated to a fragmented target polynucleotide. Preferably, the ligated adapter includes a double-stranded region ligated to the double-stranded target polynucleotide. Preferably, the double-stranded region is as short as possible without losing function. In this context, “function” refers to the ability of the double-stranded region to form a stable double helix under standard reaction conditions. In some examples, standard reaction conditions refer to reaction conditions for an enzyme-catalyzed polynucleotide ligation reaction (e.g., incubation at a temperature in the range of 4°C to 25°C in an enzyme-suitable ligation buffer) such that the two strands forming the adapter remain partially annealed while the adapter is ligated to the target molecule. The ligation method in this case utilizes a ligase enzyme, such as a DNA ligase, to achieve or catalyze the ligation of the ends of the two polynucleotide strands of the adapter oligonucleotide double helix and the target polynucleotide double helix so that a covalent bond is formed. The adapter oligonucleotide double helix may include a 5'-phosphate moiety to facilitate the ligation of the target polynucleotide to the 3'-OH group. The target polynucleotide may contain a 5'-phosphate moiety either left by a shearing process or added using an enzymatic process, may be end-repaired, and may be optionally elongated by a protruding base or multiple bases to produce a 3'-OH suitable for ligation. In this context, bonding means covalent bonding of polynucleotide chains that were not previously covalently bonded. In one embodiment, such bonding occurs by the formation of a phosphodiester linkage between two polynucleotide chains, but other means of covalent bonding (e.g., non-phosphodiester backbone linkage) may be used.

[0064] Any suitable adapter may be bound to the target polynucleotide by any suitable process, such as the process described above. The adapter contains a library-specific index tag sequence. The index tag sequence may be bound to the target polynucleotide from each library before the sample is immobilized for sequencing. The index tag is not formed by any part of the target polynucleotide itself, but becomes part of the template for amplification. The index tag may be a synthetic sequence of nucleotides that is added to the target as part of the template preparation process. Thus, the library-specific index tag is a nucleic acid sequence tag that is bound to each target molecule in a particular library, and its presence is used to indicate or identify the library from which the target molecule was isolated.

[0065] Preferably, the index tag sequence is 20 nucleotides or less in length. For example, the index tag sequence may be 1 to 10 nucleotides or 4 to 6 nucleotides in length. A 4-nucleotide index provides the feasibility of multiplexing 256 samples on the same array, while a 6-nucleotide index allows for processing 4,096 samples on the same array.

[0066] The adapter may contain two or more index tags (or identification sequences), which may increase the feasibility of redundancy.

[0067] The adapter may include a double-stranded region and a region containing two non-complementary single-stranded regions. The double-stranded region of the adapter may have any suitable number of base pairs. Preferably, the double-stranded region is a short double-stranded region typically containing five or more consecutive base pairs, formed by the annealing of two partially complementary polynucleotide chains. This “double-stranded region” of the adapter refers to a region where two strands are annealed and does not imply any particular three-dimensional structure. In some examples, the double-stranded region contains 20 or fewer consecutive base pairs, e.g., 10 or fewer or 5 or fewer consecutive base pairs.

[0068] The inclusion of non-natural nucleotides that exhibit stronger base pairing than standard Watson-Crick base pairs can increase the stability of the double-stranded region, and therefore potentially reduce its length. The two strands of the adapter may be 100% complementary in the double-stranded region.

[0069] When an adapter is bound to a target polynucleotide, the non-complementary single-stranded region may form the 5' and 3' ends of the sequenced polynucleotide. The term "non-complementary single-stranded region" refers to a region of the adapter where the sequences of the two polynucleotide strands forming the adapter exhibit a degree of noncomplementarity such that the two strands cannot fully anneal to each other under standard annealing conditions of a PCR reaction.

[0070] The non-complementary single-stranded region is provided by different portions of the same two polynucleotide chains that form the double-stranded region. The lower limit of the length of the single-stranded region is usually determined by its function of providing a sequence suitable for primer binding for primer extension, PCR, and / or sequencing. Generally, there is theoretically no upper limit to the length of the unmatched region, except that it is advantageous to minimize the total length of the adapter to facilitate the separation of the unbound adapter from the adapter-target constructor after the binding step or multiple binding steps. Therefore, the non-complementary single-stranded region of the adapter is generally preferably 50 consecutive nucleotides or less in length, for example, 40 consecutive nucleotides or less, 30 consecutive nucleotides or less, or 25 consecutive nucleotides or less.

[0071] The library-specific index tag sequence may reside in a single-stranded region, a double-stranded region, or span both the single-stranded and double-stranded regions of the adapter. Preferably, the index tag sequence resides in the single-stranded region of the adapter.

[0072] The adapter may include any other suitable sequences in addition to the index tag sequence. For example, the adapter may include a universal elongation primer sequence typically present at the 5' or 3' end of the adapter and the resulting polynucleotide for sequencing. The universal elongation primer sequence may hybridize to a complementary primer bound to the surface of a solid substrate. The complementary primer includes a free 3' end to which a nucleotide can be added to elongate the sequence using the hybridized library polynucleotide as a template by a polymerase or other suitable enzyme, resulting in the reverse chain of the library polynucleotide bound to the solid surface. Such elongation may be part of a sequencing run or cluster amplification.

[0073] In some examples, the adapter includes one or more universal sequencing primer sequences. The universal sequencing primer sequences can bind to sequencing primers to enable sequencing of an index tag sequence, a target sequence, or both an index tag sequence and a target sequence.

[0074] The precise nucleotide sequence of the adapter is generally not critical and can be selected by the user, for example, to provide binding sites for a specific set of universal extension primers and / or sequencing primers, so that the desired sequence elements are ultimately included in the template common sequence of the library derived from the adapter.

[0075] Adapter oligonucleotides may contain exonuclease resistance modifications such as phosphorothioate linkages.

[0076] Preferably, the adapters are attached to both ends of the target polypeptide to generate a polynucleotide having a first adapter-target-second adapter nucleotide sequence. The first and second adapters may be the same or different. If the first and second adapters are different, at least one of them contains a library-specific identification sequence.

[0077] The "first adapter-target-second adapter sequence" or "adapter-target-adapter" sequence refers to the relative positions of the adapters to each other and to the target, and does not necessarily mean that the sequence may not include additional sequences, such as linker sequences.

[0078] Other libraries can be prepared in a similar manner, each containing at least one library-specific index tag sequence or combination of index tag sequences that is different from the index tag sequences or combinations of index tag sequences of other libraries.

[0079] As used herein, “attached” and “bound” are used interchangeably in the context of an adapter to a target sequence. As previously stated, any preferred process may be used to bind the adapter to the target polynucleotide. For example, the adapter may be bound to the target by ligation with a ligase; a combination of ligation of part of the adapter and addition of further or remaining parts of the adapter by extension such as PCR using primers containing further or remaining parts of the adapter; rearrangement to incorporate part of the adapter and addition of further or remaining parts of the adapter by extension such as PCR using primers containing further or remaining parts of the adapter; and so on. Preferably, the bound adapter oligonucleotide is covalently bonded to the target polynucleotide.

[0080] After the adapters have bound to the target polynucleotide, the resulting polynucleotide can be subjected to a cleanup process to increase the purity of the adapter-target-adapter polynucleotide by removing at least some of the unbound adapters. Any suitable cleanup process, such as electrophoresis or size exclusion chromatography, may be used. In some examples, solid-phase reverse immobilization (SPRI) paramagnetic beads may be used to separate the adapter-target-adapter polynucleotide from the unbound adapters. While such a process can increase the purity of the resulting adapter-target-adapter polynucleotide, some unbound adapter oligonucleotides may remain.

[0081] Methods for amplifying immobilized adapter-target-adapter molecules include, but are not limited to, bridge amplification and kinetic exclusion. Amplification may be performed using one or more immobilized primers. The immobilized primer(s) may be a lawn on a flat surface.

[0082] As used herein, the term “solid-phase amplification” refers to any nucleic acid amplification reaction carried out on or in association with a solid support such that all or part of the amplification product is immobilized on the solid support when the amplification product is formed. In particular, the term encompasses solid-phase polymerase chain reactions (solid-phase PCR) and solid-phase isothermal amplification, which are reactions similar to standard solution-phase amplification, except that one or both of the forward and reverse amplification primers are immobilized on the solid support. Solid-phase PCR includes systems such as colonization in an emulsion in which one primer is anchored to a bead and the other is in free solution, and colonization in a solid-phase gel matrix in which one primer is anchored to a surface and the other is in free solution.

[0083] In some embodiments, the solid support includes a patterned surface. “Patterned surface” refers to the arrangement of different regions within or on the exposed layer of the solid support. The terms “support” or “substrate” refer to a support or substrate to which surface chemicals may be added. “Patterned substrate” refers to a support to which recesses are defined within or on the support. “Unpatterned substrate” refers to a substantially flat support. In this specification, substrate may also be referred to as “support,” “patterned support,” or “unpatterned support.” The support may be a wafer, panel, rectangular sheet, die, or any other suitable configuration. The support is generally rigid and insoluble in aqueous liquids. The support may be inert to surface chemicals used to modify recesses. For example, the support may be inert to surface chemicals used to form a polymer coating layer, to bond a primer to a deposited polymer coating layer, etc. Examples of suitable supports include epoxysiloxanes, glass and modified or functionalized glass, polyhedral oligomeric silsequioxanes (POSS) and their derivatives, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, polytetrafluoroethylene (e.g., TEFLON® from Chemors), cyclo-olefin polymers (COP) (e.g., ZEONOR® from Zeon), polyimides, etc.), nylon, ceramics / ceramic oxides, silica, fused silica, or silica-based materials, aluminum silicate, silicon and modified silicon (e.g., boron-doped p+ silicon), silicon nitride (Si3N4), silicon dioxide (SiO2), tantalum pentoxide (TaO5) or other tantalum oxides (TaO5) xExamples include ), hafnium oxide (HaO2), carbon, metals, and inorganic glass. The support may be glass or silicon or silicon-based polymer, e.g., a POSS material, having optionally a coating layer of tantalum oxide or another ceramic oxide on its surface. A POSS material may be one disclosed in Kejagoas et al., Microelectronic Engineering 86(2009)776-668, which is incorporated herein by reference in its entirety.

[0084] In one example, a recess may be a well containing an array of wells on its surface, where the patterned substrate is microwells or nanowells. The size of each well may be characterized by its volume, well opening area, depth, and / or diameter. For example, one or more of the regions may be a portion where one or more amplification primers are present. This portion may be separated by an intermediate region where no amplification primers are present. In some examples, the pattern may be an xy format of features in rows and columns. In some examples, the pattern may be a repeating arrangement of partial and / or intermediate regions. In some examples, the pattern may be a random arrangement of partial and / or intermediate regions.

[0085] In some examples, the solid support includes an array of wells or recesses on its surface. This can be manufactured using a variety of techniques, including but not limited to photolithography, stamping, molding, and microetching. The technique used may depend on the composition and shape of the array substrate.

[0086] The patterned surface features may be wells (e.g., microwells or nanowells) in an array of wells on a glass, silicon, plastic, or other suitable solid support, comprising a patterned covalent gel, e.g., poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide) (PAZAM). This process produces a gel pad for sequencing, which can be stable over sequencing cycles. Covalently bonding polymers to wells is useful in various applications for maintaining the structured features of the gel over the lifetime of the structured substrate. However, in many examples, the gel does not need to be covalently bonded to the wells. For example, under certain conditions, silane-free acrylamide that is not covalently bonded to any part of the structured substrate can be used as the gel material.

[0087] In some examples, structured substrates may be prepared by patterning a solid support material using wells (e.g., microwells or nanowells), coating the patterned support with a gel material (e.g., PAZMA, SFA, or a chemically modified variant thereof, e.g., an azidized version of SFA (azido-SFA)), and polishing the gel-coated support, for example, by chemical polishing or mechanical polishing, thereby retaining the gel within the wells, but removing or inactivating substantially all of the gel in the intermediate regions of the surface of the structured substrate between the wells. Primer nucleic acids can then be bound to the gel material. A solution of target nucleic acids (e.g., fragmented human genome) can then be brought into contact with the polished substrate so that individual target nucleic acids can be seeded into individual wells via interaction with the primers bound to the gel material. However, the target nucleic acids do not occupy the intermediate regions because the gel material is absent or inactive. Amplification of the target nucleic acids will be limited to the wells because the outward movement of the growing nucleic acid colonies is prevented due to the absence or inactivity of the gel in the intermediate regions. The process is conveniently manufacturable, scalable, and utilizes conventional micro or nano fabrication methods.

[0088] The subject matter of this disclosure includes, for example, “solid-phase” amplification methods in which only one amplification primer is immobilized (other primers are typically present in free solution), but in other examples, a solid support with both forward and reverse primers immobilized may be provided. Some examples include “multiple” identical forward primers and / or “multiple” identical reverse primers immobilized on a solid support, as the amplification process may involve excess primers to maintain amplification. References to forward and reverse primers herein should be interpreted as including “multiple” such primers unless otherwise indicated by the context.

[0089] Any given amplification reaction includes at least one type of forward primer and at least one type of reverse primer that are specific to the template being amplified. However, in certain examples, the forward and reverse primers may contain template-specific portions of the same sequence and may have completely identical nucleotide sequences and structures (including any non-nucleotide modifications). In other words, it is possible to perform solid-phase amplification using only one type of primer, and such single-primer methods are included within the scope of this disclosure. Other examples may use forward and reverse primers that contain the same template-specific sequence but differ in some other structural features. For example, one type of primer may contain non-nucleotide modifications that are not present in the other.

[0090] The terms “cluster” and “colony” are used interchangeably herein to refer to distinct sites on a solid support comprising multiple identical immobilized nucleic acid strands and multiple identical immobilized complementary nucleic acid strands. The term “clustered array” refers to an array formed by such clusters or colonies. In this context, the term “array” should not be understood as requiring an ordered arrangement of clusters.

[0091] The terms “solid phase” or “surface” are used to mean either a flat array on which primers are bonded to a flat surface, such as a glass, silica, or plastic microscope slide or similar flow cell device; beads to which one or two primers are bonded and which are amplified; or an array on the surface of beads after they have been amplified.

[0092] Clustered arrays can be prepared using either a thermal cycling process or a process in which the temperature is kept constant and extension and denaturation cycles are performed using changes in reagents. In one example, an isothermal process may advantageously involve the use of lower temperatures.

[0093] It will be understood that any amplification method described herein or commonly known in the art may be used with universal primers or target-specific primers to amplify immobilized DNA fragments. Suitable amplification methods include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA). The above amplification methods may be used to amplify one or more nucleic acids of interest. For example, PCR, including multiplex PCR, SDA, TMA, and NASBA, may be used to amplify immobilized DNA fragments. In some examples, primers specifically directed to the target polynucleotide are included in the amplification reaction.

[0094] Other suitable methods for amplifying polynucleotides include oligonucleotide extension and ligation, rolling circle amplification (RCA), and oligonucleotide ligation assay (OLA) techniques. It will be understood that these amplification methods may be designed to amplify immobilized DNA fragments. For example, in some cases, the amplification method may be a ligation probe amplification or oligonucleotide ligation assay (OLA) reaction containing primers specifically directed to the target nucleic acid. In some cases, the amplification method may be a primer extension ligation reaction containing primers specifically directed to the target nucleic acid. A non-limiting example of primers for primer extension and ligation that can be specifically designed to amplify the target nucleic acid is that the amplification may include primers used in GoldenGate assays (Illumina, Inc., San Diego, CA).

[0095] Exemplary isothermal amplification methods that may be used in the methods of this disclosure include, but are not limited to, multi-displacement amplification (MDA) or isothermal chain-displacement nucleic acid amplification. Other non-PCR-based methods that may be used in this disclosure include, for example, strand-displacement amplification (SDA) or hyperbranched-strand-displacement amplification. Isothermal amplification methods may be used, for example, with strand-displacement Phi29 polymerase or BstDNA polymerase large fragments, 5'->3' exo-, for random primer amplification of genomic DNA. The use of these polymerases takes advantage of their high processing capacity and strand-displacement activity. The high processing capacity allows the polymerase to produce fragments of 10-20 kb in length. As mentioned above, smaller fragments may be produced under isothermal conditions using polymerases with lower processing capacity and strand-displacement activity, such as Klenow polymerase.

[0096] DNA polymerases are classified into families identified as A, B, C, D, X, Y, and RT based on their structural homology. Examples of DNA polymerases in Family A include T7 DNA polymerase, eukaryotic mitochondrial DNA polymerase γ, E. coli DNA Pol I (including Klenow fragments), Thermus aquaticus Pol I, and Bacillus stearothermophilus Pol I. Examples of DNA polymerases in Family B include eukaryotic DNA polymerases a, 6, and E; DNA polymerase C; T4 DNA polymerase, Phi29 DNA polymerase, and Thermococcus sp.9. 0Examples include N-7 archaeal polymerase (also known as 9°N (trademark)) and its variants, such as those disclosed in U.S. Patent Application Publication 2016 / 0032377(A1), and RB69 bacteriophage DNA polymerase. Family C includes, for example, the E. coli DNA polymerase IIIα subunit. Family D includes, for example, polymerases derived from subdomains of the Liriarchaeota phylum of archaea. Examples of DNA polymerases in Family X include, for example, eukaryotic polymerases Polβ, Polσ, Polλ, and Polμ, and S. cerevisiae Pol4. Examples of DNA polymerases in Family Y include, for example, Polη, Polι, Polκ, E. coli Pol IV (DINB), and E. coli Pol V (UmuD'2C). Examples of the RT (reverse transcriptase) family of DNA polymerases include, for example, retroviral reverse transcriptase and eukaryotic telomerase. Examples of RNA polymerases include, but are not limited to, viral RNA polymerases such as T7 RNA polymerase; eukaryotic RNA polymerases such as RNA polymerase I, RNA polymerase II, RNA polymerase III, RNA polymerase IV, and RNA polymerase V; and archaeal RNA polymerases. Other polymerases are also included among the polymerases described herein, as are any other functional polymerases, including those having sequences modified by comparison with any of the polymerase enzymes described above, which are provided only as a non-limiting list of examples.

[0097] In some cases, isothermal amplification may be performed using kinetic exclusion amplification (KEA), also known as exclusion amplification (ExAmp). The nucleic acid libraries of this disclosure may be prepared using a method comprising the step of reacting amplification reagents to produce a number of amplification sites, each containing a substantial clonal population of amplicons from individual target nucleic acids seeded at the site. In some cases, the amplification reaction proceeds until a sufficient number of amplicons are produced to fill the volume of each amplification site. In this way, filling the already seeded site to volume inhibits the target nucleic acid from landing and amplifying at that site, thereby generating a clonal population of amplicons at that site. In some cases, apparent clonality may be achieved even if the amplification site is not filled to volume before the second target nucleic acid reaches that site. Under some conditions, amplification of the first target nucleic acid may proceed to a point where a sufficient number of copies are produced to effectively exceed or overwhelm the production of copies from the second target nucleic acid transported to that site. For example, in a case where a bridge amplification process is used on a circular feature region with a diameter of less than 500 nm, it was determined that after 14 cycles of exponential amplification of the first target nucleic acid, contamination by the second target nucleic acid at the same site generates an insufficient number of contaminating amplicons to adversely affect sequential synthesis sequencing analysis on the Illumina sequencing platform.

[0098] In some cases, kinetic exclusion can occur when a process occurs at a rate fast enough to effectively exclude the occurrence of another event or process. Consider the construction of a nucleic acid array, where array sites are randomly seeded with a target nucleic acid from solution, and an amplification process generates copies of the target nucleic acid to fill each seeded site to capacity. According to the kinetic exclusion method of this disclosure, the seeding and amplification processes may proceed simultaneously under conditions where the amplification rate exceeds the seeding rate. Therefore, the relatively fast rate at which copies of the first target nucleic acid are made at the seeded sites effectively excludes the seeding of those sites by a second nucleic acid for amplification.

[0099] Kinetic exclusion can occur due to a relatively slow rate at which amplification begins (e.g., a slow rate at which the first copy of the target nucleic acid is produced) versus a relatively fast rate at which subsequent copies of the target nucleic acid (or the first copy of the target nucleic acid) are produced. In the example in the previous paragraph, kinetic exclusion occurs because of a relatively slow rate at which target nucleic acid seeding occurs (e.g., relatively slow diffusion or transport) versus a relatively fast rate at which amplification occurs to fill the site with copies of the nucleic acid species. In another example, kinetic exclusion can occur because of a delay in the formation of the first copy of the target nucleic acid seeded at the site (e.g., a delay in activation or slow activation) versus a relatively fast rate at which subsequent copies are produced to fill the site. In this example, several different target nucleic acids may be seeded at individual sites (e.g., several target nucleic acids may be present at each site before amplification). However, the formation of the first copy of any given target nucleic acid may be activated randomly, and as a result, the average rate of the first copy formation will be relatively slow compared to the rate at which subsequent copies are produced. In this case, several different target nucleic acids may be seeded at individual sites, but kinetic exclusion ensures that only one of those target nucleic acids is amplified. More specifically, once the first target nucleic acid is activated for amplification, the site is rapidly filled to capacity with its copy, thereby preventing the production of a copy of the second target nucleic acid at the site.

[0100] Amplification reagents may contain additional components that promote amplicon formation, and in some cases, increase the rate of amplicon formation. One example is a recombinase. Recombinases can promote amplicon formation by enabling repeated entry / extension. More specifically, recombinases can promote entry of target nucleic acid by polymerase and extension of primer by polymerase using the target nucleic acid as a template for amplicon formation. This process can be repeated as a chain reaction in which the amplicon generated by each round of entry / extension serves as a template in subsequent rounds. This process can be performed more rapidly than standard PCR because it does not require a denaturation cycle (e.g., by heating or chemical denaturation). Therefore, recombinase-enhanced amplification can be performed isothermally. To promote amplification, it is generally desirable to include ATP or other nucleotides (or possibly non-hydrolyzable analogs thereof) in the recombinase-enhanced amplification reagent. A mixture of recombinase and single-stranded binding (SSB) proteins is particularly useful because SSBs can further promote amplification. A typical formulation for promoting recombinase amplification is the TwistAmp kit, commercially available from TwistDx (Cambridge, UK).

[0101] Another example of a component that can be included in amplification reagents to promote amplicon formation and, in some cases, increase the rate of amplicon formation is helicase. Helicase can promote amplicon formation by enabling a chain reaction of amplicon formation. This process can be carried out more rapidly than standard PCR because it does not require a denaturation cycle (e.g., by heating or chemical denaturation). Therefore, helicase-enhanced amplification can be performed isothermally. A mixture of helicase and single-strand binding (SSB) proteins is particularly useful because SSBs can further promote amplification. A typical formulation for helicase-enhanced amplification is the IsoAmp kit commercially available from Biohelix (Beverly, MA).

[0102] Another example of a component that may be included in amplification reagents to promote amplicon formation and, in some cases, increase the rate of amplicon formation is an origin-binding protein.

[0103] The advantage of the methods described herein is that they provide the rapid and efficient parallel detection of multiple target nucleic acids. Accordingly, this disclosure provides an integrated system that allows nucleic acids to be prepared and detected using techniques known in the art, such as those exemplified above. Thus, the integrated system of this disclosure may include a fluid component capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, and temperature controllers. A flow cell may constitute and / or be used in the integrated system for detecting target nucleic acids. As illustrated with respect to the flow cell, one or more fluid components of the integrated system may be used in amplification and detection methods. One or more fluid components of the integrated system may be used for amplification methods described herein and for the delivery of sequencing reagents in sequencing methods, such as those exemplified above. As used herein, the term “flow cell” is intended to mean a container having a chamber (i.e., a flow channel) in which a reaction can be carried out, an inlet for delivering reagents to the chamber, and an outlet for removing reagents from the chamber. In some examples, the chamber allows for the detection of a reaction or signal occurring within the chamber. For example, the chamber may include one or more transparent surfaces that allow for optical detection of arrays, optically labeled molecules, etc., within the chamber. As used herein, “flow channel” or “flow channel region” may be a region provided between two coupled components that can selectively receive a liquid sample. In some examples, the flow channel may be provided between a patterned support and a lid, thereby allowing for fluid connection with one or more recesses provided in the patterned support. In other examples, the flow channel may be provided between an unpatterned support and a lid. Other examples include dishes, plates, or wells for separating reactants, including automated fluidics for the exchange of reagents and other components of the reaction. For example, multi-well plates may be used, such as 96-well plates or 384-well plates.

[0104] Alternatively, the integrated system may include separate fluid systems for performing amplification and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and determining nucleic acid sequences include, but are not limited to, the MiSeq® platform (Illumina Inc., San Diego, CA).

[0105] Non-limiting examples of suitable primers include P5 and / or P7 primers, which are used on the surface of commercially available flow cells sold by Illumina Inc. for sequencing on HISEQ®, HISEQX®, MISEQ®, MISEQDX®, MINISEQ®, NEXTSEQ®, NEXTSEQDX®, NOVASEQ®, GENOME ANALYZER®, ISEQ®, cBot with imaging (icBot), and other instrument platforms. Furthermore, a portion of the template polynucleotide containing a nucleotide sequence corresponding to or complementary to the first or second primer disclosed above may have a P5 primer (containing the nucleotide sequence ATGATACGGCGACCACCGAGATCTACAC), a P7 primer (containing the nucleotide sequence CAAGCAGAAGACGGCATACGAGAT), or both, according to such primer sequences used on the SBS platform or elsewhere described above.

[0106] Examples of substrates include, but are not limited to, substrates used in any of the above-mentioned SBS or other platforms, such as a platform for automated cluster formation and imaging of labeled oligonucleotides hybridized to surface-bound polynucleotides, which does not necessarily have to be a platform equipped to perform the sequencing aspect of the SBS process itself. Such substrates may be flow cells.

[0107] As used herein, the term “recess” refers to an individual concave feature in a patterned support having a surface opening that is completely surrounded by an intermediate region(s) of the patterned support surface. Recesses can take any of a variety of shapes at the surface opening, including, for example, circular, elliptical, square, polygonal, or star-shaped (with any number of vertices). The cross-section of a recess taken perpendicular to the surface can be curved, square, polygonal, hyperbolic, conical, or angular. As an example, a recess may be a well. Also as used herein, “functionalized recess” refers to an individual concave feature to which a primer is bonded, in some examples being bonded to the surface of the recess by a polymer (such as PAZAM or a similar polymer).

[0108] It should be understood that the ranges provided herein include the explicitly stated ranges and any values ​​or subranges within those ranges. For example, the range of approximately 100 nm to approximately 1,000 nm should be interpreted to include not only the explicitly listed limits of approximately 100 nm to approximately 1,000 nm, but also individual values, such as approximately 708 nm, approximately 945.5 nm, etc., and subranges, such as approximately 425 nm to approximately 825 nm, approximately 550 nm to approximately 940 nm, etc. Furthermore, where "approximately" and / or "substantially" are used to express values, they mean that they include slight variations (up to ±10%) from the explicitly stated values. [Examples]

[0109] The following embodiments are intended to illustrate specific embodiments of the present disclosure, but are not intended to limit their scope.

[0110] Example 1. Evaluation of linearity with different polynucleotide sizes (human libraries of 350 bp, 450 bp, and 550 bp, and bacterial libraries of 350 bp and 550 bp) and with different GC content / AT content (bacterial library).

[0111] Methods: Different libraries with varying ratios (population length, GC content / AT content, human library or bacterial library) were used for cluster formation in HiSeqTMX flow cells. The intensity against the relative amount of the DNA library was plotted. Linear fitting was applied using JMP software, and R 2 The following was calculated. Cluster formation and hybridization of fluorescently labeled oligonucleotide probes to the identification sequences were imaged using the icBot platform. Polynucleotides were tagged with one of two complementary identification sequences to a fluorescently labeled oligonucleotide probe labeled with fluorescently labeled Alexa 647 (Probe 1: / 5Alex647N / CT ACA CAT AGA GGC ACA CTC or Probe 2: / 5Alex647N / CT ACA CGT ACT GAC ACA CTC, available from IDT). Solutions containing polynucleotides at the following concentrations from one or another population (or library) were loaded into eight flow cell (FC) lanes.

[0112] [Table 1]

[0113] The gain (40), imaging exposure time (600 ms for probe 1 and 600 ms to 900 ms for probe 2), number of exposures (3), and probe incubation time (6 minutes) were used to image the copies bonded to the surface after cluster formation.

[0114] In this embodiment, fluorescence intensity is used as reading information regarding the cluster amplification level. It is important to ensure that the fluorescence intensity detected in this assay correlates with the cluster amplification level / library load. A good correlation / linear correlation between these two factors establishes the fundamental principle of the assay (and the method for analyzing the data in this assay).

[0115] result 350bp human library: Both probe 1 and probe 2 exhibit similar linearity, and there is a linear relationship of approximately 0.99 between the library input (cluster amplification) and signal intensity. 2 The data from probe 2 is shown in Figure 3. Similar linear relationships were found using the ratios of polynucleotides in high-GC populations (e.g., Rhodobacter), low-GC populations (e.g., Bacillus cereus), or longer (550 bp) populations. This indicates a conversion of starting material concentrations to signal intensity after cluster formation.

[0116] Example 2: Assay sensitivity for determining what percentage of DNA amplification difference can be detected.

[0117] Method: DNA samples with a 10% difference were clustered using a HiSeq(trademark)X flow cell. Signal intensity was measured after hybridization with probe 1 or probe 2.

[0118] result Data from 13 flow cells clustered with 350 bp Rhodobacter DNA are summarized in Figure 4. JMP analysis showed that DNA library loading of 350 bp Rhodobacter libraries at 40%, 50%, and 60% could be separated with statistical significance (95%) by this assay. Similar assay sensitivity was observed using libraries with different loading sizes or GC content, e.g., 350 bp, 450 bp, and 550 bp human libraries, a 550 bp Rhodobacter library, and 350 bp and 550 bp Bacillus cereus libraries.

[0119] It will be understood that all combinations of the aforementioned concepts and further concepts discussed in more detail herein are intended to be part of the subject matter of the inventions disclosed herein (provided that such concepts are not mutually contradictory). In particular, all combinations of the claimed subject matter set out at the end of this disclosure are intended to be part of the subject matter of the inventions disclosed herein and may be used to achieve the benefits and advantages described herein.

Claims

1. A method for evaluating bias in polynucleotide replication, The process of creating copies of two or more populations of polynucleotides, wherein each of the two or more populations of polynucleotides contains an identification sequence unique to that population, and the process of creating the copies includes forming a cluster of copies covalently bound to a substrate, wherein the copies contain the identification sequences and their complements. After the preparation of two or more sets of polynucleotides, hybridize two or more sets of identification sequence oligonucleotide probes. Wash away the identification sequence oligonucleotide probe that does not hybridize to the aforementioned identification sequence, This includes comparing the amount of identification sequence oligonucleotide probes hybridized to the copies of two or more populations of polynucleotides, At least one characteristic differs between the clusters formed in each copy of the two or more populations of polynucleotides bound to the substrate, A method wherein the at least one feature includes at least one of the following: a difference in guanine-cytosine content between at least two of the two or more populations of polynucleotides; a difference in length between at least two of the two or more populations of polynucleotides; and / or a difference in the preparation method between at least two of the two or more populations of polynucleotides.

2. The method according to claim 1, wherein the identification sequence oligonucleotide probe comprises a fluorophore.

3. The method according to claim 1 or 2, further comprising detecting a difference between the amounts of a discriminant sequence oligonucleotide probe hybridized to the copies of two or more populations of polynucleotides bound to the substrate, wherein the difference is at least about 10%.

4. The method according to claim 3, wherein the difference is at least about 20%.

5. The method according to claim 3, wherein the difference is at least about 30%.

6. The method according to claim 1, further comprising detecting a difference between the amounts of identification sequence oligonucleotide probes hybridized to the copies of two or more populations of polynucleotides bound to the substrate, wherein the difference is at least about 10%.

7. The method according to claim 6, wherein the difference is at least about 20%.

8. The method according to claim 6, wherein the difference is at least about 30%.

9. A method for evaluating bias in polynucleotide replication, The process of creating copies of two or more groups of polynucleotides, wherein each of the two or more groups of polynucleotides contains a unique identification sequence, and the process of creating the copies includes forming a cluster of copies covalently bound to a substrate, wherein the copies contain the identification sequences and their complements. After preparing copies of two or more populations of the polynucleotide, a fluorophore-containing identification sequence oligonucleotide probe is hybridized to the identification sequence, wherein the identification sequence oligonucleotide probe of each population contains a sequence or its complement that is complementary to only one of the unique identification sequences. This includes detecting the amount of a discriminant sequence oligonucleotide probe hybridized to the copies of two or more populations of polynucleotides bound to the substrate, At least one characteristic differs between the clusters formed in each copy of the two or more populations of polynucleotides bound to the substrate, A method wherein the at least one feature includes at least one of the following: a difference in guanine-cytosine content between at least two of the two or more groups of polynucleotides; a difference in length between at least two of the two or more groups of polynucleotides; and / or a difference in the preparation method between at least two of the two or more groups of polynucleotides.

10. The method according to claim 9, further comprising detecting a difference between the amounts of a distinguishing sequence oligonucleotide probe hybridized to copies of two or more populations of polynucleotides, wherein the difference is at least about 10%.

11. The method according to claim 10, wherein the difference is at least about 20%.

12. The method according to claim 10, wherein the difference is at least about 30%.

13. The method according to claim 9, wherein the two or more groups of polynucleotides include three or more groups of polynucleotides.

14. The method according to claim 1, wherein the at least one feature comprises at least one of the following: a polymerase used to make the copies; a polymerization reagent used to make the copies; a reagent used to rehybridize the copies during cluster formation; a reagent used to linearize the copies during cluster formation; a polymer for binding the copies to a substrate; or two or more combinations thereof.

15. The method according to claim 9, wherein the at least one feature comprises at least one of the following: a polymerase used to make the copies; a polymerization reagent used to make the copies; a reagent used to rehybridize the copies during cluster formation; a reagent used to linearize the copies during cluster formation; a polymer for binding the copies to a substrate; or two or more combinations thereof.

16. The method according to claim 2, wherein each group of identification sequence oligonucleotide probes contains a fluorophore specific to each group of identification sequence oligonucleotide probes.

17. The method according to claim 9, wherein each group of discriminant sequence oligonucleotide probes comprises a fluorophore specific to each group of discriminant sequence oligonucleotide probes.

18. The method according to claim 1, wherein the two or more groups of identification sequence oligonucleotide probes are simultaneously hybridized to the identification sequence.

Citation Information

Patent Citations

  • Processes and compositions for methylation-based enrichment of fetal nucleic acid from maternal sample useful for non invasive prenatal diagnoses

    JP2015126748A

  • Polynucleotide libraries with controlled stoichiometry and their synthesis

    JP2020504709A

  • Methods and uses for molecular tags

    WO2013130512A2

  • Two-dimensional cell array device and apparatus for gene quantification and sequence analysis

    WO2014141386A1

  • Method of correcting amplification bias in amplicon sequencing

    WO2018170660A1