Analysis of polymers
By using reference sequence data to analyze the measured values of polymers in a nanopore biochemical analysis system, a similarity measure is provided and polymers of no interest are excluded, thus solving the problems of slow analysis speed and insufficient accuracy in the existing technology and achieving faster and more accurate polynucleotide sequence analysis.
Patent Information
- Application Number
- CN202510769251.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2015-05-06
- Filing Date
- 2015-10-16
- Publication Date
- 2025-09-19
AI Technical Summary
Existing nanopore biochemical analysis systems suffer from slow analysis speed and insufficient accuracy when analyzing polymers, especially polynucleotide sequences. It is difficult to effectively utilize the measurements during polymer translocation for early identification and exclusion of polymers of no interest.
The measurements are processed using the model to improve accuracy by analyzing the measurements of the polymers during their partial shift using data from a reference sequence to provide a similarity measure and rejecting polymers that are not of interest based on the similarity measure.
It significantly improves the analysis speed and accuracy, reduces the time consumption of unnecessary measurements, and improves the coverage and accuracy of sequencing, especially in polynucleotide sequence analysis, and can identify polymers of interest at an early stage and exclude polymers of no interest.
Smart Images

Figure CN120666007A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202211448003.2, entitled “Analysis of Polymers”, filed on October 16, 2015. Technical Field
[0002] The first through third aspects of the present invention relate to analyzing polymers using a biochemical analysis system comprising at least one sensor element comprising a nanopore. A fourth aspect of the present invention relates to estimating an alignment mapping between a series of measurements of a polymer comprising polymer units and a reference sequence of polymer units. In all aspects, the polymer can be, for example, but not limited to, a polynucleotide, wherein the polymer units are nucleotides. Background Art
[0003] There are various types of biochemical analysis systems that provide measurements of polymer units for sequence determination. For example, but not exclusively, one type of measurement system utilizes nanopores. Biochemical analysis systems utilizing nanopores have recently seen significant development. Typically, a sensor element comprising the nanopore collects continuous measurements of the polymer as it translocates through the nanopore. Some property of the system depends on the polymer units within the nanopore, and measurements of that property are collected. These types of measurement systems utilizing nanopores hold considerable promise, particularly in the field of sequencing polynucleotides, such as DNA or RNA.
[0004] Such biochemical analysis systems using nanopores can provide long, continuous readouts of polymers, for example, in the case of polynucleotides, ranging from hundreds to tens of thousands (and potentially more) of nucleotides. The data collected in this way include measurements, such as ionic currents, where each shift of a sequence through the sensitive portion of the nanopore results in a slight change in the measured property. Summary of the Invention
[0005] While such biochemical analysis systems using nanopores can provide significant advantages, it is also desirable to increase the speed of analysis. The first and second aspects of the present invention are directed to providing such an increase.
[0006] According to a first aspect of the present invention, there is provided a method of controlling a biochemical analysis system for analyzing a polymer, the polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to acquire continuous measurements of the polymer by the sensor element during translocation of the polymer through the nanopore of the sensor element,
[0007] wherein the method comprises analyzing a series of measurements of the polymer collected during the partial translocation of the polymer as the polymer portion translocates through the nanopore using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of the polymer units of the partially translocated polymer and the at least one reference sequence, and
[0008] Responsive to the measure of similarity, the biochemical analysis system is operated to exclude the polymer and collect measurements of additional polymers.
[0009] This method involves analyzing measurements collected from a polymer as it partially translocates through a nanopore (i.e., during translocation of the polymer through the nanopore). Specifically, a series of measurements collected from the polymer during the partial translocation are analyzed using reference data derived from at least one reference sequence of polymer units. This analysis provides a measure of similarity between the sequence of polymer units of the partially translocated polymer and the at least one reference sequence. In response to this measure of similarity, if the similarity to the reference sequence indicates that further analysis of the polymer is not necessary, for example because the measured polymer is not of interest, a polymer can be excluded in order to collect measurements of additional polymers.
[0010] Excluding polymers allows measurements of additional polymers to be taken without completing the measurement of the initially measured polymer. This provides a time saving in acquiring measurements because the operation is performed "on the fly" (i.e., while the polymer measurement is being taken). In typical applications, this time saving can be significant because biochemical analysis systems using nanopores can provide long, continuous reads of polymers, and analyses can identify at an early stage, in this read, that no further measurements are required for the currently measured polymer.
[0011] For example, in a typical application where the polymer is a polynucleotide, sequencing performed with 100% accuracy will allow for a preliminary determination after measuring approximately 30 nucleotides. Therefore, considering the actual achievable accuracy, determination can be made after measuring several hundred nucleotides, typically 250 nucleotides. This is compared to a biochemical analysis system that can measure sequences ranging in length from several hundred to tens of thousands (and possibly more) nucleotides.
[0012] This method potentially provides significantly faster times for results, wherein only those polymers of interest are measured continuously and those of no interest are excluded. The advantage of this reduction in the amount of discarded data acquisition is particularly significant for applications requiring large amounts of data acquisition. The resulting time savings are useful in themselves, or can be used, for example, to obtain larger coverage, and therefore higher sequencing accuracy can be obtained in addition through available time and resources.
[0013] The analysis of the measure of the similarity between the sequence of the polymer units of the partially shifted polymer and at least one reference sequence can itself use known techniques for comparing measured values with references. However, in contrast to this method, this known technology typically measures after completing the shift.
[0014] The methods can be applied to a wide variety of applications. Depending on the application, the measure of similarity can represent similarity to the entirety of a reference sequence or to a portion of a reference sequence.
[0015] According to a second aspect of the present invention, there is provided a method of controlling a biochemical analysis system for analyzing a polymer, the polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to acquire continuous measurements of the polymer by the sensor element during translocation of the polymer through the nanopore of the sensor element,
[0016] wherein the method comprises analyzing a series of measurements collected from a polymer during partial translocation of the polymer through a nanopore by a derived and fitted metric, the model treating the measurements as observations of a series of k-mer states of different possible types, the model comprising: transition weights for possible transitions between possible types of k-mer states relative to each transition between successive k-mer states in the series of k-mer states; and emission weights representing the probability of observing a measurement of a given k-mer relative to each type of k-mer state, and
[0017] Responsive to the fit measure, the biochemical analysis system is operated to exclude the polymer and collect measurements of additional polymers.
[0018] This method involves analyzing measurements collected from a polymer as it partially translocates through a nanopore, i.e., during the period of translocation of the polymer through the nanopore. In particular, a series of measurements collected from the polymer during the partial translocation are analyzed using reference data derived from at least one reference sequence of polymer units. This analysis provides a measure of fit of the model. In response to this measure of fit, if the measure of fit, as determined by the model, indicates that the measurement is of poor quality, such that additional translocation and measurement are not necessary, action can be taken to exclude the polymer and collect additional measurements of the polymer.
[0019] Excluding polymers allows measurements of additional polymers to be collected without completing the measurement of the initially measured polymer. This provides a time saving in collecting measurements, as the operation is performed "on the fly," i.e., during the time the polymer measurements are being collected. In typical applications, this time saving can be significant, as biochemical analysis systems using nanopores can provide long, continuous reads of polymers, even though the analysis may identify poor quality measurements early on.
[0020] The first and second aspects of the present invention are identical, except that the biochemical analysis system operates as a basis for rejecting the polymer and collecting measurements of the additional polymer. Therefore, the optional features of the first aspect of the present invention may be applied mutatis mutandis to the second aspect of the present invention. Similarly, all of the following method features also apply to the methods of the first or second aspects of the present invention.
[0021] Repelling polymers can occur in different ways.
[0022] In a first approach, at least one sensor element is operable to expel polymer displaced through the nanopore. In this case, the steps of operating the biochemical analysis system to expel the polymer and acquire measurements of the additional polymer can be performed by operating the sensor element to expel the polymer from the nanopore and to receive the additional polymer in the nanopore.
[0023] In a second approach, the biochemical analysis system includes an array of sensor elements and is operated to acquire successive measurements of the polymer using multiplexed selections of the sensor elements. In this case, operating the biochemical analysis system to exclude the polymer and acquire measurements of the additional polymer may include operating the biochemical analysis system to stop acquiring measurements from the currently selected sensor element and begin acquiring measurements from the newly selected sensor element.
[0024] These two approaches can be combined.
[0025] A third aspect of the invention relates to the application of specific forms of biochemical analysis that can be performed using nanopores.
[0026] According to a third aspect of the present invention, there is provided a method of classifying polymers, each of which comprises a sequence of polymer units, the method using a system comprising: a sample chamber comprising a sample containing the polymer, a collection chamber sealed from the sample chamber, and a sensor element comprising a nanopore communicating between the sample chamber and the collection chamber,
[0027] The method comprises causing successive polymers to translocate from a sample chamber through a nanopore, and during the translocation of each polymer:
[0028] Continuous measurements of the polymer are collected by the sensor element;
[0029] analyzing a series of measurements taken from the polymer during a partial shift of the polymer using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of the polymer units of the partially shifted polymer and the at least one reference sequence,
[0030] Based on the measure of similarity, selective displacement of the polymer to the collection chamber or expulsion of the polymer back into the sample chamber is accomplished.
[0031] Therefore, the method utilizes a measure of similarity, which is provided by analyzing a series of measurements collected from the polymer during a partial displacement. This analysis can itself use known techniques for comparing measurements with a reference. However, the measure of similarity is used to determine whether to collect the polymer. If so, the displacement of the polymer into the collection chamber is completed. Alternatively, the polymer is discharged back into the sample chamber. In this way, the selected polymer is collected in the collection chamber. For example, after the displacement of the polymer from the sample is completed, or alternatively, during the displacement of the polymer from the sample, the collected polymer can be recovered, for example, by providing a system (with a fluid system suitable therefor).
[0032] This method can be applied to a variety of applications. For example, the method can be applied to polymers of polynucleotides, such as viral genomes or plasmids. Viral genomes are typically on the order of 10-15 kB (kilobases) in length, and plasmids are typically on the order of 4 kB in length. In such instances, the polynucleotides do not need to be fragmented and can be collected in their entirety. The collected viral genomes or plasmids can be used in any manner, for example, to transfect cells.
[0033] The reference sequence of the polymer units derived from the reference data can be a desired sequence. In this case, in response to a measure of similarity that the partially displaced polymer is a desired sequence, a step of selectively completing polymer displacement into the collecting chamber is performed. However, this is not essential. In some applications, the reference sequence of the polymer units derived from the reference data can be an undesirable sequence. In this case, in response to a measure of similarity that the partially displaced polymer is not an undesirable sequence, a step of selectively completing polymer displacement into the collecting chamber is performed.
[0034] Depending on the application, the measure of similarity can represent similarity to the entirety of a reference sequence or to a portion of a reference sequence.
[0035] The system may include multiple collection chambers, and a sensor element for each collection chamber including a nanopore that provides communication between the sample chamber and the respective collection chamber. This allows the method to be performed with respect to multiple nanopores in parallel. As well as providing the ability to accelerate the sorting method, it can allow different polymers to be collected in different collection chambers. To achieve this, the reference data and standards used for collection are selected accordingly. In one example, the method can be performed using different reference data for different nanopores. In another embodiment, the method can be performed using the same reference data for different nanopores, but the step of selectively completing the displacement of the polymer into the collection chamber is performed with different dependencies on the measure of similarity for different nanopores.
[0036] According to a further aspect of the present invention, there is provided a biochemical analysis system which performs methods similar to those of the first, second or third aspects of the invention.
[0037] A fourth aspect of the invention relates to the alignment of a series of measurements of a polymer comprising polymer units to a reference sequence of polymer units.
[0038] Some types of measurement systems collect measurements of polymers that depend on k-mers, which are k polymer units of a polymer, where k is an integer. By definition, a group of k polymers is referred to as a k-mer hereinafter. In general, k can take the value 1, in which case the k-mer is a single polymer unit, or it can be a plural integer (plural integer). Depending on the nature of the polymer, each given polymer unit can be of a different type. For example, in the case where the polymer is a polynucleotide, the polymer unit is a nucleotide, and the different types are nucleotides containing different nucleic acid bases (such as cytosine, guanine, etc.). Therefore, each given k-mer can also have a different type, corresponding to different combinations of different types of each polymer unit of the k-mer.
[0039] When estimating polymer units from measured values, it's difficult to provide measurement values that depend on a single polymer unit in practical measurement systems. Instead, each measured value depends on a k-mer, where k is a complex number. Conceptually, this can be thought of as a measurement system with a "blunt readhead" that is larger than the polymer units being measured. In this case, the number of different k-mers to be resolved increases to the power of k. When measurements depend on a large number of polymer units (larger k values), it can be difficult to resolve measurements taken from different types of k-mers because they provide overlapping signal distributions, especially when considering noise and / or artifacts in the measurement system. This detracts from estimating the underlying sequence of polymer units.
[0040] When k is a complex number, it is possible to combine information from multiple measurements of overlapping k-mers (each partially dependent on the same polymer unit) to obtain a single value resolved at the polymer unit level. For example, WO-2013 / 041878 discloses a method for estimating the sequence of polymer units in a polymer from at least one series of measurements associated with the polymer using a model that treats the measurements as a series of different possible types of k-mers. The model includes: a transition weight for each transition between consecutive k-mer states in the series of k-mer states, representing a possible transition between possible types of k-mer states; and an emission weight for each type of k-mer state, representing the probability of observing a measurement of a given k-mer. The model can be, for example, a hidden Markov model (HMM). Such a model can improve the accuracy of the estimate by taking multiple measurements into account when considering the likelihood predicted by the model for a series of measurements generated by the sequence of polymer units.
[0041] In many cases, it is desirable to estimate the alignment mapping between a series of measurements of a polymer comprising polymeric units and a reference sequence of polymeric units. The estimation of such alignment mapping can be used in a variety of applications, such as providing identification or detection of the presence, absence or extent of a polymer in a sample by comparison with a reference, for example to provide a diagnosis. The range of possible specific applications is substantial and can be applied to detecting any analyte having a DNA sequence.
[0042] Existing techniques involve initially estimating the sequence of measured polymer units, then estimating their identity by comparing them to a reference sequence of polymer units. A variety of fast alignment algorithms have been developed for applications where the polymer units are nucleotides (often referred to in the literature as bases). Examples of fast alignment algorithms include BLAST (Basic Local Alignment Search Tool), FASTA, and HMMER, and their derivatives. Fast alignment algorithms typically search for small regions of high similarity, a relatively rapid process, before expanding to larger regions of low similarity, a slower process. Such algorithms have been applied in situations where they represent polymer identity by providing a similarity score that indicates whether the measured polymer matches a reference in a minimal timeframe. In these types of techniques, the identity of the polymer units in the estimated and reference sequences is directly compared. When referring to polymer units as bases, the techniques can be considered to involve comparisons of "base intervals," as opposed to comparisons between measured values as "measurement intervals."
[0043] However, this technique has limited accuracy in estimating the alignment map, or in other words, limited discriminative power, because the initial step of estimating the sequence of polymer units inherently causes a loss of information about the identity of the polymer units present in the measurements themselves.
[0044] It would be desirable to provide methods of estimating alignment maps that provide increased accuracy compared to such existing techniques.
[0045] According to a fourth aspect of the invention, there is provided a method for estimating an alignment mapping between: (a) a series of measurements of a polymer comprising polymer units, wherein the measurements depend on a k-mer, a k-mer being k polymer units of the polymer, wherein k is an integer, and (b) a reference sequence of polymer units.
[0046] The method uses a reference model that processes measurements as observations of a reference series of k-mer states corresponding to a reference sequence of polymer units, wherein the reference model comprises:
[0047] transition weights for transitions between k-mer states in a reference series of k-mer states; and
[0048] For each k-mer state, emission weights for different measurement values observed when observing the k-mer state; and
[0049] The method includes applying a reference model to a series of measurements to derive an estimate of an alignment mapping between the series of measurements and a reference series of k-mer states corresponding to a reference sequence of polymer units.
[0050] Therefore, the method uses a reference model with respect to a reference sequence. The reference model processes the measured values as the k-aggregate state of the reference series corresponding to the reference sequence of the polymer unit, and includes a conversion weighting for the conversion between the k-aggregate states in the k-aggregate state of the reference series; and about each k-aggregate state, when observing the k-aggregate state, the emission weighting for the different measured values observed. They can be, but are not limited to, HMMs. As a result, compared to the known techniques discussed above involving the initial estimation of the sequence of polymer units, and then by comparing the consistency of the polymer units to the alignment mapping of the reference sequence of the polymer units, this method can improve the evaluation accuracy of the alignment method. This is due to the following reasons.
[0051] Generally speaking, the purposes of reference model is similar to the model of the sequence of estimation polymer disclosed in WO-2013 / 041878, for example, using conversion weighting and emission weighting of similar form, and identical mathematical treatment is applied to model. However, reference model itself is different from the model disclosed in WO-2013 / 041878, and the model disclosed in WO-2013 / 041878 is the generic model of measurement system, wherein, each kind of k-aggregate state can have any one in the possible type of k-aggregate state in general. Therefore, for the various possible conversions between the possible types of k-aggregate state, about each conversion between the continuous k-aggregate state in a series of k-aggregate states provides conversion weighting. On the contrary, the reference model for present method is the model of the k-aggregate state of the reference series corresponding to the reference sequence of polymer unit. Therefore, conversion weighting is provided, for the conversion between the k-aggregate state in the k-aggregate state of the reference series.
[0052] This similarity means that the method of the present invention can utilize the power of the model disclosed in WO-2013 / 041878. Information about the identity of the polymer units (present in the measurements based on overlapping k-mers) is used to report the resulting product. Due to the different nature of the reference model itself, applying the reference model can provide an alignment mapping between a series of measurements and a reference series of k-mer states corresponding to a reference sequence of polymer units, and therefore provide an alignment mapping between a series of measurements of polymer units and the reference sequence.
[0053] In some implementations, for each measurement in the series, the derived estimate of the alignment map can include a discrete estimate of the k-mer state of the map in the reference series of k-mer states. As an example where the model is an HMM, this can be implemented by using a Viterbi algorithm to derive the estimate of the alignment map.
[0054] In other embodiments, for each measurement in the series, the derived estimate of the alignment map may include a weighting of the k-mer state of the different maps with respect to the k-mer state of the reference series. As an example where the model is an HMM, this can be achieved by deriving the estimate of the alignment map using a forward-backward algorithm.
[0055] Optionally, the method can further comprise deriving a score (indicating the likelihood that the estimated value of the alignment mapping is correct). This score provides a measure of the similarity between the measured polymer and a reference sequence of polymer units. By providing information about the identity of the measured polymer compared to the reference sequence, it can be used in a variety of applications.
[0056] In some cases, the model can be applied directly to derive the score. An example of this is when the model is an HMM and the Viterbi algorithm is applied.
[0057] In other cases, where the derived estimate of the alignment mapping may include weights of the k-mer states of the different mappings with respect to the k-mer states of the reference series for each measurement in the series, the score may be derived from those weights themselves.
[0058] The source of the reference model may vary depending on the application.
[0059] In some applications, a reference model may be pre-stored that was previously generated from a reference sequence of polymer units or measurements taken from a reference sequence of polymers.
[0060] In other applications, a reference model may be generated when performing the method, for example, as follows.
[0061] In a first example, a reference model can be generated from a reference sequence of polymer units. This can be used, for example, for applications where the reference sequence is known from a database or from earlier experiments.
[0062] In this case, the generation of the reference model can be performed using stored emission weights for a set of possible types of k-mer states. Advantageously, this allows the generation of a reference model for any reference sequence of polymer units based solely on stored data relating to emission weights for possible types of k-mer states.
[0063] For example, a reference model can be generated by a process comprising: deriving a series of k-mer states corresponding to a reference sequence of a received polymer; and generating a reference model by generating transition weights for transitions between k-mer states in the series of derived k-mer states, and by selecting an emission weight for each k-mer state in the derived series from stored emission weights based on the type of the k-mer state.
[0064] In a second example, a reference model can be generated from a series of reference measurements of polymers comprising a reference sequence of polymer units. This can be used, for example, in applications where both the reference sequence of polymer units and the target polymer are measured simultaneously. In particular, in this example, the identity of the polymer units of the reference sequence itself is not required to be known.
[0065] For example, a reference model can be generated by a method using an additional model that processes a series of reference measurements as observations of a further series of k-mer states of different possible types, wherein the additional model includes: for each transition between consecutive k-mer states in the further series of k-mer states, a transition weight for possible transitions between possible types of k-mer states; and for each type of k-mer state, when the k-mer state is of that type, an emission weight for different measurements observed. This additional model itself can be a model type disclosed in WO-2013 / 041878. In this case, the reference model can be generated by a process including the following: generating an estimate of a reference series of k-mer states by applying the additional model to a series of reference measurements; and generating the reference model by generating transition weights for transitions between k-mer states in the estimate of the reference series of k-mer states and by selecting an emission weight for each k-mer state in the estimate of the reference series according to the type of the k-mer state by the weight of the further model.
[0066] Model generation can be part of a larger framework of model training, which examines a large collection of reference measurements derived from observing a large collection of k-mer state series to find the unknown parameters of the mathematical model, such as emission and transition weights. Typically, when the model includes latent (hidden) variables, an expectation-maximization (EM) algorithm can be used to find the maximum likelihood estimates. In the specific case of HMMs, the Baum-Welch algorithm can be used. This algorithm is iterative: an initial guess is made for the model parameters, and updates are applied by examining a set of training measurements. Applying the generated HMM to a second, distinct set of measurements will produce improved results (assuming the second set can be described by the same model as the training data).
[0067] According to a further aspect of the present invention, there is provided an electronic computer program capable of implementing the method according to the fourth aspect of the present invention, or an analysis system implementing the method according to the fourth aspect of the present invention.
[0068] The present invention also relates to the following embodiments:
[0069] 1. A method of controlling a biochemical analysis system for analyzing a polymer, the polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to collect continuous measurements of the polymer from the sensor element during translocation of the polymer through the nanopore of the sensor element,
[0070] wherein the method comprises analyzing a series of measurements taken from the polymer during partial translocation of the polymer using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of polymer units of the partially translocated polymer and the at least one reference sequence, and
[0071] Responsive to the measure of similarity, the biochemical analysis system is operated to exclude the polymer and collect measurements from additional polymers.
[0072] 2. A method according to embodiment 1, wherein at least one of the sensor elements is operable to expel a polymer that is displacing through the nanopore, and the step of operating the biochemical analysis system to exclude the polymer and collect measurement values from additional polymers includes operating the sensor element to expel the polymer from the nanopore and receive additional polymers in the nanopore.
[0073] 3. A method according to embodiment 2, wherein at least one of the sensor elements is operable to expel a polymer that is being displaced through the nanopore by applying an expulsion bias sufficient to expel the polymer, the step of operating the sensor element to expel the polymer from the nanopore is performed by applying an expulsion bias, and the step of operating the sensor element to receive additional polymer in the nanopore is performed by applying a displacement bias sufficient to enable the additional polymer to be displaced therethrough.
[0074] 4. A method according to embodiment 1, wherein the biochemical analysis system includes an array of sensor elements and is operable to collect continuous measurement values of polymers from selected sensor elements in a multiplexed manner, and the step of operating the biochemical analysis system to exclude the polymer and collect measurement values from another polymer includes operating the biochemical analysis system to stop collecting measurement values from the currently selected sensor element and start collecting measurement values from a newly selected sensor element.
[0075] 5. The method of embodiment 4, wherein the measurements comprise electrical measurements collected from the sensor elements, and the biochemical analysis system is operable to collect continuous measurements of the polymer from selected sensor elements in an electrically multiplexed manner.
[0076] 6. The method according to embodiment 5, wherein the biochemical analysis system comprises:
[0077] a detection circuit comprising a plurality of detection channels, each capable of acquiring electrical measurements from the sensor elements, the number of sensor elements in the array being greater than the number of detection channels; and
[0078] A switch arrangement is provided to selectively connect the detection channels to the individual sensor elements in a multiplexed manner.
[0079] 7. A method according to any one of embodiments 4 to 7, wherein the sensor element can be controlled to expel the polymer that is displacing through the nanopore of the sensor element, and the method further includes, when operating the biochemical analysis system to stop collecting measurement values from the currently selected sensor element, also controlling the currently selected sensor element to expel the polymer, and thereby making the nanopore available to receive additional polymer.
[0080] 8. The method of any of the preceding embodiments, wherein the at least one reference sequence of polymer units is an undesirable sequence, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively operating is performed in response to a measure of similarity that indicates that the partially shifted polymer is an undesirable sequence.
[0081] 9. The method of any one of embodiments 1 to 7, wherein the at least one reference sequence of polymer units is a target, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively operating is performed in response to a measure of similarity that indicates that the partially displaced polymer is not a target.
[0082] 10. The method of any one of embodiments 1 to 7, wherein the at least one reference sequence of polymer units is a measured sequence of polymer units, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively operating is performed in response to a measure of similarity that indicates that the partially shifted polymer is a measured sequence of polymer units.
[0083] 11. A method according to any one of embodiments 1 to 7, wherein the at least one reference sequence of polymer units includes multiple targets, the reference data is derived from the at least one reference sequence of polymer units, and the step of selective operation is performed in response to a measure of similarity, wherein the measure of similarity indicates that the partially displaced polymer is one of the targets.
[0084] 12. The method according to any one of the preceding embodiments, wherein:
[0085] said reference data derived from at least one reference sequence of polymer units represent actual or simulated measurements collected from a biochemical analysis system, and
[0086] The step of analyzing a series of the measurements of the polymer taken during the partial displacement comprises:
[0087] A series of the measured values are compared with the reference data.
[0088] 13. The method according to any one of embodiments 1 to 11, wherein the reference data derived from at least one reference sequence of polymer units represents a feature vector of a temporal characteristic, the feature vector of the temporal characteristic representing a property of the measurement value collected by a biochemical analysis system, and
[0089] The step of analyzing a series of the measurements of the polymer taken during the partial displacement comprises:
[0090] deriving a feature vector representing a time series characteristic of a characteristic of the measurement value from a series of the measurement values, and
[0091] The derived feature vector is compared with the reference data.
[0092] 14. The method according to any one of embodiments 1 to 11, wherein
[0093] said reference data derived from at least one reference sequence of polymer units represents the identity of said polymer units of said at least one reference sequence, and
[0094] Said step of analyzing a series of said measurements from said polymer taken during said partial displacement comprises:
[0095] analyzing a series of said measurements to provide an estimate of the identity of said polymer units of the sequence of polymer units of the partially shifted polymer, and
[0096] The estimated value is compared to the reference data to provide a measure of the similarity.
[0097] 15. The method according to any one of embodiments 1 to 11, wherein
[0098] The measured value depends on a k-mer, wherein a k-mer is k polymer units of a polymer, where k is an integer;
[0099] The reference data represents a reference model that processes measurements as observations of a reference series of k-mer states corresponding to the reference sequence of polymer units, wherein the reference model comprises:
[0100] transition weights for transitions between said k-mer states in said reference series of k-mer states; and
[0101] For each k-mer state, the emission weights for the different measurements observed when observing that k-mer state, and
[0102] The step of analysing the series of measurements taken from the polymer during the partial displacement comprises fitting the model to the series of measurements to provide a measure of similarity as a fit of the model to the series of measurements.
[0103] 16. The method according to any one of the preceding embodiments, wherein the measured value depends on a k-mer, wherein a k-mer is k polymer units of a polymer, where k is an integer.
[0104] 17. The method according to any one of the preceding embodiments, wherein the nanopore is a biological pore.
[0105] 18. The method of any one of the preceding embodiments, wherein the polymer is a polynucleotide and the polymer units are nucleotides.
[0106] 19. The method according to any one of the preceding embodiments, wherein the translocation of the polymer through the nanopore occurs in a ratchet manner.
[0107] 20. The method of any preceding embodiment, wherein the measurements comprise electrical measurements.
[0108] 21. A biochemical analysis system for analyzing a polymer, the polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to collect continuous measurements of the polymer from the sensor element during translocation of the polymer through the nanopore of the sensor element,
[0109] wherein the biochemical analysis system is arranged to analyse a series of measurements of the polymer taken during the partial translocation of the polymer, when the polymer has been partially translocated through the nanopore, using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of polymer units of the partially translocated polymer and the at least one reference sequence, and
[0110] The biochemical analysis system is arranged to reject the polymer and collect measurements from further polymers in response to the measure of similarity.
[0111] 22. A method of controlling a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to collect continuous measurements of the polymer from the sensor element during translocation of the polymer through the nanopore of the sensor element,
[0112] wherein the method comprises analyzing a series of said measurements collected from the polymer during the partial translocation of the polymer through the nanopore by deriving a metric fitted to a model, said model treating said measurements as observations of a series of k-mer states of different possible types and comprising: a transition weight, for each transition between successive k-mer states in said series of k-mer states, for a possible transition between possible types of k-mer states; and an emission weight, for each type of k-mer state, representing the probability of observing a given measurement of said k-mer, and
[0113] Responsive to the measure of the fit, the biochemical analysis system is operated to exclude the polymer and acquire measurements from additional polymers.
[0114] 23. A biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to collect continuous measurements of the polymer from the sensor element during translocation of the polymer through the nanopore of the sensor element,
[0115] wherein the biochemical analysis system is arranged to analyse a series of said measurements collected from the polymer during the partial translocation of said polymer through said nanopore by deriving a metric fitted to a model, said model treating said measurements as observations of a series of k-mer states of different possible types and comprising: a transition weight, for each transition between successive k-mer states in said series of k-mer states, for possible transitions between said possible types of k-mer states; and an emission weight, for each type of k-mer state, representing the probability of observing a given measurement of said k-mer, and
[0116] In response to the measure of fit, the biochemical analysis system is arranged to exclude the polymer and collect measurements from further polymers.
[0117] 24. A method of classifying polymers, each of the polymers comprising a sequence of polymer units, the method using a system comprising: a sample chamber comprising a sample containing a polymer, a collection chamber isolated from the sample chamber, and a sensor element comprising a nanopore, the nanopore communicating between the sample chamber and the collection chamber,
[0118] The method comprises causing successive polymers to translocate from the sample chamber through the nanopore, and during the translocation of each polymer:
[0119] collecting continuous measurements of the polymer from the sensor element;
[0120] analyzing a series of said measurements taken from said polymer during a partial translocation of said polymer using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of said polymer units of said partially translocated polymer and said at least one reference sequence,
[0121] Based on the measure of similarity, selectively translocation of the polymer to the collection chamber is accomplished, or expulsion of the polymer back into the sample chamber is accomplished.
[0122] 25. A method according to embodiment 24, wherein the system includes multiple collection chambers and, relative to each collection chamber, includes a sensor element containing a nanopore, wherein the nanopore provides communication between the sample chamber and each collection chamber, and the method is performed with respect to multiple sensor elements in parallel.
[0123] 26. The method of embodiment 25, wherein the method is performed using different reference data for different nanopores.
[0124] 27. A method according to embodiment 25, wherein the method is performed using the same reference data regarding different nanopores, and the step of selectively completing the shifting of the polymer to the collection chamber or additionally expelling the polymer back to the sample chamber is performed using different dependencies on the measure of the similarity regarding different nanopores.
[0125] 28. A method according to any one of embodiments 24 to 27, wherein the step of causing a continuous polymer to shift from the sample chamber through the nanopore includes applying a bias sufficient to initiate the shift, and the step of expelling the polymer back to the sample chamber includes applying an expulsion bias sufficient to expel the polymer.
[0126] 29. A method according to any one of embodiments 24 to 28, wherein, after the step of analyzing the series of measurements collected from the polymer during the partial displacement to provide a measure of similarity, no further analysis of the measurements is performed when the step of completing the displacement of the polymer to the collection chamber is performed.
[0127] 30. A method according to any one of embodiments 24 to 29, wherein, after the step of analyzing the series of measurements collected from the polymer during the partial shifting to provide a measure of similarity, the shifting is performed at an increased rate while the step of completing the shifting of the polymer to the collection chamber is performed.
[0128] 31. A method according to any one of embodiments 24 to 30, wherein the at least one reference sequence of the polymer unit is a desired sequence, the reference data is derived from the at least one reference sequence of the polymer unit, and the step of selectively completing the shift of the polymer to the collection chamber is performed in response to a measure of similarity, wherein the measure of similarity indicates that the partially shifted polymer is the desired sequence.
[0129] 32. The method according to any one of embodiments 24 to 31, wherein
[0130] said reference data derived from at least one reference sequence of polymer units represent actual or simulated measurements collected by a biochemical analysis system, and
[0131] Said step of analyzing a series of said measurements of said polymer taken during a partial displacement comprises:
[0132] A series of the measured values are compared with the reference data.
[0133] 33. The method of any one of embodiments 24 to 31, wherein the reference data derived from at least one reference sequence of polymer units represents a feature vector of temporal characteristics, the feature vector of temporal characteristics representing characteristics of the measurements collected by a biochemical analysis system, and
[0134] Said step of analyzing a series of said measurements taken from said polymer during a partial displacement comprises:
[0135] deriving a feature vector representing a time series characteristic of a characteristic of the measurement value from a series of the measurement values, and
[0136] The derived feature vector is compared with the reference data.
[0137] 34. The method according to any one of embodiments 24 to 31, wherein
[0138] said reference data derived from at least one reference sequence of polymer units represents the identity of said polymer units of said at least one reference sequence, and
[0139] Said step of analyzing a series of said measurements taken from said polymer during a partial displacement comprises:
[0140] analyzing a series of said measurements to provide an estimate of the identity of said polymer units of a sequence of polymer units of said polymer that has been partially shifted, and
[0141] The estimated value is compared to the reference data to provide a measure of the similarity.
[0142] 35. The method according to any one of embodiments 24 to 31, wherein
[0143] The measured value depends on a k-mer, wherein a k-mer is k polymer units of a polymer, where k is an integer;
[0144] The reference data represents a reference model that processes measurements as observations of a reference series of k-mer states corresponding to the reference sequence of polymer units, wherein the reference model comprises:
[0145] transition weights for transitions between said k-mer states in said reference series of k-mer states; and
[0146] For each k-mer state, the emission weights for the different measurements observed when observing that k-mer state, and
[0147] The step of analysing the series of measurements taken from the polymer during the partial displacement comprises fitting the model to the series of measurements to provide the measure of similarity as a fit of the model to the series of measurements.
[0148] 36. The method of any one of embodiments 24 to 35, wherein the measured value is dependent on a k-mer, wherein a k-mer is k polymer units of a polymer, where k is an integer.
[0149] 37. The method of any one of embodiments 24 to 36, wherein the nanopore is a biological pore.
[0150] 38. The method of any one of embodiments 24 to 37, wherein the polymer is a polynucleotide and the polymer units are nucleotides.
[0151] 39. The method of any one of embodiments 24 to 38, wherein translocation of the polymer through the nanopore occurs in a ratchet manner.
[0152] 40. The method of any one of embodiments 24 to 39, wherein the measurements comprise electrical measurements.
[0153] 41. A system for classifying polymers, each of which comprises a sequence of polymer units, the system comprising:
[0154] a sample chamber for containing a sample, said sample comprising said polymer;
[0155] a collection chamber isolated from the sample chamber; and
[0156] a sensor element comprising a nanopore communicating between the sample chamber and the collection chamber,
[0157] wherein the system is arranged to cause successive polymers to translocate from the sample chamber through the nanopore, and during translocation of each polymer:
[0158] The system is arranged to collect continuous measurements of the polymer from the sensor element;
[0159] The system is arranged to analyse a series of measurements taken from a polymer during a partial displacement of the polymer using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of the polymer units of the partially displaced polymer and the at least one reference sequence, and depending on the measure of similarity the system is arranged to selectively complete the displacement of the polymer to the collection chamber or to otherwise expel the polymer back into the sample chamber.
[0160] 42. The system of embodiment 41 further comprises a sensor electrode disposed in each collection chamber, wherein the collection chamber is detachable from the electrode.
[0161] 43. A method of estimating an alignment mapping between: (a) a series of measurements of a polymer comprising polymer units, wherein the measurements depend on a k-mer, wherein the k-mer is k polymer units of the polymer, where k is an integer, and (b) a reference sequence of polymer units;
[0162] The method uses a reference model that processes measurements as observations of a reference series of k-mer states corresponding to the reference sequence of polymer units, wherein the reference model comprises: transition weights for transitions between the k-mer states in the reference series of k-mer states;
[0163] for each k-mer state, the emission weights for the different measurement values observed when observing the k-mer state; and
[0164] The method comprises applying the reference model to a series of the measurements to derive an estimate of an alignment mapping between the series of the measurements and the reference series of k-mer states corresponding to the reference sequence of polymer units.
[0165] 44. An analysis system for estimating an alignment mapping between: (a) a series of measurements of a polymer comprising polymer units, wherein the measurements depend on a k-mer, the k-mer being k polymer units of the polymer, where k is an integer, and (b) a reference sequence of polymer units;
[0166] The analysis system comprises an analysis unit arranged to use a reference model, the reference model processing the measurement values as observations of a reference series of k-mer states corresponding to the reference sequence of polymer units, wherein the reference model comprises: transition weights for transitions between the k-mer states in the reference series of k-mer states; and, for each k-mer state, emission weights for different measurement values observed when observing the k-mer state; and the analysis unit is arranged to perform the following steps:
[0167] The method comprises applying said reference model to a series of said measurements to derive an estimate of an alignment mapping between a series of said measurements and said reference series of k-mer states corresponding to said reference sequence of polymer units. BRIEF DESCRIPTION OF THE DRAWINGS
[0168] For a better understanding, embodiments of the present invention will now be described by way of non-limiting examples with reference to the accompanying drawings, in which:
[0169] Figure 1 is a schematic diagram of the biochemical analysis system;
[0170] Figure 2 is a cross-sectional view of the sensor equipment of the system;
[0171] Figure 3 is a schematic diagram of a sensor element of a sensor device;
[0172] Figure 4 is a graph of a signal of an event measured over time by a measurement system;
[0173] Figure 5 is a block diagram of the electronic circuitry of the system in a first arrangement;
[0174] Figure 6is a block diagram of the electronic circuitry of the system in the second arrangement;
[0175] Figure 7 is a flow chart of a method for controlling a biochemical analysis system to analyze a polymer;
[0176] Figure 8 It is a flow chart of the state detection steps;
[0177] Figure 9 It is a detailed flow chart of an instance of the state detection step;
[0178] Figure 10 is a graph of a sequence of raw measurements and a sequence of obtained measurements that undergo a state detection step;
[0179] Figure 11 is a flow chart of an alternative method for controlling a biochemical analysis system for analyzing polymers;
[0180] Figure 12 is a flow chart of a method of controlling a biochemical analysis system to classify polymers;
[0181] Figures 13 to 16 is a flow chart of different methods used to analyze different forms of reference data;
[0182] Figure 17 is a state diagram of an instance of the k-mer state of the reference series;
[0183] Figure 18 is a state diagram of a reference series of k-mer states illustrating possible types of transitions between k-mer states;
[0184] Figure 19 is a flow chart of a first process for generating a reference model;
[0185] Figure 20 is a flow chart of a second process for generating a reference model; and
[0186] Figure 21 is a flow chart of a method for estimating an alignment mapping; and
[0187] Figure 22 is a block diagram of the alignment map. DETAILED DESCRIPTION
[0188] A variety of nucleotide and amino acid sequences can be used in the described embodiments. In particular:
[0189] SEQ ID NO: 1 is a nucleotide sequence encoding pore MS-(B1)8 (=MS-(D90N / D91N / D93N / D118R / D134R / E139K)8);
[0190] SEQ ID NO: 2 is the amino acid sequence encoding pore MS-(B1)8 (=MS-(D90N / D91N / D93N / D118R / D134R / E139K)8);
[0191] SEQ ID NO: 3 is a nucleotide sequence encoding pore MS-(B2)8 (=MS-(L88N / D90N / D91N / D93N / D118R / D134R / E139K)8);
[0192] SEQ ID NO: 4 is an amino acid sequence encoding pore MS-(B2)8 (=MS-(L88N / D90N / D91N / D93N / D118R / D134R / E139K)8). The amino acid sequence of B2 is identical to that of B1, except for the mutation L88N.
[0193] SEQ ID NO: 5 is the sequence for wild-type Escherichia coli exonuclease I (WT EcoExo I), a preferred polynucleotide handling enzyme;
[0194] SEQ ID NO: 6 is the sequence for E. coli exonuclease III, preferably a polynucleotide processing enzyme;
[0195] SEQ ID NO: 7 is the sequence for Thermomyces truncatula RecJ, preferably a polynucleotide processing enzyme;
[0196] SEQ ID NO: 8 is the sequence for a bacteriophage lambda exonuclease, preferably a polynucleotide handling enzyme; and
[0197] SEQ ID NO: 9 is the sequence for Phi29 DNA polymerase, preferably a polynucleotide handling enzyme.
[0198] The various features described below are examples rather than limitations. Likewise, the features described do not have to be applied together and can be applied in any combination.
[0199] The nature of the polymers to which the present invention can be applied is first described.
[0200] A polymer consists of a sequence of polymer units. Depending on the properties of the polymer, each given polymer unit can be of a different type (or identity).
[0201] A polymer can be a polynucleotide (or nucleic acid), a polypeptide such as a protein, a polysaccharide, or any other polymer. A polymer can be natural or synthetic. The polymer units can be nucleotides. Nucleotides can be of different types containing different nucleic acid bases.
[0202] A polynucleotide can be deoxyribonucleic acid (DNA), ribonucleic acid (RNA), cDNA, or a synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains. A polynucleotide can be single-stranded, double-stranded, or contain both single-stranded and double-stranded regions. Typically, cDNA, RNA, GNA, TNA, or LNA is single-stranded.
[0203] Nucleotides can be of any type. Nucleotides can be naturally occurring or artificial. Nucleotides typically comprise a nucleic acid base (which may be referred to herein simply as a "base"), a sugar, and at least one phosphate group. Nucleobases are typically heterocyclic. Suitable nucleic acid bases include purines and pyrimidines, and more specifically adenine, guanine, thymine, uracil, and cytosine. The sugar is typically a pentose. Suitable sugars include, but are not limited to, ribose and deoxyribose. Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically comprise monophosphates, diphosphates, or triphosphates.
[0204] Nucleotides can include damaged bases or epigenetic bases. Nucleotides can be labeled or modified to act as markers with a distinct signal. This technology can be used to identify absent bases, for example, abasic units or gaps in a polynucleotide.
[0205] Of particular use when considering measurements of modified or damaged DNA (or similar systems) are methods in which complementary data are taken into account. The additional information provided enables discrimination between a larger number of basis states.
[0206] The polymer may also be a type of polymer other than a polynucleotide, some non-limiting examples of which are listed below.
[0207] The polymer may be a polypeptide, in which case the polymer units may be naturally occurring or synthetic amino acids.
[0208] The polymer may be a polysaccharide, in which case the polymer units may be monosaccharides.
[0209] In particular, when the biochemical analysis system 1 comprises a nanopore and the polymer comprises a polynucleotide, the polynucleotide may be long, for example at least 5 kB (kilobases), ie at least 5,000 nucleotides, or at least 30 kB (kilobases), ie at least 30,000 nucleotides.
[0210] As used herein, the term 'k-mer' refers to a group of k polymer units, where k is a positive integer, including the case where k is 1, in which case the k-mer is a single polymer unit. In some cases, reference is made to a k-mer (where k is plural), which is a subset of k-mers, generally excluding the case where k is 1.
[0211] Thus, each given k-mer may also be of different types corresponding to different combinations of different types of each polymer unit of the k-mer.
[0212] Figure 1 A biochemical analysis system 1 for analyzing polymers is shown, which can also be used to classify polymers. Figure 1 , the biochemical analysis system 1 comprises a sensor device 2 connected to an electronic circuit 4 , which in turn is connected to a data processor 6 .
[0213] Some examples will first be described in which the sensor device 2 comprises an array of sensor elements each comprising a biological nanopore.
[0214] In a first form, the sensor device 2 may have a Figure 2 , which comprises a body 20 in which an array of wells 21 are formed, each of which is a recess having a sensor electrode 22 disposed therein. A large number of wells 21 are provided to optimize the data collection rate of the system 1. In general, there may be any number of wells 21, typically 256 or 1024, but in Figure 2 Only a few of the recesses 21 are shown. The body 20 is covered by a cover 23 that extends over the body 20 and is hollow to define a sample chamber 24 into which each recess 21 opens. A common electrode 25 is disposed within the sample chamber 24. In this first form, the sensor device 2 may be a device described in further detail in WO-2009 / 077734, the teachings of which may be applied to the biochemical analysis system 1 and which is incorporated herein by reference.
[0215] In a second form, the sensor device 2 may have the construction described in detail in WO-2014 / 064443, the teachings of which may be applied to the biochemical analysis system 1 and incorporated herein by reference. In this second form, the sensor device 2 has a generally similar construction to the first form, including an array of compartments generally similar to the recess 21, but having a more complex construction and each including a sensor electrode 22.
[0216] To facilitate collection of sample from the collection chamber, the sensor device may be arranged so that the collection chamber 21 can be removed from the underlying respective electrodes 22 to expose the sample contained therein. This device configuration is described in more detail in UK Patent Application No. 1418512.8.
[0217] The sensor device 2 is prepared to form an array of sensor elements 30, Figure 3 One of these is schematically shown in FIG. Each sensor element 30 is fabricated by forming a membrane 31 across each recess 21 in the first form of the sensor device 2 or across each compartment in the second form of the sensor device 2, and then embedding a hole 32 into the membrane 31. The membrane 31 seals each recess 21 from the sample chamber 24. The membrane 31 can be made of an amphiphilic molecule, such as a lipid.
[0218] The pore 32 is a biological nanopore and connects the sample chamber 24 and the recess 21 in a known manner.
[0219] Such preparation may be carried out using the techniques and materials described in detail in WO-2009 / 077734 for the first form of sensor device 2 or using the techniques and materials described in detail in WO-2009 / 077734 for the second form of sensor device 2 .
[0220] Each sensor element 30 is operable to acquire electrical measurements of the polymer 33 during its displacement through the aperture 32 using the sensor electrodes 22 and the common electrode 25 for each sensor element 30. The displacement of the polymer 33 through the aperture 32 generates a characteristic signal of a measured characteristic that can be observed and generally referred to as an "event."
[0221] In this example, the pore is a biological pore, which may have the following properties.
[0222] The biological pore can be a transmembrane protein pore. The transmembrane protein pore used in the methods described herein can be derived from a β-barrel pore or an α-helical bundle pore. A β-barrel pore comprises a barrel or channel formed by β strands. Suitable β-barrel pores include, but are not limited to, α-toxins such as α-hemolysin, anthrax toxin, and leukocidin, as well as bacterial outer membrane proteins / porins such as Mycobacterium smegmatis porins (Msp), such as MspA, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP). An α-helical bundle pore comprises a barrel or channel formed by α-helices. Suitable α-helical bundle pores include, but are not limited to, inner and outer membrane proteins such as WZA and ClyA toxins. The transmembrane pore can be derived from Msp or from α-hemolysin (α-HL).
[0223] Suitable transmembrane protein pores can be derived from Msp, preferably from MspA. Such pores are oligomeric and typically comprise 7, 8, 9, or 10 monomers derived from Msp. The pore can be a homo-oligomeric pore derived from Msp comprising identical monomers. Alternatively, the pore can be a hetero-oligomeric pore derived from Msp comprising at least one monomer that is different from the other monomers. The pore can also comprise one or more constructs comprising two or more covalently linked monomers derived from Msp. Suitable pores are described in WO-2012 / 107778. The pore can be derived from MspA or a homologue or paralog thereof.
[0224] Biological pores can be naturally occurring pores or can be mutant pores. Typical pores are described in Stoddart D et al., Proc Natl Acad Sci, 12;106(19):7702-7, Stoddart D et al., Angew Chem Int Ed Engl. 2010;49(3):556-9, Stoddart D et al., Nano Lett. 2010 Sep 8;10(9):3633-7, Butler TZ et al., Proc Natl Acad Sci 2008;105(52):20647-52, and WO-2012 / 107778.
[0225] The biological pore may be MS-(B1) 8. The nucleotide sequences encoding B1 and the amino acid sequence of B1 are Seq ID: 1 and Seq ID: 2.
[0226] More preferably, the biological pore is MS-(B2)8. The amino acid sequence of B2 is identical to that of B1 except for the mutation L88N. The nucleotide sequence encoding B2 and the amino acid sequence of B2 are Seq ID: 3 and Seq ID: 4.
[0227] The biological pore can be embedded in a membrane, such as an amphiphilic layer, for example, a lipid bilayer. An amphiphilic layer is formed from amphiphilic molecules, such as phospholipids, that possess both hydrophilic and lipophilic properties. The amphiphilic layer can be a monolayer or a bilayer. The amphiphilic layer can be a co-block polymer, such as that disclosed by (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450) or by PCT / GB2013 / 052767, published as WO2014 / 064444. Alternatively, the biological pore can be embedded in a solid-state layer.
[0228] The pore 32 is an example of a nanopore. More generally, the sensor device 2 may be of any form comprising at least one sensor element 30 operable to acquire measurements of the polymer during its translocation through the nanopore.
[0229] A nanopore is typically a hole with nanometer-scale dimensions that allows a polymer to pass through it. Properties dependent on polymer units that migrate through the pore can be measured. The properties can be related to the interaction between the polymer and the nanopore. The polymer interaction can occur within the constricted region of the nanopore. The biochemical analysis system 1 measures the properties, generating a measurement value dependent on the polymer units of the polymer.
[0230] Alternatively, the nanopore may be a solid-state pore comprising a hole formed in a solid-state layer. In this case, it may have the following characteristics.
[0231] The solid-state layer can be formed by organic and inorganic materials, including, but not limited to, microelectronic materials, insulating materials, such as Si3N4, Al2O3 and SiO, organic and inorganic polymers such as polyamides, plastics, such as Teflon® or elastomers, such as two-component addition-cured silicone rubber, and glass. The solid-state layer can be formed by Graphene. Suitable Graphene is disclosed in WO-2009 / 035647 and WO-2011 / 046706.
[0232] When the solid-state pore is a hole in the solid-state layer, the hole may be chemically or otherwise modified to enhance its properties as a nanopore.
[0233] Solid-state pores can be used with additional components that provide alternative or additional measurements of the polymer, such as tunnel electrodes (Ivanov AP et al., Nano Lett. 2011 Jan 12;11(1):279-85), or field effect transistor (FET) devices (WO-2005 / 124888). Solid-state pores can be formed by known methods, including, for example, those described in WO 00 / 79257.
[0234] In such Figure 1In the example of the biochemical analysis system 1 shown, the measurements are electrical measurements, in particular current measurements of the ionic current flowing through the pore 32. In general, these and other electrical measurements can be performed using standard single-channel recording devices such as those described in Stoddart Det al., Proc Natl Acad Sci, 12;106(19):7702-7, Lieberman KR et al, J Am Chem Soc. 2010;132(50):17961-72 and WO-2000 / 28312. Alternatively, electrical measurements can be performed using multichannel systems such as those described in WO-2009 / 077734 and WO-2011 / 067559.
[0235] In order to allow the collection of measurements as the polymer shifts through the nanopore 32, the displacement rate can be controlled by a polymer binding moiety. Typically, the moiety can cause the polymer to shift through the hole 32 with the aid of or in response to an applied field. The moiety can be a molecular motor, which uses, for example, enzymatic activity in the case where the moiety is an enzyme, or can act as a molecular brake. In the case where the polymer is a polynucleotide, a variety of methods have been proposed to control the displacement rate, including the use of polynucleotide binding enzymes. Suitable enzymes for controlling the displacement rate of polynucleotides include, but are not limited to, polymerases, helicases, exonucleases, single-stranded and double-stranded binding proteins and topoisomerases, such as gyrases. For other polymer types, moieties that interact with the polymer type can be used. The polymer interaction moiety may be any of those disclosed in WO-2010 / 086603, WO-2012 / 107778 and Lieberman KR et al, J Am Chem Soc. 2010;132(50):17961-72) and any of those disclosed for voltage gating schemes (Luan B et al., Phys Rev Lett. 2010;104(23):238103).
[0236] The polymer-binding moiety can be used in a variety of ways to control polymer motion. The moiety can cause the polymer to translocate through the pore 32 using or in response to an applied field. The moiety can function as a molecular motor, for example, for enzymatic activity in the case where the moiety is an enzyme, or as a molecular brake. The translocation of the polymer can be controlled by a molecular ratchet that controls the translocation of the polymer through the pore. The molecular ratchet can be a polymer-binding protein.
[0237] For polynucleotide, polynucleotide binding protein is preferably a polynucleotide treatment enzyme. Polynucleotide treatment enzyme is a polypeptide that can interact with polynucleotide and modify at least one characteristic of polynucleotide. Enzyme can modify polynucleotide by cutting it to form a shorter chain of single nucleotide or nucleotide, such as di- or trinucleotide. Enzyme can modify polynucleotide by directing it or shifting it to a specific position. Polynucleotide treatment enzyme does not need to show enzymatic activity, as long as it can bind target polynucleotide and control its displacement by hole. For example, can modified enzyme to remove its enzymatic activity, or, can use under the condition of preventing it from acting as enzyme. Such conditions are discussed in more detail below.
[0238] The polynucleotide handling enzyme may be derived from a nucleolytic enzyme. The polynucleotide handling enzyme used to construct the enzyme is more preferably derived from a member of any one of the enzyme classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31. The enzyme may be any of those disclosed in WO-2010 / 086603.
[0239] Preferred enzymes are polymerases, exonucleases, helicases, and topoisomerases, such as gyrase. Suitable enzymes include, but are not limited to, exonuclease I from Escherichia coli (Seq ID: 5), exonuclease III from Escherichia coli (Seq ID: 6), RecJ from T. thermophilus (Seq ID: 7), and bacteriophage lambda exonuclease (Seq ID: 8), and variants thereof. Three subunits comprising the sequence shown in Seq ID: 8, or variants thereof, interact to form a trimeric exonuclease. The enzyme is preferably derived from Phi29 DNA polymerase. Enzymes derived from Phi29 polymerase include the sequences shown in Seq ID: 9, or variants thereof.
[0240] The variant of Seq IDs: 5, 6, 7, 8 or 9 is an enzyme having an amino acid sequence that varies from the amino acid sequence in Seq IDs: 5, 6, 7, 8 or 9 and maintains polynucleotide binding ability. The variant may include modifications that promote polynucleotide binding and / or promote its activity at high salt concentrations and / or room temperature.
[0241] For the entire length of the amino acid sequence of Seq IDs: 5, 6, 7, 8 or 9, the variant will preferably be at least 50% homologous to the above sequence based on amino acid identity. More preferably, for the entire sequence, the variant polypeptide may be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% and more preferably at least 95%, 97% or 99% homologous to the amino acid sequence of Seq IDs: 5, 6, 7, 8 or 9 based on amino acid identity. For a stretch of 200 or more, for example 230, 250, 270 or 280 or more contiguous amino acids, there may be at least 80%, for example at least 85%, 90% or 95% amino acid identity ("hard homology"). Homology is determined as described above. The variant may differ from the wild-type sequence in any of the ways discussed above with reference to SEQ ID NO: 2. As discussed above, the enzyme may be covalently attached to the pore.
[0242] A suitable strategy for single-stranded DNA sequencing is to shift the DNA through the pore 32 in a cis-to-trans and trans-to-cis manner using or in response to an applied potential. The most advantageous mechanism for strand sequencing is the controlled displacement of single-stranded DNA through the pore 32 under an applied potential. Nucleases that act gradually or continuously on double-stranded DNA can be used on the cis side of the pore to feed the remaining single strands through under an applied potential or on the trans side under a reverse potential. Similarly, helicases that unwind double-stranded DNA can also be used in a similar manner. Sequencing applications that require strand displacement in response to an applied potential are also possible, but the DNA must first be "captured" by the enzyme under a reverse or no potential. After binding, by switching the potential back, the chain will pass through the pore in a cis-to-trans manner and be held in an extended configuration by the current. Single-stranded DNA exonucleases or single-stranded DNA-dependent polymerases can act as molecular motors to pull the recently displaced single strand back through the pore in a controlled, step-by-step manner in response to an applied potential in a trans-to-cis manner. Alternatively, the single-stranded DNA dependent polymerase may act as a molecular brake that slows down translocation of the polynucleotide through the pore.Any of the moieties, techniques or enzymes described in WO-2012 / 107778 or WO-2012 / 033524 may be used to control polymer movement.
[0243] Generally speaking, when the measurement is a current measurement of ion flow through the aperture 32, the ion flow may typically be a DC ion flow, but in principle AC current (ie the amplitude of an AC current flowing under an applied AC voltage) could alternatively be used.
[0244] The biochemical analysis system 1 can collect types of electrical measurements other than amperometric measurements of ion flow through the nanopore described above.
[0245] Other possible electronic measurements include current measurements, impedance measurements, tunneling measurements (e.g. as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12;11(1):279-85) and field effect transistor (FET) measurements (e.g. as disclosed in WO2005 / 124888).
[0246] As an alternative to electrical measurements, the biochemical analysis system 1 can acquire optical measurements. A suitable optical method involving the measurement of fluorescence is disclosed in J. Am. Chem. Soc. 2009, 131 1652-1653.
[0247] The measurement system 8 can collect electrical measurements other than current measurements of ion flow through the nanopore. Possible electrical measurements include current measurements, impedance measurements, tunneling measurements (e.g., as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12;11(1):279-85), and field effect transistor (FET) measurements (e.g., as disclosed in WO 2005 / 124888).
[0248] Optical measurements can be combined with electrical measurements (Soni GV et al., Rev Sci Instrum. 2010 Jan;81(1):014301).
[0249] The biochemical analysis system 1 can simultaneously collect measurements of different properties. The measurements can be of different properties because they are measurements of different physical properties, which can be any of those described above. Alternatively, the measurements can be of different properties because they are measurements of the same physical property under different conditions, such as electrical measurements, such as current measurements, under different bias voltages.
[0250] The signal output by many types of sensor devices 2 as a series of raw measurement values 11 is typically in the form of a "noisy staircase wave", but is not limited to this signal type. In the case of ion current measurements obtained using a type of measurement system 8 comprising a nanopore, Figure 4 An example of a series of raw measurement values 11 of this form is shown.
[0251] Typically, each measurement value collected by the biochemical analysis system 1 is based on a k-mer, which is a sequence of k polymer units, where k is a positive integer. While ideally, a measurement value would be based on a single polymer unit (i.e., where k is 1), for many typical types of biochemical analysis systems 1, each measurement value is based on a k-mer of multiple polymer units (i.e., where k is a plural number). That is, each measurement value is based on the sequence of each polymer unit in the k-mer, where k is a plural number.
[0252] In a series of measurements acquired by the biochemical analysis system 1, successive groups of multiple measurements are based on the same k-mer. The multiple measurements in each group have a constant value, subject to some variation discussed below, and thus form a "level" in the series of raw measurements. This level is typically formed by measurements based on the same k-mer (or consecutive k-mers of the same type) and thus corresponds to the normal state of the biochemical analysis system 1.
[0253] The signal moves between a set of levels (which can be a large set). Given the sampling rate of the instrument and the noise on the signal, it can be assumed that the transitions between levels are instantaneous, so the signal can be approximated by an idealized step trace.
[0254] The measured value corresponding to each state is constant on the time scale of events, but for most types of biochemical analysis systems 1, will experience changes over short time frames. The changes may be caused by measurement noise, such as that generated from circuits and signal processing, especially in the special case of electrophysiology, from amplifiers. Due to the small amplitude of the characteristics to be measured, this measurement noise is unavoidable. The changes may also come from inherent changes or expansions in the basic physical or biological systems of the biochemical analysis system 1. Most types of biochemical analysis systems 1 will experience such inherent changes to a greater or lesser extent. For any given type of biochemical analysis system 1, two sources of variation may play a role or one of these noise sources may be dominant.
[0255] Additionally, there is typically no prior knowledge of the number of measurements in a group, which varies unpredictably.
[0256] The two aforementioned factors of variation and the lack of knowledge of the number of measurements may make it difficult to distinguish some groups, for example if the groups are short and / or the levels of the measurements of two consecutive groups are close to each other.
[0257] A series of raw measurements may take this form as a result of a physical or biological process occurring in the biochemical analysis system 1. Therefore, in some cases, each set of measurements may be referred to as a "state."
[0258] For example, in some types of biochemical analysis systems 1, events consisting of the displacement of a polymer through the pore 32 can occur in a ratcheting manner. During each step of the ratcheting movement, at a given voltage across the pore 32, the ion flow through the nanopore is constant and undergoes the changes discussed above. Therefore, each set of measurements is associated with a step of the ratcheting movement. Each step corresponds to a state in which the polymer is in a corresponding position relative to the nanopore 32. Although there may be some variation in the precise position between states, there is a large-scale displacement of the polymer between states. Depending on the nature of the biochemical analysis system 1, the states can occur due to binding events in the nanopore.
[0259] The duration of a single state can depend on a variety of factors, such as the potential applied across the pore, the type of enzyme used to ratchet the polymer, whether the polymer is pushed or pulled through the pore by the enzyme present, pH, salt concentration, and the type of nucleoside triphosphate. The duration of a state can typically vary between 0.5 ms and 3 s, depending on the biochemical assay system 1, and for any given nanopore system, with some random variation between states. For any given biochemical assay system 1, the expected distribution of durations can be determined experimentally.
[0260] It is possible to experimentally examine the extent to which a given biochemical analysis system 1 provides a measurement value that depends on the k-mers and the size of the k-mers. A possible approach to this is disclosed in WO-2013 / 041878.
[0261] Returning to the biochemical analysis system 1 , it can acquire types of electrical measurements other than current measurements of ion flow through the nanopore described above.
[0262] Other possible electrical measurements include current measurements, impedance measurements, tunneling measurements (e.g., as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12;11(1):279-85), and field effect transistor (FET) measurements (e.g., as disclosed in WO2005 / 124888).
[0263] Return to Figure 1 The arrangement of the electronic circuit 4 will now be discussed. The electronic circuit 4 is connected to the sensor electrodes 22 associated with each sensor element 30 and to the common electrode 25. The electronic circuit 4 may have an overall arrangement as described in WO 2011 / 067559. The electronic circuit 4 is arranged as follows to control the application of a bias voltage across each sensor element 3 and to acquire measurements from each sensor element 3.
[0264] exist Figure 5A first arrangement for the electronic circuit 4 is shown, which shows components in relation to a single sensor element 30, the components being duplicated for each sensor element 30. In this first arrangement, the electronic circuit 4 comprises a detection channel 40 and a bias control circuit 41, each connected to a sensor electrode 22 of the sensor element 30.
[0265] Detection channel 40 collects measurements from sensor electrodes 22. Detection channel 40 is configured to amplify electrical signals from sensor electrodes 22. Detection channel 40 is therefore designed to amplify very small currents with sufficient resolution to detect changes in properties caused by the interactions of interest. Detection channel 40 is also designed to have a sufficiently high bandwidth to provide the temporal resolution required to detect each such interaction. These constraints require sensitive, and therefore expensive, components. Specifically, detection channel 40 can be configured as described in detail in WO-2010 / 122293 or WO 2011 / 067559 (each of which is incorporated herein by reference).
[0266] The bias control circuit 41 supplies a bias voltage to the sensor electrode 22 for biasing the sensor electrode 22 with respect to the input of the detection channel 40 .
[0267] During normal operation, the bias voltage supplied by the bias control circuit 41 is selected to enable polymer displacement through the pore 32. This bias voltage may typically be up to a level of -200 mV.
[0268] The bias voltage supplied by the bias control circuit 41 can also be selected so as to be sufficient to displace polymer from the pore 32. By causing the bias control circuit 41 to supply this bias voltage, the sensor element 30 can be operated to displace polymer that is displaced through the pore 32. To ensure reliable displacement, the bias voltage is typically a reverse bias, but this is not always necessary. When this bias voltage is applied, the input to the detection circuit 40 is designed to remain at a constant bias potential, even when a negative current (of a magnitude similar to normal current, typically on the order of -50 pA to -100 pA) is present.
[0269] Used for Figure 5 The first arrangement of the electronic circuit 4 , illustrated in , requires a separate detection channel 40 for each sensor element 30 , which is expensive to implement. Figure 6 A second arrangement for the electronic circuit 4 is shown which reduces the number of detection channels 40 .
[0270] In this arrangement, the number of sensor elements 30 in the array is greater than the number of detection channels 40, and the biochemical sensing system is operable to acquire polymer measurements by selecting sensor elements in a multiplexed manner, particularly in an electrical measurement multiplexed manner. This is achieved by providing a switch arrangement 42 between the sensor electrodes 23 of the sensor elements 30 and the detection channels 40. Figure 6 A simplified example with four sensor cells 30 and two detection channels 40 is shown, but the number of sensor cells 30 and detection channels 40 may be larger, typically much larger. For example, for some applications, the sensor device 2 may include a total of 4096 sensor cells 30 and 1024 detection channels 40.
[0271] The switch arrangement 42 may be arranged as described in detail in WO-2010 / 122293. For example, the switch arrangement 42 may include a plurality of 1 to N multiplexers each connected to a group of N sensor elements 30, and may include appropriate hardware such as latches to select the state of the switches.
[0272] Thus, by switching the switch arrangement 42 , the biochemical analysis system 1 can be operated to acquire measurement values of the polymer via the sensor elements 30 selected in an electrically measurement multiplexed manner.
[0273] The switching arrangement 42 can be controlled in a manner as described in WO-2010 / 122293 so as to selectively connect the detection channel 40 to individual sensor elements 30 that have performance of acceptable quality based on the amplified electrical signal output by the detection channel 40, but in addition the switching arrangement is controlled as further described below.
[0274] As in the first arrangement, the second arrangement further comprises a bias control circuit 41 in relation to each sensor element 30 .
[0275] Although in this example the sensor elements 30 are selected in an electrical measurement multiplexing manner, other types of biochemical analysis systems 1 can be configured to switch between the sensor elements in a spatial multiplexing manner, for example by moving a probe for performing the electrical measurements, or by controlling an optical system for acquiring optical measurement values from different spatial positions of different sensor elements 30.
[0276] The data processor 5 connected to the electronic circuit 4 is arranged as follows. The data processor 5 can be a computer device that runs an appropriate program, which can be performed by dedicated hardware devices, or by any combination thereof. The computer device used can be any type of computer system, but typically has a conventional structure. The computer program can be written in any suitable programming language. The computer program can be stored in a computer-readable storage medium, which can be of any type, such as a recording medium that can be inserted into a drive of a computing system and can store information magnetically, optically, or magneto-optically; a fixed recording medium of a computing system, such as a hard drive; or a computer memory. The data processor 5 can include a circuit board that is inserted into a computer, such as a desktop or laptop computer. The data used by the data processor 5 can be stored in its memory 10 in a conventional manner.
[0277] The data processor 5 controls the operation of the electronic circuit 3. As well as controlling the operation of the detection channels 41, the data processor controls the bias control circuit 41 and controls the switching of the switch arrangement 31. The data processor 5 also receives and processes a series of measurements from each detection channel 40. As described further below, the data processor 5 stores and analyzes the series of measurements.
[0278] The data processor 5 controls the bias control circuit 41 to apply a bias voltage sufficient to cause the polymer to shift through the pores 32 of the sensor element 30. This operation of the biochemical sensor element 41 allows a series of measurements to be collected from different sensor elements 30, which can be analyzed by the data processor 5 or by another data processing unit to estimate the sequence of polymer units in the polymer, for example using techniques such as those described in WO-2013 / 041878. The data from the different sensor elements 30 can be collected and combined.
[0279] The data processor 5 receives and analyzes the series of measurements 11 acquired by the sensor device 2 and supplied by the electronic circuit 4. The data processor 5 may also provide control signals to the electronic circuit 5, for example to select a voltage to be applied across the biological pore 1 in the sensor device 2. The series of raw measurements 11 may be applied over any suitable connection, such as a direct connection if the data processor 5 and sensor device 2 are physically located together, or any type of network connection if the data processor 5 and sensor device 2 are physically remote from each other.
[0280] We will now describe Figure 7A method for controlling a biochemical analysis system 1 to analyze polymers is shown. The method is in accordance with a first aspect of the present invention and is performed in a manner that increases the speed of analysis by rejecting polymers that do not require further analysis. The method is implemented in a data processor 5. The method is performed in parallel for each sensor element 30 that acquires a series of measurements, i.e., each sensor element 30 in a first arrangement for electronic circuit 4, and each sensor element 30 connected to a detection channel 40 via a switch arrangement 42 in a second arrangement for electronic circuit 4.
[0281] In step C1, the biochemical analysis system 1 is operated by controlling the bias control circuit 30 to apply a bias across the pore 32 of the sensor element 30 sufficient to enable displacement of the polymer. Based on the output signal from the detection channel 40, displacement is detected and measurement acquisition begins. A series of measurements are acquired over time.
[0282] In some cases, the following steps operate on a series of raw measurement values 11 collected by the sensor device 2, i.e., a series of measurement values of the above type, which comprises consecutive groups of multiple measurement values depending on the same k-mer without a priori knowledge of the number of measurements of any group.
[0283] In other cases, such as Figure 8 As shown, the raw measurement values 11 are pre-processed using a state detection step SD to derive a series of measurement values 12 that are used in the following steps instead of the raw measurement values.
[0284] In this state detection step SD, a series of raw measurement values 11 is processed to identify consecutive groups of raw measurement values and to derive a series of measurement values 12 consisting of a predetermined number of measurement values for each identified group. Thus, a series of measurement values 12 is derived for each sequence of measured polymer units. The purpose of the state detection step SD is to reduce the series of raw measurement values to a predetermined number of measurement values associated with each k-mer to simplify the subsequent analysis. For example, a noisy step wave signal, such as Figure 4 As shown, it can be reduced to states where the single measurement value associated with each state can be the average current. Such states can be called levels.
[0285] Figure 9 An example of such a state detection step SD is shown as follows, looking for short-term increases in the derivatives of a series of raw measured values 11 .
[0286] In step SD-1, a series of raw measurements 11 are differenced to derive their derivatives.
[0287] In step SD-2, the derivative of step SD-1 is subjected to low-pass filtering to suppress high-frequency noise, and the difference in step SD-1 tends to be amplified.
[0288] In step SD-3, the filtered derivatives from step SD-2 are thresholded to detect transition points between groups of measurements, thereby identifying groups of original measurements.
[0289] In step SD-4, a predetermined number of measurements are derived from each group of raw measurements identified in step SD-3. The measurements output from step SD-4 form a series of measurements 12.
[0290] The predetermined number of measurement values may be one or more.
[0291] In the simplest approach, a single measured value is derived from each group of raw measured values, such as the mean, median, standard deviation or number of raw measured values in each determined group.
[0292] In other approaches, a predetermined plurality of measurements of different properties are derived from each group, such as any two or more of the mean, median, standard deviation, or number of raw measurements in each determined group. In this case, the predetermined plurality of measurements of different properties are collected based on the same k-mer because they are different measures in the same group of raw measurements.
[0293] The state detection step SD can be used Figure 9 Different methods from those shown. For example, Figure 9 A common simplification of the method shown is to use a sliding window analysis, which compares the means of two adjacent windows of data. A threshold can then be set directly based on the mean difference, or based on the variance of the data points in the two windows (e.g., by calculating the Student's t-statistic). A distinct advantage of these methods is that they can be applied without imposing numerous assumptions about the data.
[0294] Other information related to the measurement level can be stored for later analysis. Such information may include, but is not limited to: signal variance; asymmetry information; confidence level of the observation; and group length.
[0295] For example, Figure 10 An example is given of a series of raw measurements reduced by a moving window t-test as determined experimentally11. In particular, Figure 10 A series of raw measurements 11 are shown as thin lines. The levels after state detection are shown superimposed as dark lines.
[0296] Step C2 is performed while the polymer is partially translocated through the nanopore, i.e., during translocation. At this point, a series of measurements collected from the polymer during the partial translocation period are collected for analysis, referred to herein as a "chunk" of measurements. Step C2 can be performed after a predetermined number of measurements have been collected such that the chunk of measurements has a certain size, or alternatively, after a predetermined amount of time. In the former case, the size of the chunk of measurements can be defined by parameters that are initialized at the start of the run, but dynamically changed such that the size of the chunk of measurements changes.
[0297] In step C3, the blocks of measurements collected in step C2 are analyzed. This analysis uses reference data 50. As discussed in more detail below, reference data 50 is derived from at least one reference sequence of polymer units. The analysis performed in step C3 provides a measure of similarity between (a) the sequence of polymer units of the partially shifted polymer for which measurements were collected and (b) a reference sequence. A variety of techniques are possible for performing this analysis, some examples of which are described below.
[0298] The measure of similarity can represent the similarity to the entire reference sequence or to a portion of the reference sequence, depending on the application. The technique applied in step C3 of deriving the measure of similarity, such as a global or local approach, can be selected accordingly.
[0299] Additionally, a measure of similarity may represent a variety of different measures of similarity, provided that it generally provides a measure of how similar the sequences are. Set forth below are some examples of specific measures of similarity that can be determined from sequences in different ways.
[0300] In step C4, a decision is made in response to the measure of similarity determined in step C3 to (a) reject the measured polymer, (b) require additional measurements to make the decision, or (c) continue collecting measurements until the end of the polymer.
[0301] If the determination made in step C4 is that (a) the polymer being measured is to be excluded, the method proceeds to step C5 , where the biochemical analysis system 1 is controlled to exclude the polymer so that measurement values can be collected from another polymer.
[0302] Step C5 is performed differently between the first and second arrangements of the electronic circuit 4 as follows.
[0303] In the case of the first arrangement of the electronic circuit 4, then in step C5, the bias control circuit 30 is controlled to apply a bias voltage across the aperture 32 of the sensor element 30 sufficient to displace the currently displaced polymer. This displaces the polymer and thereby makes the aperture 32 available to receive further polymer. After this displacement in step C5, the method returns to step C1, so that the bias control circuit 30 is controlled to apply a bias voltage across the aperture 32 of the sensor element 30 sufficient to displace further polymer through the aperture 32.
[0304] In the case of the second arrangement of the electronic circuit 4, then in step C5, the analytical biochemical analysis system 1 is caused to stop acquiring measurements by the currently selected sensor element 30 by controlling the switch arrangement 42 to disconnect the detection channel 40 currently connected to the sensor element 30 and selectively connect the detection channel 40 to a different sensor element 30. Simultaneously, in step C5, the bias control circuit 30 is controlled to apply a bias voltage across the aperture 32 of the sensor element 30 sufficient to expel the polymer currently displaced through the currently selected sensor element 30, so that the sensor element 30 is available to receive further polymer in the future.
[0305] The method then returns to step C1 , where it is applied to the newly selected sensor element 30 so that the biochemical analysis system 1 starts acquiring measurement values therefrom.
[0306] If the determination made in step C4 is (b) that additional measurements are required to make the determination, the method returns to step C2. Thus, measurements of the displaced polymer continue to be collected until the next block of measurements is controlled in step C2 and analyzed in step C3. When step C2 is performed again, the block of measurements collected may be simply new measurements analyzed in isolation, or may be new measurements combined with previous blocks of measurements.
[0307] If the determination in step C4 is (c) to continue acquiring measurements until the end of the polymer, the method proceeds to step C6 without repeating steps C2 and C3, thereby not analyzing additional blocks of data. In step C6, sensor element 1 continues operating, thereby continuing to acquire measurements until the end of the polymer. Thereafter, the method returns to step C1, allowing for analysis of additional polymers.
[0308] As represented by a measure of similarity, the degree of similarity used as the basis for the decision in step C4 can vary depending on the nature of the application and reference sequences. Thus, if the decision is responsive to a measure of similarity, then generally there is no limit to the degree of similarity used to make different decisions.
[0309] Some examples of how the dependence of the measure of similarity may change are as follows.
[0310] In an application where the reference sequence of polymer units is an undesirable sequence and a determination is made in step C4 that a polymer is excluded based on a measure of similarity indicating that the partially displaced polymer is an undesirable sequence, a relatively high degree of similarity can be used as the basis for the exclusion of the polymer. Similarly, in the context of an application, the degree of similarity can change according to the properties of the reference sequence. When it is intended to distinguish similar sequences, a higher degree of similarity can be required as the basis for exclusion.
[0311] In contrast, in applications where the reference sequence of polymer units from which reference data 50 is derived is the target and a decision to exclude the polymer is made in step C4 in response to a measure of similarity indicating that the partially displaced polymer is not the target, a relatively low similarity can be used as a basis for excluding the polymer.
[0312] As another example, if the application is to determine whether a known gene from a known bacterium is present in samples of multiple bacteria, the degree of similarity required to determine whether a polynucleotide has the same sequence as the target will be higher if the gene has a conserved sequence across different strains than if the sequence is not conserved.
[0313] Similarly, in some embodiments of the invention, the measure of similarity will be equivalent to the degree of identity of a polymer to a target polymer, while in other embodiments, the measure of similarity will be equivalent to the probability that a polymer is identical to the target polymer.
[0314] The similarity required as the basis for repulsion can also be changed according to possible time saving, which itself depends on the application described below. An acceptable false positive rate can depend on time saving. For example, when the possible time saving of repulsion of an undesirable polymer is relatively high, it is acceptable to repulse an increased proportion of polymers as a target, provided that there is an overall time saving of repulsion of an undesirable polymer.
[0315] Now back Figure 7 In the method, if at any point during the collection of measurements for a polymer, it is detected that no further measurements are being collected, indicating that the end of the polymer has been reached, the method immediately returns to step C1 so that further polymer can be analyzed. After measurements for the entire polymer have been collected in this manner, those measurements can be analyzed, for example, to derive an estimate of the sequence of the polymer units, as disclosed in WO-2013 / 041878.
[0316] The source of the reference data 50 may vary depending on the application. The reference data 50 may be generated from a reference sequence of polymer units or from measurements taken from a reference sequence of polymer units.
[0317] In some applications, previously generated reference data 50 may be pre-stored. In other applications, reference data 50 is generated while the method is being performed.
[0318] Reference data 50 can be provided for a single reference sequence of polymer units or for multiple reference sequences of polymer units. In the latter case, either step C3 is performed for each sequence, or one of the multiple reference sequences is selected for use in step C3. In the latter case, the selection can be made based on various criteria. For example, reference data 50 can be applied to different types of biochemical analysis systems 1 (e.g., different nanopores) and / or external conditions. In such cases, the reference model 70 described below is selected based on the type of biochemical analysis system 1 actually used and / or the actual external conditions.
[0319] Figure 7 The method shown may vary depending on the application. For example, in some variations, the decision in step C4 is never (c) continue collecting measurements to the end of the polymer, such that the method repeatedly collects and analyzes chunks of measurements to the end of the polymer.
[0320] In another variant, in step C3 , instead of using reference data 50 and determining a measure of similarity, the decision to exclude a polymer in step C4 may be based on another analysis of a series of measured values, in general any analysis based on a block of measured values.
[0321] In one possibility, step C3 may analyze whether the block of measured values is of insufficient quality, for example has a noise level exceeding a threshold, has wrong ratios, or has characteristics of damaged polymers.
[0322] Based on this analysis, the decision in step C4 is made to exclude the polymer based on an internal quality control check. This still involves making a decision to exclude the polymer based on a block of measurements, i.e., a series of measurements collected from the polymer during a partial displacement period. Therefore, in contrast to the expelled polymer that caused the blockage, in the case of the excluded polymer, the polymer is no longer displaced, so no k-mer-dependent measurements are collected.
[0323] In another possibility, wherein the method is according to the second aspect of the invention, Figure 11 The modified method is shown in Figure 7The method is the same as that of , except that step C3 is modified. In step C3, instead of using reference data 50 derived from at least one reference sequence of polymer units and determining a measure of similarity, the measurements are processed using observations that are a series of k-mer states of different possible types, and the following general model 60 is included: a transition weight 61, for each transition between consecutive k-mer states in the series of k-mer states, for a possible transition between possible types of k-mer states; and an emission weight 62, for each type of k-mer state, which represents the probability of observing a measurement value of a given k-mer. Step C3 is modified to include a measure of fitting the reference model 60.
[0324] The general model 60 may be of the type described in WO-2013 / 041878. For details of the model, refer to WO-2013 / 041878. Figure 13 , the general model 60 is further described below. A measure of fit is derived, for example as the likelihood of the measurement being observed by the most similar sequence of k-mer states. This measure of fit indicates the quality of the measurement.
[0325] When step C3 is modified in this way, the decision in step C4 is made based on the measure of fit, thereby rejecting polymers based on internal quality control checks.
[0326] Thus, if similarity to a reference sequence of polymer units indicates that further analysis of the polymer is unnecessary, or if measurements taken from the polymer are of poor quality, as determined by the model, such that additional shifts and measurements are unacceptable, the method results in the polymer being rejected. The extent to which the model poorly represents the data depends on the complexity of the model itself. For example, more complex models may have parameters that account for some of the conditions that could lead to rejection.
[0327] Conditions that may cause rejection may include, for example: incoming unacceptable signals; high noise; non-model behavior; irregular system errors such as temperature fluctuations; and / or errors due to the electro-physical system.
[0328] For example, one possibility is that polymer or other debris has become lodged in the nanopore, producing a slowly varying, fairly static current. Models typically expect well-separated (time-constant) steps in the data, so such measurements will have poor fits to the model's metric.
[0329] The second possibility is transient noise, such as large changes in current between otherwise closely spaced steps. If this noise occurs at a high frequency, the data may be of little use for practical purposes. Due to the high frequency of the undesired measurements, the metric fit to the model will be low.
[0330] These "errors" can occur in a non-instantaneous manner. Indeed, it is often observed that measured sections exhibit a shift in their average current relative to neighboring sections. Possible explanations for this are variations in pore morphology and polymer molecules. Regardless of the cause, this behavior is not captured by the model, making the data of little use for practical purposes.
[0331] The effects of this error can be mitigated to some extent by increasing the complexity of the model. However, this is not desirable and can lead to increased computational costs in modeling the data and decoding the polymer sequence.
[0332] Since such polymer chains are excluded, only those polymer sequences with strong homology for which the model transition weights and emission weights are derived give measurements with good fit to the model metric.
[0333] After completing the collection of measurements for the entire polymer, those measurements may be analysed as disclosed in WO-2013 / 041878, for example to derive an estimate of the sequence of the polymer units.
[0334] Can be applied independently or in combination Figure 7 and Figure 11 In such cases, the two methods can be performed simultaneously (e.g., step C3 of the two methods is performed in parallel, and the other steps are performed together) or continuously (e.g., in parallel). Figure 7 Before the method Figure 11 method) to apply them.
[0335] We will now describe Figure 12 The method for controlling the biochemical analysis system 1 to classify polymers is shown. This method is a third aspect of the present invention. In this case, the sample chamber 24 contains a sample containing polymers of different types, and the groove 21 serves as a collection chamber for collecting the classified polymers.
[0336] The method is implemented in the data processor 5. The method is performed in parallel with respect to a plurality of sensor elements 30, for example each sensor element 30 in the first arrangement for the electronic circuit 4 and each sensor element 30 connected to a detection channel 40 via a switch arrangement 42 in the second arrangement for the electronic circuit 4.
[0337] In step D1, the biochemical analysis system 1 is operated by controlling the bias control circuit 30 to apply a bias voltage across the pore 32 of the sensor element 30 sufficient to enable polymer displacement. This causes the polymer to begin displacing through the nanopore and the following steps are performed during displacement. Based on the output signal from the detection channel 40, the displacement is detected and measurement acquisition begins. A series of polymer measurements are collected by the sensor element 30 over time.
[0338] In some cases, the following steps operate on a series of raw measurement values 11 collected by the sensor device 2, i.e., a series of measurement values of the above type, which comprises consecutive groups of multiple measurement values depending on the same k-mer without knowing in advance the number of measurement values of any group.
[0339] In other cases, the raw measurement values 11 are pre-processed using the state detection step SD to derive a series of measurement values 12 for the following steps instead of the raw measurement values. Figure 8 and Figure 9 , the state detection state SD is performed in the same manner as step C1 described above.
[0340] Step D2 is performed while the polymer is partially translocated through the nanopore, i.e., during translocation. At this point, a series of measurements collected from the polymer during the partial translocation period are collected for analysis, referred to herein as a "chunk" of measurements. Step D2 can be performed after a predetermined number of measurements have been collected such that the chunk of measurements has a certain size, or alternatively, after a predetermined amount of time. In the former case, the size of the chunk of measurements can be defined by parameters that are initialized at the start of the run, but dynamically changed such that the size of the chunk of measurements changes.
[0341] In step D3, the blocks of measurements collected in step D2 are analyzed. This analysis uses reference data 50. As discussed in more detail below, reference data 50 is derived from at least one reference sequence of polymer units. The analysis performed in step D3 provides a measure of similarity between (a) the sequence of polymer units of the partially shifted polymer for which measurements were collected and (b) a reference sequence. A variety of techniques are possible for performing this analysis, some examples of which are described below.
[0342] The measure of similarity may represent similarity to the entire reference sequence or to a portion of the reference sequence, depending on the application. The technique applied in step D3 of deriving the measure of similarity, such as a global or local approach, may be selected accordingly.
[0343] Additionally, a measure of similarity may represent a variety of different measures of similarity, provided that it generally provides a measure of how similar the sequences are. Set forth below are some examples of specific measures of similarity that can be determined from sequences in different ways.
[0344] In step D4, based on the measure of similarity determined in step D3, a determination is made as to whether (a) additional measurements are needed to make the determination, (b) the displacement of the polymer into recess 21 is complete, or (c) the measured polymer is expelled back into sample chamber 24. If the determination in step D4 is that (a) additional measurements are needed to make the determination, the method returns to step D2. Thus, measurements of the displaced polymer continue to be collected until the next block of measurements is collected in step D2 and analyzed in step D3. The block of measurements collected when step D2 is repeated may be simply new measurements analyzed in isolation or may be new measurements combined with the previous block of measurements.
[0345] If the decision made in step D4 is that (b) the displacement of the polymer into the groove 21 is complete, the method proceeds to step D6 without repeating steps D2 and D3 , so that no further analysis of the measured values is performed.
[0346] In step D6, the displacement of the polymer into the groove 21 is completed. As a result, the polymer is collected in the groove 21.
[0347] Step D6 may be performed by applying the same bias voltage across the aperture 32 of the sensor element 30 that enables displacement of the polymer.
[0348] Alternatively, in step D6, the bias voltage can be changed to carry out the remaining shift of the polymer at an increased rate to reduce the time spent on the shift. This is advantageous because it increases the overall rate of the classification process. Increasing the shift rate is acceptable because it is no longer necessary to analyze the polymer. Typically, the change in bias voltage can be elevated. In a typical system, the increase can be significant. For example, in one embodiment, the shift rate can be increased from about 30 bases / second to about 10,000 bases / second. The possibility of changing the shift rate can depend on the configuration of the sensor element. For example, when a polymer binding moiety such as an enzyme is used to control the shift, this can depend on the polymer binding moiety used. Advantageously, a polymer binding moiety that can control the rate can be selected.
[0349] During step D6 , the sensor element 1 can continue to be operated so that measurement values continue to be acquired up to the end of the polymer, but this is optional since the remaining sequence does not need to be determined.
[0350] After step D6, the method returns to step D1 so that further polymer can be displaced.
[0351] If the determination made in step D4 is (c) to discharge the polymer, the method proceeds to step D5 , where the biochemical analysis system 1 is controlled to discharge the measured polymer back to the sample chamber 24 so that additional polymer measurement values can be collected.
[0352] In step D5, bias control circuit 30 is controlled to apply a bias voltage across aperture 32 of sensor element 30 sufficient to displace the currently displaced polymer. This displaces the polymer and thereby makes aperture 32 available to receive additional polymer. After this displacing in step D5, the method returns to step D1, whereupon bias control circuit 30 is controlled to apply a bias voltage across aperture 32 of sensor element 30 sufficient to displace additional polymer through aperture 32.
[0353] The method repeats upon returning to step D1. The repetitive nature of the method results in a continuous displacement and processing of the polymer from the sample chamber 24.
[0354] Thus, the method utilizes a measure of similarity provided by analyzing a series of measurements taken from a polymer during a partial displacement as a basis for determining whether to collect a continuous polymer in recess 21. In this way, the polymers of the sample in sample chamber 24 are sorted and the desired polymers are selectively collected in recess 21.
[0355] The collected polymer can be recovered by removing the sample from the sample chamber 24 and then recovering the polymer from the recess 21. This can be done after repeating the method. Alternatively, this can be done during the displacement of the polymer from the sample, for example by providing the biochemical analysis system 1 with a fluidic system for extracting the polymer from the recess 21.
[0356] This method can be applied to a wide variety of applications. For example, it can be applied to polymers that are polynucleotides, such as viral genomes or plasmids. Viral genomes are typically on the order of 10-15 kB (kilobases) in length, and plasmids are typically on the order of 4 kB in length. In this embodiment, the polynucleotides do not need to be fragmented and can be collected in their entirety. The collected viral genomes or plasmids can be used in any manner, for example, to transfect cells. Transfection is the process of introducing DNA into the cell nucleus and is an important tool for studying gene function and the regulation of gene expression, thereby promoting advances in basic cell research, drug discovery, and target validation. RNA and proteins can also be transfected.
[0357] As indicated by the measure of similarity, the degree of similarity used as the basis for the decision in step D4 can vary depending on the application and the nature of the reference sequence. Therefore, if the decision is based on a measure of similarity, then generally speaking, there is no restriction on the degree of similarity used to make different decisions.
[0358] Some examples of how the dependence of the measure of similarity may change are as follows.
[0359] In many applications, the reference sequence of polymer units from which reference data 50 is derived is the desired sequence. In that case, in step D4, a determination is made that the shift is complete in response to a measure of similarity indicating that the partially shifted polymer is the desired sequence, and a relatively high degree of similarity can be used as a basis for completing the shift.
[0360] However, this is not necessary. In some applications, the reference sequence of polymer units is an undesirable sequence. In that case, in step D4, a determination is made to complete the shift in response to indicate that the partially shifted polymer is not a measure of the similarity of an undesirable sequence.
[0361] Similarly, in the context of application, the degree of similarity may vary depending on the nature of the reference sequence. When aiming to distinguish between similar sequences, a higher degree of similarity may be required as a basis for exclusion.
[0362] The method is carried out in step D4 using the same reference data 50 and the same standard for each sensor element 30. In that case, each groove 21 collects the same polymer in parallel.
[0363] Alternatively, a method can be performed to collect different polymers in different recesses 21. In this case, differential sorting is performed. In one embodiment, different reference data 50 are used for different sensor elements 30. In another embodiment, the same reference data 50 is used for different sensor elements 30, but step D4 is performed with different reliance on the similarity measure for the different sensor elements.
[0364] Can be changed according to application Figure 7 、 Figure 11 and Figure 12 The method shown.
[0365] Various different types of reference sequences for polymer units can be used depending on the application. Without limitation, where the polymer is a polynucleotide, the reference sequence for the polymer unit can comprise one or more reference genomes or, one or more regions of interest of a genome to which measurements are compared.
[0366] The source of the reference data 50 may vary depending on the application.The reference data may be generated from a reference sequence of polymer units or from measurements taken from a reference sequence of polymer units.
[0367] In some applications, previously generated reference data 50 may be pre-stored. In other applications, the reference data 50 is generated while the method is being performed.
[0368] Reference data 50 can be provided for a single reference sequence of polymer units or for multiple reference sequences of polymer units. In the latter case, either step D3 is performed for each sequence, or one of the multiple reference sequences is selected for use in step D3. In the latter case, the selection can be made based on various criteria, depending on the application. For example, the reference data 50 may be applicable to different types of biochemical analysis systems 1 (e.g., different nanopores) and / or external conditions. In such cases, the reference model 70 described below is selected based on the type of biochemical analysis system 1 actually used and / or the actual external conditions.
[0369] The biochemical analysis system 1 described above is an example of a biochemical analysis system comprising an array of sensor elements each comprising a nanopore. However, the above method can generally be applied to any biochemical analysis system operable to continuously measure values of a collection polymer, possibly without using a nanopore.
[0370] An example of such a biochemical analysis system that does not include a nanopore is a scanning probe microscope, which can be an atomic force microscope (AFM), a scanning tunneling microscope (STM), or another form of scanning microscope. In this case, the biochemical analysis system can be operable to collect continuous measurements of selected polymers in a spatially multiplexed manner. For example, the polymers can be disposed on a substrate at different spatial locations, and spatial multiplexing can be provided by movement of the scanning probe microscope.
[0371] In the case where the reader is an AFM, the resolution of the AFM tip can be less fine than the size of a single polymer unit. Thus, the measurement can be a function of multiple polymer units. The AFM tip can be functionalized so as to interact with the polymer unit in an alternative manner as if it were not functionalized. The AFM can be operated in contact mode, non-contact mode, tapping mode, or any other mode.
[0372] Where the reader is an STM, the resolution of the measurement may be less fine than the size of a single polymer unit, such that the measurement is a function of multiple polymer units.The STM may be operated conventionally or in any other mode or to perform spectroscopic measurements (STS).
[0373] The form of the reference data 50 used in any of the above methods will now be discussed. The reference data 50 can take a variety of forms derived from the reference sequence of polymer units in different ways. The analysis performed in step C4 or D4 to provide a measure of similarity depends on the form of the reference data 50. Some non-limiting examples will now be described.
[0374] In a first example, the reference data 50 represents the identity of polymer units of at least one reference sequence. In that case, step C4 or D4 comprises the following Figure 13 The process shown.
[0375] In step C4a-1, the block of measurements 63 is analyzed to provide an estimate of the identity of the polymer units of the sequence of polymer units of the partially shifted polymer 64. In general, step C4a-1 can be performed using any method for analyzing measurements collected by a biochemical analysis system.
[0376] Step C4a-1 can in particular be carried out using the method described in detail in WO-2013 / 041878, which is incorporated herein by reference. Reference is made to WO-2013 / 041878 for details of the method, but an outline is given below.
[0377] The method refers to a general model 60 comprising transition weights 61 and emission weights 62 for a series of k-mer states corresponding to chunks 63 of measurement values.
[0378] About every kind of conversion between the continuous k aggregate state in a series of k aggregate states provide conversion weighting 61.Every kind of conversion can be considered as from starting point k aggregate state to terminal point k aggregate state.Conversion weighting 61 represents the relative weighting of the possible conversion between the k aggregate state of possible type, and it is from the starting point k aggregate state of any type to the terminal point k aggregate state of any type.In general, this comprises the weighting for the conversion between two k aggregate states of same type.
[0379] An emission weighting 62 is provided for each type of k-aggregate state. The emission weighting 62 is the weighting of different measurement values for observation when the k-aggregate state is of that type. Conceptually, the emission weighting 62 can be considered to represent the probability of observing a given value of the measurement value of the k-aggregate state, but they do not need to be probabilities.
[0380] Conceptually, the transition weights 61 can be thought of as representing the probability of a possible transition, although they need not be probabilities. Thus, the transition weights 61 take into account the probability of a k-mer state at which a measurement value depends on its transition between different k-mer states, which can be more or less likely depending on the types of the starting and ending k-mer states.
[0381] By way of example and not limitation, the model may be an HMM where the transition weights 61 and transmit weights 62 are probabilities.
[0382] Step C4a-1 uses the reference model 60 to derive an estimate 64 of the identity of the sequence of polymer units of the partially shifted polymer. This can be performed using known techniques applicable to the properties of the reference model 60. Typically, such techniques derive an estimate 64 of the sequence observations of the k-mer state based on the likelihood of the measurements predicted by the reference model 60. As described in WO-2013 / 041878, such techniques can be performed on a series of original measurements 11 or a series of measurements 12.
[0383] This approach can also provide a measure of the fit of the measurement to the model, such as a quality score representing the likelihood that the measurement is predicted by the reference model 60 of the most likely sequence observation of the k-mer state. Such measures are typically derived because they are used to derive the estimate 64.
[0384] As an example, in the case where the general model is an HMM, analytical techniques can be used to solve known algorithms for HMMs, such as the Viterbi algorithm well known in the art. In that case, an estimate 64 is derived based on the likelihood generated by the total sequence of k-aggregate states predicted by the general model.
[0385] As another example, where general model 60 is an HMM, the analysis technique may be of the type disclosed in Fariselli et al., “The posterior-Viterbi: a new decoding algorithm for hidden Markov models,” filed January 4, 2005, at the Department of Biology, Casadio University, Cornell University. In this method, rather than simply selecting the most likely k-mer for each event, a posterior matrix (representing the probability of a measurement value observed for each k-mer state) and consistent paths (paths where adjacent k-mer states tend to overlap) are derived. Essentially, this allows the recovery of the same information directly obtained by applying the Viterbi algorithm.
[0386] The above description is given in terms of a general model 60 that is an HMM in which the transition weights 61 and the emission weights 62 are probabilities, and the method uses a probabilistic technique referring to the general model 60. However, it is alternatively possible that the general model 60 uses a framework in which the transition weights 61 and / or the emission weights 62 are not probabilities, but represent the chance of a transition or measurement in some other way. In this case, the method can use an analytical technique rather than a probabilistic technique that is based on the likelihood predicted by the general model 60 of a series of measurements produced by a sequence of polymer units. The analytical technique can explicitly use a likelihood function, but in general this is not necessary.
[0387] In step C4a-2, the estimated value 64 is compared with the reference data 50 to provide a measure of similarity 65. This comparison can use any known technique for comparing two sequences of polymer units, typically an alignment algorithm that derives an alignment map between the polymer units, along with a score for the accuracy of the alignment map (and therefore the measure of similarity 65). Any number of available fast alignment algorithms can be used, such as the Smith-Waterman alignment algorithm, BLAST or their derivatives, or k-mer counting techniques.
[0388] This embodiment of the reference data 50 in this form has the advantage that the process for deriving the measure of similarity 65 is rapid, but other forms of reference data are possible.
[0389] In a second embodiment, the reference data 50 represent actual or simulated measurement values acquired by the biochemical analysis system 1. In that case, step C4 or D4 comprises Figure 14 The process shown simply comprises a step C4b of comparing a block 63 of measurement values (in this case collected from a series of raw measurement values 11) with the reference data 50 to derive a measure of similarity 65. Any suitable comparison may be made, for example using a distance function to provide a measure of the distance between the two series of measurement values as the measure of similarity 65.
[0390] In a third embodiment, the reference data 50 represents a feature vector of a time series characteristic, which represents a characteristic of the measurement values collected by the biochemical analysis system 1. Such a feature vector can be derived as described in detail in WO-2013 / 121224, which is incorporated herein by reference. In that case, step C4 or D4 comprises the following steps: Figure 15 The process shown.
[0391] In step C4c-1, a block 63 of measured values, in this case acquired from a series of raw measured values 11, is analyzed to derive a feature vector 66 representing a temporal characteristic of the behavior of the measured values.
[0392] In step C4c-2, the feature vector 66 is compared with the reference data 50 to derive a measure of similarity 65. The comparison may be performed using the method described in detail in WO-2013 / 121224.
[0393] In the fourth embodiment, the reference data 50 represent the reference model 70. In that case, step C4 or D4 comprises Figure 16 The process shown comprises a step C4d of fitting a model to a block 63 of a series of measurements to provide a measure 65 of the similarity of fit of a reference model 70 to the block 63 of measurements. The block 63 of measurements may be a series of raw measurements 11 or a series of measurements 12.
[0394] Step C4d can be performed as follows.
[0395] Reference model 70 is a model of a reference sequence of polymer units in biochemical analysis system 1. Reference model 70 treats measured values as observations of k-mer states corresponding to a reference series of the reference sequence of polymer units. The k-mer states of reference model 70 can model the actual k-mers on which the measured values depend, but this is not mathematically necessary, so the k-mer states can be an abstract concept of the actual k-mers. Therefore, different types of k-mer states can correspond to different types of k-mers present in the reference sequence of polymer units.
[0396] Reference model 70 can be considered an adaptation of general model 60 of the type described above and in WO-2013 / 041878 to model the measurements specifically obtained when measuring a reference sequence. Thus, reference model 70 treats the measurements as observations of k-mer states 73 corresponding to a reference series of polymer units. Thus, reference model 70 has the same form as general model 60, in particular including conversion weights 71 and emission weights 72, which will now be described.
[0397] The conversion weight 71 represents the conversion between the k-polymer states 73 of the reference series. Those k-polymer states 73 correspond to the reference sequence of polymer units. Therefore, the continuous k-polymer states 73 in the reference series correspond to the continuous overlapping groups of k polymer units. Thus, there is an intrinsic mapping between the k-polymer states 73 of the reference series and the polymer units of the reference sequence. Similarly, each k-polymer state 73 has a type of combination of different types of each polymer unit in the group corresponding to the k polymer units.
[0398] refer to Figure 17 The state diagram is used to illustrate this. Figure 17 An example of three consecutive k-mer states 73 in a reference series of estimated k-mer states 73 is shown. In this example, k is 3, and the reference series of polymer units includes consecutive polymer units labeled A, A, C, G, T (although those specific types of k-mer states 73 are of course not limited). Thus, the consecutive k-mer states 73 corresponding to the reference series of polymer units are of types AAC, ACG, CGT, which correspond to the measured sequence of polymer units, AACGT.
[0399] Figure 18The state diagram shows transitions between the reference series of k-mer states 73 represented by transition weights 71. In this embodiment, the state may only allow the reference series of k-mer states 73 to travel through in the forward direction (but may also allow backward travel in general). Three different types of transitions 74, 75, and 76 are shown below.
[0400] From each given k-mer state 73 in the reference series, a transition 74 to the next k-mer state 73 is allowed. This models the likelihood of successive measurements of a series of measurements 12 collected from successive k-mers of polymer units of the reference sequence. In the case of preprocessing the blocks of measurements 63 to identify successive groups of measurements, and deriving a series of process measurements consisting of a predetermined number of measurements relative to each identified group for further analysis, the transition weight 71 indicates such a transition 74 with a relatively high likelihood.
[0401] For each given k-mer state 73 in the reference series, a transition 75 to the same k-mer state is allowed. This models the likelihood of a series of consecutive measurements 12 collected from the same k-mer of the polymer unit of the reference sequence. This can be called a "stay". In the case of pre-processing the block 63 of measurement values to identify consecutive groups of measurement values and deriving a series of process measurement values consisting of a predetermined number of measurement values (for each identified group) for further analysis, the transition weight 71 represents such a transition 75 that has a relatively high likelihood compared to the transition 74.
[0402] For each given k-mer state 73 in the reference series, a transition 76 to a subsequent k-mer state 73 is allowed to skip the next k-mer state 73. This models the likelihood of no measurement value being collected from the next k-mer state so that consecutive measurements in a series of measurements 12 collected by the k-mer of the reference sequence of polymer units are separated. This can be called a "skip". In the case of preprocessing the block 63 of measurement values to identify consecutive groups of measurement values and deriving a series of process measurement values consisting of a predetermined number of measurement values (about each identified group) for further analysis, the transition weight 71 represents such a transition 76 with a relatively high likelihood compared to the transition 74.
[0403] The levels of the transitions 75 and 76 representing jump and stay with respect to the level of the transition weight 71 representing the transition 74 can be obtained in the same manner as the transition weight 61 for jump and stay in the general model 31 described above.
[0404] In an alternative embodiment, where there is no pre-processed block 63 of measurements to identify consecutive groups of measurements and derive a series of processed measurements, such that additional analysis is performed on the block 63 of measurements itself, then the transition weights 71 are similar, but rewritten to increase the likelihood of transitions 75 representing jumps to represent consecutive measurements collected by the same k-mer. The level of transition weight 71 used for transitions 75 depends on the expected number of measurements collected by any given k-mer and can be determined experimentally for the particular biochemical analysis system 1 being used.
[0405] About every kind of k-aggregate state, provide emission weighting 72.Emission weighting 72 is the weighting for the different measured values observed when observing the k-aggregate state.Therefore emission weighting 72 depends on the type of the k-aggregate state in question.Especially, the emission weighting 72 for the k-aggregate state of any given type is identical with the emission weighting 62 for the k-aggregate state of those types in the general model 60 mentioned above.
[0406] The same reference model as above is used except that the reference model 70 replaces the general model 60. Figure 13 The same technique described carries out step C4d, fitting the model to a series of blocks of measurements 63 to provide a measure 65 of similarity to the fit of the reference model 70 to the block 63 of measurements.
[0407] Due to the form of the reference model 70, in particular the representation of the transitions between reference series of k-mer states 73, the application of the model inherently derives an estimate of the alignment mapping between the block 63 of measurements and the reference series of k-mer states 73. This can be understood as follows. Since the general model 60 represents the transitions between possible types of k-mer states, the application of this model provides an estimate of the type of k-mer state observed for each measurement value in particular. Since the reference model 70 represents the transitions between reference series of k-mer states 73, the application of this reference model 70 instead estimates the k-mer state 73 of the reference series by which each measurement value is observed, which is the alignment mapping between the k-mer state 73 of the series of measurements and the reference series.
[0408] Additionally, the algorithm derives a score for the accuracy of the alignment, e.g., representing the likelihood that the estimate of the alignment is correct, e.g., because the algorithm derives the alignment based on such scores for different pathways in the model. Thus, this score for the accuracy of the alignment is a measure of similarity 65.
[0409] As an example, where the reference model 70 is an HMM and the analysis technique applied is the Viterbi algorithm described above, then the score is simply the likelihood of being predicted by the reference model 70 with respect to the derivative estimate of the alignment map.
[0410] As another example, where the general model 60 is an HMM, the analysis technique may be of the type disclosed by Fariselli et al., supra, which again derives a score which is a measure of similarity 65 .
[0411] The reference model 70 can be generated from a reference sequence of polymer units or from measurements acquired from a reference sequence of polymer units as follows.
[0412] This can be done as follows Figure 19 The process shown generates a reference model 70 from a reference sequence of polymer units 80. This can be used for applications where the reference sequence is known from a library or earlier experiments. Input data representing the reference sequence of polymer units 80 may already be stored in the data processor 5 or may be input thereto.
[0413] The process uses stored emission weights 81, which include emission weights e1 to en for a set of possible types of k-mer states type-1 to type-n. Advantageously, this allows a reference model for any reference sequence of polymer units 80 to be generated based solely on the emission weights 81 for the possible types of k-mer states.
[0414] The process proceeds as follows.
[0415] In step P1, a reference sequence of polymer units 80 is received and from it a reference sequence of k-mer states 73 is generated. This is a simple process to establish, for each k-mer state in the reference sequence, the type of those k-mer states 73 based on the combination 73 of the types of polymer units 80 to which those k-mer states 73 correspond.
[0416] In step P2, a reference model is generated as follows.
[0417] Transition weights 71 are derived for transitions between the reference series of k-mer states 73 derived in step P1. Transition weights 71 take the form defined above with respect to the reference series of k-mer states 73.
[0418] In step P1, an emission weight 72 is derived for each k-mer state 73 in a series of k-mer states 73 by selecting a stored emission weight 81 according to the type of the k-mer state 73. For example, if a given k-mer state 73 is of type Type-4, then emission weight e4 is selected.
[0419] As follows Figure 20The illustrated process generates a reference model 70 from a series of reference measurements 93 collected from a reference sequence of polymer units. This can be used, for example, in applications where both the reference sequence of polymer units and the target polymer are measured simultaneously. In particular, in this embodiment, the identity of the polymer units of the reference sequence itself is not required to be known. A series of reference measurements 93 can be collected by the biochemical analysis system 1 from polymers containing polymer units of the reference sequence.
[0420] The process uses an additional model 90 that processes a series of reference measurements as observations of different possibly similar further series of k-polymer states. This additional model 90 is a model of the biochemical analysis system 1 for collecting a series of reference measurements 93 and can be the same as the general model described above, for example, type 60 disclosed in WO-2013 / 041878. Therefore, the additional model includes a transition weight 91 for each transition between successive k-polymer states in the further series of k-polymer states, which is a transition weight 91 for possible transitions between possible types of k-polymer states; and an emission weight 92 for each type of k-polymer state, which is an emission weight 92 for different measurements observed when the k-polymer state is of that type.
[0421] The process was performed as follows.
[0422] In step Q1, the further model 90 is applied to a series of reference measurements 93 to estimate the reference series of k-mer states 73 as discrete estimated k-mer states. This can be done using the techniques described above.
[0423] In step Q2, a reference model 70 is generated as follows.
[0424] Transition weights 71 are derived for transitions between the reference series of k-mer states 73 derived in step D1 . Transition weights 71 take the form defined above with respect to the reference series of k-mer states 73 .
[0425] In step Q1, an emission weighting 72 is derived for each k-mer state 73 in a series of k-mer states 73 by selecting an emission weighting from the weighting of the further model 50 according to the type of the k-mer state 73. Thus, the emission weighting for each type of k-mer state 73 in the reference model is the same as the emission weighting for the k-mer state 73 of that type in the further model 50.
[0426] We will now describe Figure 7The illustrated methods, and more generally, various applications of the first aspect of the present invention, illustrate the nature of the reference sequence of polymer units, the basis for the determination in step C4, and the potential time savings. In the following examples, the polymer is a polynucleotide, and it is assumed that measuring the first 250 nucleotides and comparing them to the reference sequence will be sufficient to determine (a) whether they are related to the reference sequence and (b) their position relative to the overall sequence. However, this number may be greater or less than this. The number of polymer units required for determination does not necessarily have to be fixed. Typically, measurements will be performed continuously until such a determination is made.
[0427] For each of the application types, there may be Figure 7 This is a slightly different use of the method shown. A mixture of application types can also be used. The analysis performed in step C3 and / or the basis for the decision in step C4 can also be dynamically adjusted as the run progresses. For example, the decision logic may not be present initially, and then applied to the run later when sufficient data has been accumulated to make a decision. Alternatively, the decision logic can be changed during the run.
[0428] In a first class of applications, the reference sequence of polymer units from which reference data 50 is derived is an undesirable sequence, and in step C4 a decision to reject the polymer is made in response to a measure of similarity indicating that the partially shifted polymer is an undesirable sequence.
[0429] This first class of applications has a variety of possible uses. For example, this application can be used to sequence incomplete portions of an organism's genome. If the genome of an organism is partially defined, but the sequence is incomplete, the incomplete portion of the sequence can be determined using the method of the present invention. In this embodiment, the reference sequence can be the sequence of the complete portion of the genome. The polymer can be a fragment of a polynucleotide from the organism. If the measure of similarity indicates that the polymer is a reference sequence (i.e., the sequence of an already defined portion of the genome), the polymer is rejected and a new polymer can be accepted by the nanopore. This can be repeated until the polymer portion that is not similar to the reference sequence is displaced through the nanopore, and this polymer will correspond to the previously undefined portion of the genome and can be retained in the nanopore and fully sequenced. This method allows for rapid sequencing of undefined portions of a genome.
[0430] The first type of application can also be advantageously used to sequence polymers from a polymer sample containing human DNA. Sequencing human DNA presents ethical issues. Therefore, it is useful to be able to sequence a sample of polymers while ignoring sequences that contain human DNA (e.g., identifying bacteria in samples extracted from human patients). In this case, the reference sequence (the undesired sequence) can be the human genome. Any polymers with a measure indicating their similarity to a portion of the human genome can be rejected, while polymers with a measure indicating their similarity to a portion of the human genome can be retained in the nanopore and fully sequenced. Thus, this is an example of a method in which the measure of similarity indicates similarity to a portion of the reference sequence. In this application, the method avoids sequencing human DNA but allows sequencing of bacterial DNA. If bacteria are present in a sample from the human gut, we assume that bacterial DNA (the DNA we want to sequence, or "target" DNA) constitutes approximately 5% of the DNA and that 95% of the DNA in the sample is human DNA ("off-target DNA"). If we assume that approximately 250 base pairs (bp) of sequence per fragment will be sufficient to provide the required measure of similarity, and that polymers can translocate through the pore at a rate of 25 bases per second, then polymers that are not target DNA (i.e., DNA similar to the human DNA reference sequence ("off-target" polymers)) will translocate through the nanopore for approximately 10 seconds before being expelled. Therefore, the relative amount of time that the nanopore contains off-target polymers can be considered to be 95% x 10 = 9.5. On the other hand, assuming that the DNA is fragmented into 10 kB fragments, the amount of time it takes to sequence one fragment of the target DNA will be 10,000 / 25, which is 400 seconds. Therefore, the relative amount of time that the nanopore contains the target polymer can be considered to be 5% x 10 = 9.5. 400, is 20 seconds.So the ratio of the time that wherein nanopore comprises target strand can be considered as the time that wherein nanopore comprises target strand / the time that wherein nanopore comprises off-target strand+the time that wherein nanopore comprises target strand, it is 20 / 29.5.On the other hand, if it is necessary to sequence off-target strand with their overall structure, then the relative amount of the time that wherein nanopore comprises off-target strand will be 95% x 400, which is 380, and so the ratio of the time that nanopore comprises target strand can be considered as 20 / 380.This represents the efficiency of about 13.6 times.
[0431] The first type of application can also be advantageously used to sequence contaminants in a sample. In this embodiment, the reference sequence would be the sequence of a known component present in the sample. For example, this could be used to detect contaminants in food products, such as meat products like beef. In this case, the reference sequence would be the sequence of a polynucleotide from an organism derived from the food (e.g., the genome of that organism). The reference sequence could be the sequence of the cow genome. Any polymers in the sample with a measure indicating their similarity to the cow genome can be rejected, while polymers with a measure indicating their similarity to the cow genome can be retained in the nanopore and fully sequenced. This allows for the rapid and simple definition of the nature of the contaminant without requiring knowledge of the contaminant's identity. This is advantageous compared to prior art methods, such as quantitative PCR, which require knowledge of suspected contaminants. Assuming 99% of the DNA is off-target (meat DNA) and 1% is target (e.g., a contaminant), the method of the present invention would be approximately 29 times more effective than if the nanopore were unable to reject the undesirable polymers.
[0432] In a second class of applications, the reference sequence of polymer units from which reference data 50 is derived is the target, and in step C4 a decision is made to reject the polymer in response to a measure of similarity indicating that the partially displaced polymer is not the target.
[0433] This second type of application can be advantageously used to sequence a gene of interest from a DNA sample. In this application, the reference sequence is the target, which can be a portion of a polynucleotide such as the gene of interest, and the polymers can comprise fragments of polynucleotides, such as DNA, from the sample. Any polymers in the sample with a measure of similarity indicating that they are not similar to the target (gene of interest) can be rejected. The remaining polymers can be retained and sequenced. This allows for rapid sequencing of the gene of interest and is advantageous over existing techniques, which require isolation of the target gene of interest prior to sequencing (e.g., by hybridizing the gene of interest to a probe attached to a solid surface). This isolation technique is time-consuming and unnecessary when using the methods of the present invention. An example of this application would be sequencing the human genome. The human genome contains 50 Mb (megabases) of coding sequence. It would be ideal to be able to sequence this 50 Mb, rather than the remaining 3,000 Mb. Therefore, the amount of "off-target" DNA (to be rejected) is 3,000 Mb. The DNA will be fragmented into fragments of approximately 10 kB in length, and thus 3,000 Mb would represent approximately 300,000 fragments. Assuming that approximately 250 base pairs of sequence per fragment are sufficient to provide the required measure of similarity, and that polymers can be translocated through the pore at a rate of 25 bases per second, polymers that are not similar to the target polymer ("off-target" human DNA) will translocate through the nanopore for approximately 10 seconds before being ejected. Since there are 300,000 off-target fragments, the off-target fragments will be retained in the pore for approximately 3,000,000 seconds per nanopore (the number of fragments multiplied by the time each fragment remains in the pore - approximately 10 seconds). The remaining 50 Mb ("target") that are similar to the target polymer will take 2,000 seconds (the time spent at 25 bases per second is equal to 50,000,000 / 25, or 2,000,000 seconds). The total time to sequence the described 50Mb target polymer is the sum of the time taken to sequence the off-target polymer and the time taken to sequence the target polymer, which is 3,000,000 + 2,000,000, or 5,000,000 seconds per nanopore. On the other hand, if you sequence the entirety of each of the 300,000 off-target fragments, it will take 3,000,000,000 / 25 (3,000Mb at a rate of 25 base pairs / second) + 2,000,000 (the time taken to sequence the target polymer), which is 122,000,000 seconds per nanopore (over 50 times longer) to sequence a single genome.
[0434] This second application can also be advantageously used to identify whether bacteria in a sample (e.g., from a hospitalized patient) are antibiotic-resistant. Here, the reference sequence would be the target, which can be a polynucleotide corresponding to a specific antibiotic-resistance gene. Any polymers in the sample that have a similarity metric to the target antibiotic-resistance gene can be rejected. If no polymers are detected with a similarity metric to the antibiotic-resistance gene, this would indicate that the bacteria are losing the specific antibiotic-resistance gene. Alternatively, if polymers are detected with a similarity metric to the antibiotic-resistance gene, they can be retained and sequenced, and the sequence used to determine whether the antibiotic-resistance gene is functional. In this case, the off-target polymer (the bacterial genome) would be approximately 5000 kB, and the target polymer (the region of the sensing zone) would be approximately 5 kB. Making the same assumptions as above, this means that the method of the present invention will sequence DNA approximately 40 times faster than if the nanopore were unable to eject unwanted polymers.
[0435] This second type of application can also be advantageously used for sequencing total bacterial mRNA. In this case, it is desirable to sequence mRNA, but the sequence of rRNA or tRNA can be ignored. Here, the reference sequence can be an annotated version of the target sequence such as a bacterial genome. The polymer can comprise RNA from a sample of bacteria. Any polymer with a measure of similarity that indicates they are not similar to the target bacterial genome in the sample will be associated with rRNA or tRNA and can be excluded. The remaining polymer will correspond to mRNA and can be sequenced to provide the sequence of total bacterial mRNA. In this case, the target polymer will be mRNA (which is approximately 5% of the total RNA), and the polymer that is off-target will be tRNA and rRNA, which is approximately 95% of the total RNA. Using the same assumptions as those defined above, we expect that sequencing efficiency will increase by approximately 8.4 times.
[0436] This second type of application can also be advantageously used to identify strains for phenotypic or SNP (single nucleotide polymorphism) detection, where the bacterial strain is unknown. For example, in this case, the polymers can be fragments of polynucleotides from a bacterial sample. Initially, polymers are not rejected (no reference sequence is used) and polymers that have translocated through the pore are sequenced, but when sufficient sequence information has been obtained to allow the user to determine the bacterial strain, a reference sequence is then selected. The reference sequence will correspond to the target region of interest and will depend on the bacterial species that has been defined. Once a reference sequence has been defined, any polymers (the target region of interest) that partially translocate through the pore and have a measure of similarity indicating their similarity to the reference sequence are retained and fully sequenced, while other polymers can be rejected. This allows detection of the presence of a phenotype or SNP.
[0437] Similarly, this second type of application would be useful for phenotyping cancer. In this application, the polymers could be fragments of polynucleotides obtained from cancer patients. Initially, the reference sequence could be the target sequence. These target sequences could be polynucleotides, such as the sequences of genes associated with different types of cancer. Any polymers that have a measure of similarity to these target sequences would be retained, while other polymers would be excluded. However, once the type of cancer is identified, the reference sequence could be refined so that it now includes targets with sequences of polynucleotides associated with subtypes of cancer.
[0438] In a third class of applications, the reference sequence of polymer units from which reference data 50 is derived is a sequence of polymer units that has already been measured, and in step C4, a decision to reject the polymer is made in response to a measure of similarity indicating that the partially shifted polymer is a sequence of polymer units that has already been measured.
[0439] This type of application can be used to accurately sequence genomes. Determining the sequence of a genome requires sequencing multiple strands of DNA and, for accuracy, determining a consensus sequence for that portion of DNA. Therefore, polymers corresponding to the same portion of the sequence should be sequenced a sufficient number of times to define an accurate consensus sequence. To this end, the methods of the present invention can be used to quickly and accurately sequence genomes. For example, the polymers can comprise DNA from a sample of the organism whose genome is to be defined. A reference sequence is a portion of DNA for which sufficient measurements have been collected (in this case, sufficient sequence data has been obtained to provide an accurate consensus sequence). Initially, no sequence is excluded. However, once sufficient sequence has been calculated for a portion of the genome to allow calculation of an accurate consensus sequence, that consensus sequence becomes the target (reference sequence). Any polymers that partially displace through the pore and have a measure indicating their similarity to the reference sequence (the portion of DNA for which an accurate consensus sequence has been defined) can be excluded, freeing the nanopore to sequence other portions of the genome for which sufficient information has not yet been collected.
[0440] In a fourth class of applications, the reference sequence of polymer units from which reference data 50 is derived comprises multiple targets, and in step C4 a decision to reject a polymer is made in response to a measure of similarity indicating that the partially displaced polymer is one of the targets.
[0441] This is a counting method that can be used to quantify the proportion of each target polymer in a sample of target polymers. For example, the targets can represent different polymers. As the polymer moieties shift through the nanopore, any polymers that have a measure of similarity that represents their similarity to a reference sequence can be assigned to a "bucket" and the number of polymers that fall into each "bucket" can be quantitatively detected. In this embodiment, once sufficient information is obtained about the polymer to determine whether it has a measure of similarity that represents its similarity to one in the reference sequence, the polymer will be excluded. One example of the use of this technology is to quantify contaminants. For example, the polymers can be a sample of a food such as a beef product. In this case, the reference sequence can contain a target with a sequence found in cow DNA and a target with a sequence found in horse DNA. The proportion of polymers that are similar to the cow DNA target and the proportion of polymers that are similar to horse DNA can be calculated using this method, and this will indicate the level of contamination of the beef product with horse meat.
[0442] Similarly, if the reference sequence used contains targets with sequences found in different bacteria, the technique can be used to determine the proportions of different bacteria present in a sample, such as a sample from an infected patient.
[0443] Figure 16 The method shown results in an alignment map. The method can be applied more generally as follows.
[0444] Figure 21 A method for estimating an alignment mapping between (a) a series of measurements of a polymer comprising a polymer unit and (b) a reference sequence of the polymer unit is shown. The method is performed as follows.
[0445] like Figure 21 As shown, the input to the method may be a series of measurements 12 derived by collecting a series of raw measurements of the sequence of polymer units by the biochemical analysis system 1 and subjecting them to a pretreatment as described above. Alternatively, the input to the method may be a series of raw measurements 11.
[0446] The method uses a reference model 70 of a reference sequence of polymer units, which is stored in the memory 10 of the data processor 5. The reference model 70 takes the same form as described above, treating the measurements as observations of the reference sequence of k-mer states corresponding to the reference sequence of polymer units.
[0447] The reference model 70 is used for the alignment step S1. Specifically, in the alignment step S1, the reference model 70 is applied to the series of measurements 12. The alignment step S1 is performed in the same manner as in the above step C4d. In other words, the reference model 70 is used to replace the general model 60 by using the same method as in the above reference model 70. Figure 13The same technique described for step C4d is used to fit the model to a series of blocks of measurements 63 to provide a measure of similarity 65 of the fit of the reference model 70 to the block of measurements 63 to perform the alignment step S1 .
[0448] Due to the form of the reference model 70, and in particular the representation of transitions between reference series of k-mer states 73, applying the model inherently derives an estimate of the alignment mapping between a series of measurements and a reference series of k-mer states 73. This can be understood as follows. Since the general model 60 represents transitions between possible types of k-mer states, applying the model provides an estimate of the type of k-mer state observed for each measurement, i.e., an estimate of the initial series of k-mer states 34 and a discrete estimated k-mer state 35, each estimate of which is observed for each measurement by the type of k-mer state. Since the reference model 70 represents transitions between reference series of k-mer states 73, applying the reference model 70 instead estimates the k-mer state 73 of the reference series of observations for each measurement, which is an alignment mapping between a series of measurements and the reference series of k-mer states 73.
[0449] Since there is an inherent mapping between the reference series of k-mer states 73 and the polymer units of the reference sequence, the alignment mapping between the series of measurements of k-mer states 73 and the reference series also provides an alignment mapping between the series of measurements of polymer units and the reference sequence.
[0450] Figure 22 An embodiment of an alignment map is shown to illustrate its properties. In particular, Figure 22 The alignment mapping between the polymer units p0 to p7 of the reference sequence, the k-polymer states k1 to k6 of the reference series, and the measured values m1 to m7 is shown. By way of example, in this embodiment, k is three. The horizontal line represents the alignment between the k-polymer states and the measured values, or the alignment of the gaps in the other series in the case of the dashed lines. Therefore, inherently, the polymer units p0 to p7 of the reference sequence as illustrated are aligned to the k-polymer states k1 to k6 of the reference series. The k-polymer state k1 corresponds to and maps to the polymer units p1 to p3, and so on. As for the mapping between the k-polymer states k1 to k6 of the reference series and the measured values m1 to m7: the k-polymer state k1 maps to the measured value m1, the k-polymer state k2 maps to the measured value m2, the k-polymer state k3 maps to the gaps in the series of measured values, the k-polymer state k4 maps to the measured value m3, and the measured values m4 and m5 map to the gaps in the series of k-polymer states.
[0451] Depending on the method applied, the form of the estimate 13 of the alignment map may be changed as follows.
[0452] As described above, the analysis technique applied in the alignment step S1 can take various forms that are suitable for the form of the reference model 70. For example, in the case where the reference model 70 is an HMM, the analysis technique can be a known algorithm for solving HMMs, such as the forward-backward algorithm or the Viterbi algorithm, which are well known in the art. Generally speaking, such an algorithm can avoid brute-force calculation of the likelihood (probability) of all possible paths through the sequence of states, and instead utilize a simplified likelihood-based approach to determine the state sequence.
[0453] By some technique applied in the alignment step S1, the derived estimate 13 of the alignment map comprises, for each measurement 12 in the series, a weighting of the different k-mer states 73 in the reference series with respect to the k-mer state 73. For example, it can be obtained by M i,j Denote this alignment map by where index i denotes the measured value and index j denotes the k-mer state in the reference series, so that when there are K k-mer states, M i,1 To M i,K The value of represents the weight used for the i-th measurement value of each k-mer state 73 in the reference series of k-mer states 73. In this case, the estimated value 13 does not represent a single k-mer state 73 as it is mapped to each measurement value, but instead provides the weights of the different possible k-mer states 73 so mapped to each measurement value.
[0454] As an embodiment in the case where the reference model 70 is an HMM, when the analysis technique applied is the forward-backward algorithm described above, the derived estimated value can be of this type. In the forward-backward algorithm, the total likelihood of all sequences ending with a given k-mer state is calculated in a loop using conversion and emission weighting in the forward and backward directions. Combining these forward and backward probabilities and together with the total likelihood of the data, the probability of each measurement from a given k-mer state is calculated. This probability matrix, called the back matrix, is the estimated value 13 of the alignment map.
[0455] In this case, in a subsequent scoring step S2 (which is optional), there is a score 14 that represents the likelihood that the estimate 13 of the alignment map is correct. This can be derived from the estimate 13 of the alignment map using simple probabilistic techniques, or alternatively can be derived as an intrinsic part of the alignment step S1.
[0456] By means of the other techniques applied in the alignment step S1, the derived estimate 13 of the alignment map comprises, for each measurement in the series, a discrete estimate of the k-mer state in the reference series of k-mer states. For example, such an alignment map may be obtained by Mi Indicates that the index i indicates the measured value and M i The values 1 to K may be used to represent K k-mer states. In this case, the estimate 13 represents a single k-mer state 73 mapped to each measurement value.
[0457] As an embodiment in the case where the reference model 70 is an HMM, when the applied analysis technique is the Viterbi algorithm described above, the derived estimate can be of a type in which the analysis technique estimates the sequence of k-mers based on the likelihood of the model's expectations for a series of measurements produced by a reference series of k-mer states.
[0458] In the case where the derived estimate 13 of the alignment map includes discrete estimates of the k-mer states, the algorithm inherently derives a score 14 representing the likelihood that the estimate of the alignment map is correct, because the algorithm derives the alignment map based on scores for different paths through the model. Therefore, in this case, a separate scoring step S2 is not performed. As an example, where the reference model 70 is an HMM and the analysis technique applied is the Viterbi algorithm described above, the score is simply the likelihood of the derived estimate 13 of the alignment map predicted by the reference model 70.
[0459] Figure 21 The method shown has a wide range of applications, wherein it is expected that the alignment mapping between a series of measurements of its estimation polymer and the reference sequence of polymer units and / or the score representing the accurate likelihood of the alignment mapping. The assessment of this alignment mapping can be used for various applications, such as providing identification or detection of the presence, absence or degree of the polymer in the sample with reference to a comparison, for example, providing diagnosis. The specific application of possible scope is a large amount and can be applied to detect any analyte with a dna sequence.
[0460] The above embodiments relate to a single reference model 70. In various applications, multiple reference models 70 may be used. Figure 21 The illustrated method can be applied using each reference model 70, or one of the reference models 70 can be selected. Depending on the application, the selection can be based on a variety of criteria. For example, a reference model 70 can be applied to different types of sensor devices 2 (e.g., different nanopores) and / or ambient conditions. In such cases, the reference model 8 described below is selected based on the type of sensor device 2 actually used and / or the actual ambient conditions. In another embodiment, the selection can be based on the analyte to be detected, such as a specific G / C enrichment or whether specific epigenetic information is experimentally determined.
[0461] Thus, according to a fourth aspect of the present invention, there is provided a method of estimating an alignment mapping between: (a) a series of measurements of a polymer comprising polymer units, wherein the measurements depend on a k-mer, a k-mer being k polymer units of the polymer, wherein k is an integer, and (b) a reference sequence of polymer units;
[0462] The method uses a reference model that processes reference data as observations of a series of reference k-mer states corresponding to a reference sequence of polymer units, wherein the reference model comprises:
[0463] transition weights for transitions between k-mer states in a reference series of k-mer states; and
[0464] For each k-mer state, emission weights for different measurement values observed when observing the k-mer state; and
[0465] The method includes applying a reference model to a series of measurements to derive an estimate of an alignment mapping between the series of measurements and a reference series of k-mer states corresponding to a reference sequence of polymer units.
[0466] The following features may optionally be applied to the fourth aspect of the present invention in any combination:
[0467] For each measurement in the series, the derived estimate of the alignment map may include a discrete estimate of the k-mer state of the map in the reference series of k-mer states.
[0468] For each measurement in the series, the derived estimate of the alignment map may include a weighting of the k-mer state with respect to the different mapped k-mer states in the reference series.
[0469] The method may further comprise deriving a score representing the likelihood that the estimate of the alignment mapping is correct.
[0470] The method may further include generating a reference model from a reference sequence of polymer units using stored emission weights for a set of possible types of k-mer states by a process comprising:
[0471] deriving a series of k-mer states corresponding to a reference sequence of the received polymer;
[0472] A reference model is generated by generating transition weights for transitions between k-mer states in a derived series of k-mer states and by selecting an emission weight for each k-mer state in the derived series from stored emission weights according to the type of the k-mer state.
[0473] The method may further comprise generating a reference model from a series of reference measurements of a polymer comprising a reference sequence of polymer units.
[0474] The step of generating a reference model may use a further model that processes the series of reference measurements as observations of a further series of k-mer states of different possible types, wherein the further model comprises:
[0475] for each transition between successive k-mer states in the further series of k-mer states, a transition weight for a possible transition between possible types of k-mer states; and
[0476] For each type of k-mer state, the emission weights for the different measurement values observed when the k-mer state is of that type.
[0477] The steps to generate a reference model include:
[0478] generating an estimate of the k-mer state of the reference series by applying the additional model to the series of reference measurements; and
[0479] A reference model is generated by transition weighting the transitions between k-mer states in the reference series of estimates of k-mer states and by emission weighting for each k-mer state in the reference series of estimates selected by a further model according to the type of k-mer state.
[0480] Reference models can be pre-stored.
[0481] One or both of the transition weights and the transmit weights may be probabilities.
[0482] The model may be a Hidden Markov Model.
[0483] The integer k may be a complex number.
[0484] The measurements may be measurements taken during translocation of the polymer through the nanopore.
[0485] The translocation of the polymer through the nanopore may occur in a ratchet manner.
[0486] The nanopore may be a biological pore.
[0487] The polymer may be a polynucleotide, and the polymer units may be nucleotides.
[0488] A single measurement value may depend on a k-mer, or a predetermined plurality of measurements of different properties may depend on the same k-mer.
[0489] The measurements may include one or more of current measurements, impedance measurements, tunneling measurements, field effect transistor measurements, and optical measurements.
[0490] The reference model may be stored in memory.
[0491] Prior to the step of applying the reference model to the series of measurements, the method may further comprise deriving the series of measurements by:
[0492] receiving a series of raw measurements by the aggregate without previously knowing the number of measurements in the group, wherein the series of raw measurement groups of a plurality of raw measurements are dependent on the same k-mer, and
[0493] A series of raw measurements is processed to identify successive groups of measurements and with respect to each identified group a single measurement value or multiple measurements of a different type is derived to form the series of measurements.
[0494] The method may further comprise collecting a series of raw measurements from the polymer.
[0495] In each of the plurality of series of measurements, the plurality of groups of measurements may be dependent on the same k-mer, with the number of measurements in the group being unknown.
[0496] The method may further comprise acquiring the series of measurements from the polymer.
[0497] Sequence Listing
[0498]
[0499]
[0500]
[0501] .
Claims
1. A method for controlling a biochemical analysis system for analyzing a polymer, the polymer comprising a sequence of polymer units, wherein The biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to collect continuous measurements of a polymer from the sensor element during translocation of the polymer through the nanopore of the sensor element, wherein the method comprises analyzing a series of measurements taken from the polymer during partial translocation of the polymer using reference data derived from at least one reference sequence of polymer units to provide a measure of similarity between the sequence of polymer units of the partially translocated polymer and the at least one reference sequence, and Responsive to the measure of similarity, the biochemical analysis system is operated to exclude the polymer and collect measurements from additional polymers.
2. The method according to claim 1, wherein At least one of the sensor elements is operable to expel polymer being translocated through the nanopore, and the step of operating the biochemical analysis system to expel the polymer and collect measurements from additional polymer includes operating the sensor element to expel the polymer from the nanopore and receive additional polymer in the nanopore.
3. The method according to claim 2, wherein: At least one of the sensor elements is operable to expel a polymer being translocated through the nanopore by applying an expulsion bias sufficient to expel the polymer, the step of operating the sensor element to expel the polymer from the nanopore being performed by applying an expulsion bias, and the step of operating the sensor element to receive additional polymer in the nanopore being performed by applying a translocation bias sufficient to enable additional polymer to be translocated therethrough.
4. The method according to claim 1, wherein The biochemical analysis system comprises an array of sensor elements and is operable to collect successive measurements of polymers from selected sensor elements in a multiplexed manner, and the step of operating the biochemical analysis system to exclude the polymer and collect measurements from additional polymers comprises operating the biochemical analysis system to stop collecting measurements from the currently selected sensor element and begin collecting measurements from a newly selected sensor element.
5. The method according to claim 4, wherein The measurements include electrical measurements collected from the sensor elements, and the biochemical analysis system is operable to collect successive measurements of the polymer from selected sensor elements in an electrically multiplexed manner.
6. The method according to claim 5, wherein: The biochemical analysis system comprises: a detection circuit comprising a plurality of detection channels, each capable of acquiring electrical measurements from the sensor elements, the number of sensor elements in the array being greater than the number of detection channels; and A switch arrangement is provided to selectively connect the detection channels to the individual sensor elements in a multiplexed manner.
7. The method according to any one of claims 4 to 7, wherein The sensor element is controllable to expel polymer that is displacing through the nanopore of the sensor element, and the method further comprises, when operating the biochemical analysis system to stop acquiring measurements from the currently selected sensor element, also controlling the currently selected sensor element to expel the polymer and thereby making the nanopore available to receive additional polymer.
8. A method according to any one of the preceding claims, wherein The at least one reference sequence of polymer units is an undesirable sequence, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively manipulating is performed in response to a measure of similarity that indicates that the partially shifted polymer is an undesirable sequence.
9. The method according to any one of claims 1 to 7, wherein The at least one reference sequence of polymer units is a target, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively manipulating is performed in response to a measure of similarity that indicates that the partially shifted polymer is not a target.
10. The method according to any one of claims 1 to 7, wherein The at least one reference sequence of polymer units is a measured sequence of polymer units, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively operating is performed in response to a measure of similarity indicating that the partially shifted polymer is the measured sequence of polymer units.
Citation Information
Patent Citations
Analysis of polymers
CN115851894B
Electrical device with detachable components
GB201418512D0
A miniature support for thin films containing single channels or nanopores and methods for using same
WO2000028312A1
Molecular and atomic scale evaluation of biopolymers
WO2000079257A1
Suspended carbon nanotube field effect transistor
WO2005124888A1