Analysis of polymers

By analyzing the similarity of measured values ​​using reference data of reference sequences in polymer analysis systems, the operating system rejects polymers that do not require further analysis, solving the problem of slow analysis speed in the prior art and achieving a faster measurement process.

CN115851894BActive Publication Date: 2025-06-27OXFORD NANOPORE TECH LTD

Patent Information

Application Number
CN202211448003.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-05-06
Filing Date
2015-10-16
Publication Date
2025-06-27
Estimated Expiration
2035-10-16

AI Technical Summary

Technical Problem

Existing biochemical analysis systems using nanopores are difficult to improve the analysis speed when analyzing polymers, resulting in a longer measurement time.

Method used

The collected measurements are analyzed using reference data derived from the reference sequence of the polymer unit when the polymer part is partially shifted through the nanopore, providing a measure of similarity between the sequence of the polymer and the reference sequence, and operating the system based on similarity measures to reject polymers that do not require further analysis.

Benefits of technology

The rapid repulsion and acquisition of additional polymer measurements without completing the initial measurement of the polymer is achieved, significantly improving the analysis speed and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115851894B_ABST
    Figure CN115851894B_ABST
Patent Text Reader

Abstract

The present invention relates to the analysis of polymers. The present invention also relates to a biochemical analysis system that analyzes a polymer by acquiring measurement values of the polymer during translocation of the polymer through a sensor element including a nanopore. When the polymer has been partially translocated, reference data from a reference sequence is used to analyze a series of measurement values to provide a measure of similarity. In response to the measure of similarity, the sensor element can be selectively operated to eject the polymer and thereby make the nanopore available for receiving additional polymers. The biochemical analysis system includes an array of sensor elements and acquires measurement values from selected sensor elements in a multiplexed manner. In response to the measure of similarity, the biochemical analysis system stops acquiring measurement values from the currently selected sensor element and begins acquiring measurement values from a newly selected sensor element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 201580069073.9, titled "Analysis of Polymers", with an application date of October 16, 2015. Technical Field

[0002] The first to third aspects of the present invention relate to analyzing polymers using a biochemical analysis system that includes at least one sensor element comprising a nanopore. The fourth aspect of the present invention relates to estimating an alignment mapping between a series of measurements of a polymer comprising polymer units and a reference sequence of the polymer units. In all aspects, the polymer can be, for example but not limited to, a polynucleotide, where the polymer units are nucleotides. Background Art

[0003] There are various types of biochemical analysis systems that provide measurements of polymer units for determining sequences. For example, but without limitation, one class of measurement systems uses nanopores. Biochemical analysis systems using nanopores have been the subject of much recent development. Generally, during the translocation of a polymer through a nanopore, consecutive measurements of the polymer are acquired by a sensor element comprising the nanopore. Some characteristics of the system depend on the polymer units in the nanopore, and measurements of such characteristics are acquired. Such measurement systems using nanopores have considerable promise, particularly in the field of polynucleotide sequencing, such as DNA or RNA sequencing.

[0004] Such biochemical analysis systems using nanopores can provide long consecutive reads of polymers, for example, in the case of polynucleotides, in the range from hundreds to tens of thousands (and possibly more) of nucleotides. The data collected in this way includes measurements such as measurements of ionic current, where each translocation of the sequence through the sensitive part of the nanopore results in a slight change in the measured characteristic. Summary of the Invention

[0005] While such biochemical analysis systems using nanopores can provide significant advantages, it is also desirable to increase the analysis speed. The first and second aspects of the present invention relate to providing such an increase.

[0006] According to a first aspect of the present invention, there is provided a method of controlling a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, wherein the biochemical analysis system includes at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to acquire consecutive measurements of the polymer by the sensor element during translocation of the polymer through the nanopore of the sensor element.

[0007] Wherein, the method includes, when a polymer portion is translocating through a nanopore, analyzing a series of measurements of the polymer acquired during the partial translocation of the polymer using reference data from at least one reference sequence derived from polymer units to provide a measure of similarity between the sequence of polymer units of the partially translocated polymer and the at least one reference sequence, and

[0008] operating a biochemical analysis system to reject the polymer and acquire measurements of additional polymers in response to the measure of similarity.

[0009] This method involves analyzing measurements taken from a polymer when a polymer portion is translocating through a nanopore (i.e., during the translocation of the polymer through the nanopore). In particular, a series of measurements of the polymer acquired during the partial translocation are analyzed using reference data from at least one reference sequence derived from polymer units. This analysis provides a measure of similarity between the sequence of polymer units of the partially translocated polymer and the at least one reference sequence. In response to this measure of similarity, if the similarity to the reference sequence indicates that no further analysis of the polymer is needed, for example because the measured polymer is not of interest, the polymer can be rejected to acquire measurements of additional polymers.

[0010] Rejecting the polymer allows measurements of additional polymers to be made without completing the measurements of the initially measured polymer. This provides a time savings in acquiring measurements, as it is done "on-the-fly" (i.e., during the measurement of the polymer). In typical applications, this time savings can be significant, as biochemical analysis systems using nanopores can provide long continuous reads of polymers, and the analysis can identify at an early stage in such a read that no further measurement of the currently measured polymer is needed.

[0011] For example, in a typical application where the polymer is a polynucleotide, sequencing performed at 100% accuracy would allow for a preliminary determination after measuring approximately 30 nucleotides. Thus, considering the practically achievable accuracy, a determination can be made after measuring several hundred nucleotides, typically 250 nucleotides. Compare this to a biochemical analysis system capable of measuring sequences in the range of hundreds to tens of thousands (and potentially more) of nucleotides in length.

[0012] The method potentially provides a significantly faster time to results, where continued measurements are made only on those polymers determined to be of interest and those determined to be uninteresting are rejected. This advantage of reducing the amount of wasted data acquisition is particularly significant for applications that require a large amount of data acquisition. The resulting time savings are useful in themselves or can be used, for example, to obtain greater coverage and thus potentially higher sequencing accuracy with the available time and resources.

[0013] Analysis providing a measure of the similarity between the sequence of polymer units of a partially translocated polymer and at least one reference sequence can itself use known techniques for comparing measurements with a reference. However, contrary to the present method, such known techniques typically make measurements after translocation is complete.

[0014] The method can be applied to a wide variety of applications. Depending on the application, the measure of similarity can represent similarity to the whole of the reference sequence or to a part of the reference sequence.

[0015] According to a second aspect of the invention, there is provided a method of controlling a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, wherein the biochemical analysis system includes at least one sensor element comprising a nanopore and the biochemical analysis system is operable to acquire successive measurements of the polymer by the sensor element during translocation of the polymer through the nanopore of the sensor element,

[0016] wherein the method includes, when the polymer is partially translocated through the nanopore, analyzing a series of measurements acquired from the polymer during the partial translocation of the polymer by a measure of derivation and fitting, the model treating the measurements as observations of a series of k-mer states of different possible types, the model including: transition weighting for possible transitions between possible types of k-mer states, relative to each transition between successive k-mer states in the series of k-mer states; and emission weighting representing the probability of observing a given k-mer for each type of k-mer state, and

[0017] in response to the measure of fitting, operating the biochemical analysis system to reject the polymer and acquire measurements of another polymer.

[0018] This method involves analyzing measurements acquired from the polymer when the polymer is partially translocated through the nanopore, i.e., during translocation of the polymer through the nanopore. In particular, a reference data analysis derived from at least one reference sequence of polymer units is used to analyze a series of measurements acquired from the polymer during the partial translocation. This analysis provides a measure of fitting of the model. In response to this measure of fitting, if the measure of fitting as determined by the model indicates that the measurements have poor quality such that no further translocation and measurement are required, action can be taken to reject the polymer and acquire measurements of another polymer.

[0019] Exclusion polymers allow additional polymer measurements to be taken without completing the measurement of the initial polymer. This provides a time savings in taking measurements because the operation is performed "on-the-fly", i.e., during the taking of polymer measurements. In typical applications, this time savings can be significant because biochemical analysis systems using nanopores can provide long, continuous reads of polymers, even though the analysis can identify early on that a measurement has poor quality.

[0020] The first and second aspects of the invention are the same, except for the basis on which the biochemical analysis system operates to exclude polymers and take additional polymer measurements. Thus, the optional features according to the first aspect of the invention can be applied with the necessary modifications to the second aspect of the invention. Similarly, all of the following features of the method apply equally to the method according to the first or second aspect of the invention.

[0021] Exclusion of polymers can occur in different ways.

[0022] In a first approach, at least one sensor element is operable to expel a polymer that translocates through the nanopore. In this case, the step of operating the biochemical analysis system to exclude polymers and take additional polymer measurements can be performed by operating the sensor element to expel the polymer from the nanopore and receive another polymer in the nanopore.

[0023] In a second approach, the biochemical analysis system includes an array of sensor elements and operates to take consecutive measurements of a polymer by sensor elements selected in a multiplexed manner. In this case, the step of operating the biochemical analysis system to exclude polymers and take additional polymer measurements can include causing the biochemical analysis system to stop taking measurements by the currently selected sensor element and start taking measurements by a newly selected sensor element.

[0024] These two approaches can be used in combination.

[0025] A third aspect of the invention relates to the application of a particular form of biochemical analysis that can be performed using nanopores.

[0026] According to a third aspect of the invention, there is provided a method of classifying polymers, each of which comprises a sequence of polymer units, the method using a system comprising: a sample chamber containing a sample containing a polymer, a collection chamber sealed to the sample chamber, and a sensor element comprising a nanopore that communicates between the sample chamber and the collection chamber,

[0027] The method includes causing consecutive polymers to translocate from the sample chamber through the nanopore, and during the translocation of each polymer:

[0028] Continuous measurements of the polymer are acquired by a sensor element;

[0029] Using reference data from at least one reference sequence of polymer units, a series of measurements acquired from the polymer during a partial shift of the polymer are analyzed to provide a measure of the similarity between the sequence of polymer units of the partially shifted polymer and the at least one reference sequence.

[0030] Based on the measure of similarity, the shift of the polymer into the collection chamber is selectively completed or the polymer is discharged back into the sample chamber.

[0031] Thus, the method utilizes a measure of similarity that is provided by analyzing a series of measurements acquired from the polymer during a partial shift. This analysis can itself use known techniques for comparing the measurements with a reference. However, the measure of similarity is used to determine whether to collect the polymer. If so, then the shift of the polymer into the collection chamber is completed. Additionally, the polymer is discharged back into the sample chamber. In this way, selected polymers are collected into the collection chamber. For example, after completion of the shift of the polymer from the sample, or alternatively, during the shift of the polymer from the sample, the collected polymer can be recovered, for example, by providing a system (with a fluid system suitable for it).

[0032] The method can be applied to a wide variety of applications. For example, the method can be applied to polynucleotides, such as polymers of viral genomes or plasmids. Viral genomes typically have a length on the order of 10 - 15 kB (kilobases), and plasmids typically have a length on the order of 4 kB. In such instances, the polynucleotide does not need to be fragmented and can be collected in its entirety. The collected viral genome or plasmid can be used in any way, such as for transfecting cells.

[0033] The reference sequence of polymer units from the reference data can be a desired sequence. In this case, in response to a measure of similarity indicating that the partially shifted polymer is the desired sequence, the step of selectively completing the shift of the polymer into the collection chamber is performed. However, this is not necessary. In some applications, the reference sequence of polymer units from the reference data can be an undesired sequence. In this case, in response to a measure of similarity indicating that the partially shifted polymer is not the undesired sequence, the step of selectively completing the shift of the polymer into the collection chamber is performed.

[0034] Depending on the application, the measure of similarity can represent similarity to the whole of the reference sequence or to a part of the reference sequence.

[0035] The system may include a plurality of collection chambers, and for each collection chamber, a sensor element including nanopores providing communication between the sample chamber and the respective collection chambers. This allows the method to be carried out with respect to a plurality of parallel nanopores. As well as providing the ability to accelerate the sorting method, it may allow different polymers to be collected in different collection chambers. To achieve this, the reference data and criteria for collection are accordingly selected. In one example, different reference data regarding different nanopores may be used for the method. In another embodiment, the same reference data may be used regarding different nanopores, but the step of selectively effecting the translocation of the polymer into the collection chamber is carried out with different dependencies on the measure of similarity regarding different nanopores.

[0036] According to a further aspect of the invention, there is provided a biochemical analysis system which carries out methods similar to those of the methods of the first, second or third aspects of the invention.

[0037] A fourth aspect of the invention relates to the alignment between a series of measurements of a polymer comprising polymer units and a reference sequence of the polymer units.

[0038] Some types of measurement systems acquire measurements of polymers depending on k-mers, where a k-mer is a group of k polymer units of the polymer and k is an integer. By definition, a group of k polymers will hereinafter be referred to as a k-mer. In general, k may take the value 1, in which case the k-mer is a single polymer unit, or it may be plural (a plural integer). Depending on the nature of the polymer, each given polymer unit may be of a different type. For example, in the case where the polymer is a polynucleotide, the polymer unit is a nucleotide and the different types are nucleotides containing different nucleic acid bases (such as cytosine, guanine, etc.). Thus, for each different combination of different types of each polymer unit corresponding to a k-mer, each given k-mer may also be of a different type.

[0039] For estimating polymer units from measurements, in practical types of measurement systems, it is difficult to provide measurements depending on individual polymer units. Instead, the value of each measurement depends on a k-mer, where k is plural. Conceptually, this may be considered a measurement system with a "blunt reading head" larger than the polymer unit being measured. In this case, the number of different k-mers to be resolved increases to the power of k. When the measurements depend on a large number of polymer units (larger k values), it may be difficult to resolve the measurements taken from different types of k-mers because they provide overlapping signal distributions, especially when considering noise and / or artefacts in the measurement system. This is detrimental to estimating the underlying sequence of the polymer units.

[0040] When k is complex, it may be possible to combine information from multiple measurements of overlapping k-mers (each depending in part on the same polymer units) to obtain a single value resolved at the level of the polymer units. For example, WO-2013 / 041878 discloses a method of estimating the sequence of polymer units in a polymer from at least one series of measurements related to the polymer using a model that processes the measurements using observations of a series of different possible types of k-mers. The model includes: transition weighting, for each transition between consecutive k-mer states in a series of k-mer states, for possible transitions between possible types of k-mer states; and emission weighting, for each type of k-mer state representing the chance of a measurement of a given k-mer being observed. The model can be, for example, a Hidden Markov Model (HMM). Such a model can improve the accuracy of the estimate by taking into account multiple measurements when considering the likelihood predicted by a model of a series of measurements generated by the sequence of polymer units.

[0041] In many cases, it is desirable to estimate an alignment map between a series of measurements of a polymer containing polymer units and a reference sequence of polymer units. Estimation of such an alignment map can be used in various applications, such as comparing to a reference to provide identification or detection of the presence, absence, or degree of a polymer in a sample, for example to provide a diagnosis. The specific applications of possible scope are numerous and can be applied to detecting any analyte having a DNA sequence.

[0042] Existing techniques involve initially estimating the sequence of polymer units that have been measured and then estimating the alignment map to a reference sequence of polymer units by comparing the identity (identity, unity, property) of the polymer units. A variety of fast alignment algorithms have been developed for the case where the polymer units are nucleotides (often referred to as bases in the literature). Examples of fast alignment algorithms are BLAST (Basic Local Alignment Search Tool), FASTA, and HMMER, and their derivatives. Fast alignment algorithms typically look for smaller regions of high similarity, which is a relatively rapid process, and then extend to larger regions of low similarity, which is a slow process. Such algorithms have been applied to situations where they represent the identity of a polymer by providing a similarity score as to whether the measured polymer matches the reference in a minimum time frame. In these types of techniques, the identity of the polymer units in the estimated sequence and the reference sequence is directly compared. When referring to polymer units as bases, this technique can be considered a comparison of "base gaps" compared to a comparison between measurements as "measurement gaps".

[0043] However, this technique has limited accuracy in estimating the alignment mapping, or in other words, has limited discriminative ability. This is because the initial step of estimating the sequence of polymer units inherently causes loss of information on the consistency of the polymer units (with respect to what is present in the measurements themselves).

[0044] There is a desire to provide a method for estimating the alignment mapping that provides increased accuracy compared to this existing technique.

[0045] According to a fourth aspect of the present invention, there is provided a method for estimating the alignment mapping between: (a) a series of measurements of a polymer comprising polymer units, where the measurements depend on k-mers, the k-mers being k polymer units of the polymer, where k is an integer, and (b) a reference sequence of polymer units.

[0046] The method uses a reference model that processes the measurements as observations of a reference series of k-mer states corresponding to the reference sequence of polymer units, where the reference model includes:

[0047] Transition weights for transitions between k-mer states in the reference series of k-mer states; and

[0048] Emission weights for different measurements when observing a k-mer state, for each k-mer state; and

[0049] The method includes applying the reference model to the series of measurements to derive an estimate of the alignment mapping between the series of measurements and the reference series of k-mer states corresponding to the reference sequence of polymer units.

[0050] The method thus uses the reference model with respect to the reference sequence. The reference model processes the measurements as a reference series of k-mer states corresponding to the reference sequence of polymer units and includes transition weights for transitions between k-mer states in the reference series of k-mer states; and emission weights for different measurements when observing a k-mer state, for each k-mer state. They can be, but are not limited to, HMMs. As a result, compared to the known techniques discussed above that involve initially estimating the sequence of polymer units and then estimating the alignment mapping to the reference sequence of polymer units by comparing the consistency of the polymer units, this method can improve the evaluation accuracy of the alignment method. This is for the following reasons.

[0051] Generally speaking, the use of the reference model is similar to the model for estimating the sequence of polymers disclosed in WO-2013 / 041878. For example, it uses a conversion weighting sum and an emission weighting of a similar form, and applies the same mathematical processing to the model. However, the reference model itself is different from the model disclosed in WO-2013 / 041878. The model disclosed in WO-2013 / 041878 is a generic model of a measurement system, where each k-mer state can generally have any one of the possible types of k-mer states. Therefore, for various possible conversions between the possible types of k-mer states, a conversion weighting is provided for each conversion between consecutive k-mer states in a series of k-mer states. In contrast, the reference model used in this method is a model of the k-mer states of a reference series corresponding to the reference sequence of polymer units. Therefore, a conversion weighting is provided for the conversion between k-mer states in the k-mer states of the reference series.

[0052] This similarity means that the method of the present invention can utilize the power of the model disclosed in WO-2013 / 041878. The information about the consistency of polymer units (present in the measurements depending on overlapping k-mers) is used to report the product. Due to the different nature of the reference model itself, applying the reference model can provide an alignment mapping between a series of measurements and the k-mer states of a reference series corresponding to the reference sequence of polymer units, and thus provide an alignment mapping between a series of measurements of polymer units and the reference sequence.

[0053] In some embodiments, for each measurement in the series, the estimated value of the derived alignment mapping may include a discrete estimated value of the mapped k-mer state in the k-mer states of the reference series. As an example where the model is an HMM, it can be achieved by using the Viterbi algorithm to derive the estimated value of the alignment mapping.

[0054] In other embodiments, for each measurement in the series, the estimated value of the derived alignment mapping may include a weighting of the different mapped k-mer states in the k-mer states of the reference series. As an example where the model is an HMM, it can be achieved by using the forward-backward algorithm to derive the estimated value of the alignment mapping.

[0055] Optionally, the method may further include deriving a score (representing the likelihood that the estimated value of the alignment mapping is correct). This score provides a measure of the similarity between the measured polymer and the reference sequence of polymer units. By providing information about the consistency of the measured polymer compared to the reference sequence, it can be used for a variety of applications.

[0056] In some cases, the model can be directly applied to derive this score. An example of this is when the model is an HMM and the Viterbi algorithm is applied.

[0057] In other cases, where the estimated alignment mapping can include, for each measurement in the series, a weighting of the k-mer states of different mappings in the k-mer state of the reference series, the score can be derived from those weightings themselves.

[0058] The source of the reference model can vary according to the application.

[0059] In some applications, a reference model previously generated from measurements taken from a reference sequence of polymer units or from a reference sequence of a polymer can be pre-stored.

[0060] In other applications, the reference model can be generated, for example, as follows when performing the method.

[0061] In a first instance, the reference model can be generated from a reference sequence of polymer units. This can be used, for example, in applications where the reference sequence is known from a database or from earlier experiments.

[0062] In this case, the generation of the reference model can be performed using stored emission weightings for a set of possible types of k-mer states. Advantageously, this allows the generation of a reference model for any reference sequence of polymer units based only on stored data involving the emission weightings for the possible types of k-mer states.

[0063] For example, the reference model can be generated by a process including: deriving a series of k-mer states corresponding to the reference sequence of the received polymer; and generating a transition weighting for transitions between the k-mer states in the derived series of k-mer states, and generating the reference model by selecting, according to the type of k-mer state, an emission weighting for each k-mer state in the derived series from the stored emission weightings.

[0064] In a second instance, the reference model can be generated from a series of reference measurements of a polymer containing a reference sequence of polymer units. This can be used, for example, in applications where the reference sequence of the polymer units is measured simultaneously with a target polymer. In particular, in this instance, no consistency of the polymer units with a reference sequence known per se is required.

[0065] For example, a reference model can be generated by using the method of another model that processes a series of reference measurements as observations of a further series of k-mer states of different possible types, where the other model includes: for each transition between consecutive k-mer states in the further series of k-mer states, a transition weight for a possible transition between possible types of k-mer states; and for each type of k-mer state, an emission weight for different measurements for observation when the k-mer state is of that type. This other model itself can be of the model type disclosed in WO-2013 / 041878. In this case, a reference model can be generated by a process including: generating an estimate of a reference series of k-mer states by applying the other model to a series of reference measurements; and generating a reference model by generating transition weights between k-mer states in the generated estimate of the reference series of k-mer states and by selecting emission weights for each k-mer state in the generated estimate of the reference series according to the type of k-mer state by weighting of a further model.

[0066] The generation of the model can be part of a larger framework of model training that examines a large collection of reference measurements derived from observing a large collection of series of k-mer states to find unknown parameters of a mathematical model, such as emission and transition weights. Typically, when the model includes latent (hidden) variables, the expectation-maximisation (EM) algorithm can be used to find maximum likelihood estimates. In the specific case of an HMM, the Baum-Welch algorithm can be used. This algorithm is iterative: an initial guess is made for the model parameters and updates are applied by examining a set of training measurements. Applying the resulting HMM to a second distinct set of measurements will yield improved results (assuming the second set can be described by the same model as the training data).

[0067] According to a further aspect of the invention, there is provided an electronic computer program capable of implementing the method according to the fourth aspect of the invention, or an analysis system implementing the method according to the fourth aspect of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] For better understanding, embodiments of the present invention will now be described by way of non-limiting examples with reference to the accompanying drawings, in which:

[0069] Figure 1 is a schematic diagram of a biochemical analysis system;

[0070] Figure 2 is a cross-sectional view of a sensor device of the system;

[0071] Figure 3Schematic diagram of the sensor element of the sensor device;

[0072] Figure 4 Graph of the signal of the event measured over time by the measurement system;

[0073] Figure 5 Block diagram of the electronic circuit of the system in the first arrangement;

[0074] Figure 6 Block diagram of the electronic circuit of the system in the second arrangement;

[0075] Figure 7 Flowchart of the method for controlling a biochemical analysis system to analyze a polymer;

[0076] Figure 8 Flowchart of the state detection step;

[0077] Figure 9 Detailed flowchart of an example of the state detection step;

[0078] Figure 10 Graph of a series of raw measurements and a series of measured values obtained through the state detection step;

[0079] Figure 11 Flowchart of an alternative method for controlling a biochemical analysis system to analyze a polymer;

[0080] Figure 12 Flowchart of the method for controlling a biochemical analysis system to classify a polymer;

[0081] Figures 13 to 16 Flowchart of different methods for analyzing different forms of reference data;

[0082] Figure 17 State diagram of an example of the k-mer state of a reference series;

[0083] Figure 18 State diagram of the k-mer state of a reference series illustrating the possible types of transitions between k-mer states;

[0084] Figure 19 Flowchart of the first process for generating a reference model;

[0085] Figure 20 Flowchart of the second process for generating a reference model; and

[0086] Figure 21 Flowchart of the method for estimating an alignment map; and

[0087] Figure 22 Block diagram of the alignment map. Detailed Description

[0088] In the described embodiments, a variety of nucleotide and amino acid sequences can be used. In particular:

[0089] SEQ ID NO:1 is a nucleotide sequence that encodes pore MS-(B1)8 (= MS-(D90N / D91N / D93N / D118R / D134R / E139K)8);

[0090] SEQ ID NO:2 is an amino acid sequence that encodes pore MS-(B1)8 (= MS-(D90N / D91N / D93N / D118R / D134R / E139K)8);

[0091] SEQ ID NO:3 is a nucleotide sequence that encodes pore MS-(B2)8 (= MS-(L88N / D90N / D91N / D93N / D118R / D134R / E139K)8);

[0092] SEQ ID NO:4 is an amino acid sequence that encodes pore MS-(B2)8 (= MS-(L88N / D90N / D91N / D93N / D118R / D134R / E139K)8). Except for the mutation L88N, the amino acid sequence of B2 is the same as that of B1;

[0093] SEQ ID NO:5 is a sequence for wild-type Escherichia coli exonuclease I (WT EcoExo I), a preferred polynucleotide handling enzyme;

[0094] SEQ ID NO:6 is a sequence for Escherichia coli exonuclease III, a preferred polynucleotide handling enzyme;

[0095] SEQ ID NO:7 is a sequence for Thermus thermophilus RecJ, a preferred polynucleotide handling enzyme;

[0096] SEQ ID NO:8 is a sequence for λ phage exonuclease, a preferred polynucleotide handling enzyme; and

[0097] SEQ ID NO:9 is a sequence for Phi29 DNA polymerase, a preferred polynucleotide handling enzyme.

[0098] The various features described below are examples and not restrictive. Similarly, the described features need not be applied together and can be applied in any combination.

[0099] First, the nature of the polymers to which the present invention can be applied will be described.

[0100] A polymer comprises a sequence of polymer units. Depending on the nature of the polymer, each given polymer unit can be of a different type (or species (identity)).

[0101] The polymer can be a polynucleotide (or nucleic acid), a polypeptide such as a protein, a polysaccharide, or any other polymer. The polymer can be natural or synthetic. The polymer units can be nucleotides. Nucleotides can be of different types containing different nucleic acid bases.

[0102] The polynucleotide can be deoxyribonucleic acid (DNA), ribonucleic acid (RNA), cDNA, or a synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains. The polynucleotide can be single-stranded, double-stranded, or contain single-stranded and double-stranded regions. Typically, cDNA, RNA, GNA, TNA, or LNA is single-stranded.

[0103] The nucleotides can be of any type. Nucleotides can be naturally occurring or artificial. Nucleotides typically contain a nucleic acid base (which can be abbreviated herein as “base”), a sugar, and at least one phosphate group. Nucleic acid bases are typically heterocyclic. Suitable nucleic acid bases include purines and pyrimidines and more specifically adenine, guanine, thymine, uracil, and cytosine. The sugar is typically a pentose. Suitable sugars include, but are not limited to, ribose and deoxyribose. Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain a monophosphate, diphosphate, or triphosphate.

[0104] Nucleotides can include damaged bases or epigenetic bases. Nucleotides can be labeled or modified to act as markers with distinct signals. Such techniques can be used to identify absent bases, e.g., abasic units or spacers in a polynucleotide.

[0105] When considering measurements of modified or damaged DNA (or similar systems), of particular use are methods in which complementary data are considered. The additional information provided allows for discrimination between a larger number of base states.

[0106] The polymer can also be a class of polymers other than polynucleotides, some non-limiting examples of which are as follows.

[0107] The polymer can be a polypeptide, in which case the polymer units can be naturally occurring or synthetic amino acids.

[0108] The polymer can be a polysaccharide, in which case the polymer units can be monosaccharides.

[0109] In particular, when the biochemical analysis system 1 includes a nanopore and the polymer comprises a polynucleotide, the polynucleotide can be long, for example at least 5 kB (kilobases), i.e., at least 5,000 nucleotides, or at least 30 kB (kilobases), i.e., at least 30,000 nucleotides.

[0110] As used herein, the term 'k-mer' refers to a group of k polymer units, where k is a positive integer, including the case where k is 1, in which case the k-mer is a single polymer unit. In some instances, when referring to k-mers (where k is plural), the k-mers are a subgroup of k-mers and generally do not include the case where k is 1.

[0111] Thus, for each different combination of different types of each polymer unit corresponding to a k-mer, each given k-mer can also have a different type.

[0112] Figure 1 A biochemical analysis system 1 for analyzing polymers is shown, which can also be used for classifying polymers. Turning to Figure 1 , the biochemical analysis system 1 includes a sensor device 2 connected to an electronic circuit 4, which in turn is connected to a data processor 6.

[0113] Some examples will first be described, where the sensor device 2 includes an array of sensor elements each including a biological nanopore.

[0114] In a first form, the sensor device 2 can have a configuration as shown in the cross-section of Figure 2 , which includes a body 20 in which an array of wells 21 are formed, each of which is a recess having a sensor electrode 22 disposed therein. A large number of wells 21 are provided to optimize the data collection rate of the system 1. Generally, any number of wells 21 can be present, typically 256 or 1024, but only a few wells 21 are shown in Figure 2 . The body 20 is covered by a lid 23 that extends over the body 20 and is hollow to define a sample chamber 24 to which each well 21 opens. A common electrode 25 is disposed within the sample chamber 24. In this first form, the sensor device 2 can be the device described in further detail in WO-2009 / 077734, the teachings of which can be applied to the biochemical analysis system 1 and are incorporated herein by reference.

[0115] In a second form, the sensor device 2 can have a configuration described in detail in WO-2014 / 064443, the teachings of which can be applied to the biochemical analysis system 1 and are incorporated herein by reference. In this second form, the sensor device 2 has a configuration generally similar to the first form, including an array of compartments generally similar to the wells 21, but they have a more complex configuration and each of them includes a sensor electrode 22.

[0116] To facilitate the collection of a sample from the collection chamber, the sensor device can be arranged such that the collection chamber 21 can be detached from the respective electrodes 22 below to expose the sample contained therein. Such a device configuration is described in more detail in UK patent application number 1418512.8.

[0117] The sensor device 2 is prepared to form an array of sensor elements 30, Figure 3 one of which is schematically shown therein. Each sensor element 30 is fabricated by forming a membrane 31 that traverses the respective grooves 21 in the first form of the sensor device 2 or the respective compartments in the second form of the sensor device 2, and then embedding holes 32 into the membrane 31. The membrane 31 seals the respective grooves 21 from the sample chamber 24. The membrane 31 can be made of an amphiphilic molecule, such as a lipid.

[0118] The holes 32 are biological nanopores. The holes 32 communicate the sample chamber 24 and the grooves 21 in a known manner.

[0119] For the first form of the sensor device 2, such preparation can be carried out using the techniques and materials described in detail in WO-2009 / 077734, or for the second form of the sensor device 2 using the techniques and materials described in detail in WO-2009 / 077734.

[0120] Each sensor element 30 is capable of operating to acquire electrical measurements of a polymer during the translocation of the polymer 33 through the hole 32 using the sensor electrodes 22 and the common electrode 25 for each sensor element 30. The translocation of the polymer 33 through the hole 32 generates a characteristic signal of the measurement characteristics that can be observed and can generally be referred to as an "event".

[0121] In this example, the pores are biological pores, which can have the following characteristics.

[0122] The biological pores can be transmembrane protein pores. The transmembrane protein pores used in the methods described herein can be derived from β-barrel pores or α-helical bundle pores. β-barrel pores include barrels or channels formed by β-strands. Suitable β-barrel pores include, but are not limited to, α-toxins such as α-hemolysin, anthrax toxin, and leukocidin, as well as outer membrane proteins / porins of bacteria such as Mycobacterium smegmatis porin (Msp) such as MspA, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP). α-helical bundle pores include barrels or channels formed by α-helices. Suitable α-helical bundle pores include, but are not limited to, inner membrane proteins and outer membrane proteins, such as WZA and ClyA toxins. The transmembrane pores can be derived from Msp or from α-hemolysin (α-HL).

[0123] Suitable transmembrane protein pores can be derived from Msp, preferably from MspA. Such pores are oligomeric and typically contain 7, 8, 9, or 10 monomers derived from Msp. The pores can be homo-oligomeric pores derived from Msp containing the same monomers. Alternatively, the pores can be hetero-oligomeric pores derived from Msp, which contain at least one monomer different from the other monomers. The pore can also contain one or more constructs that contain two or more covalently linked monomers derived from Msp. Suitable pores are described in WO-2012 / 107778. The pores can be derived from MspA or its homologs or paralogs.

[0124] Biological pores can be naturally occurring pores or can be mutant pores. Typical pores are described in the following: Stoddart D et al., Proc Natl Acad Sci, 12; 106(19):7702-7, Stoddart D et al., Angew Chem Int Ed Engl. 2010; 49(3):556-9, Stoddart D et al., Nano Lett. 2010 Sep 8; 10(9):3633-7, Butler TZ et al., Proc Natl Acad Sci 2008; 105(52):20647-52, and WO-2012 / 107778.

[0125] The biological pore can be MS-(B1)8. The nucleotide sequences encoding the amino acid sequences of B1 and B1 are Seq ID:1 and Seq ID:2.

[0126] The biological pore is more preferably MS-(B2)8. The amino acid sequence of B2 is the same as that of B1 except for the mutation L88N. The nucleotide sequence encoding B2 and the amino acid sequence of B2 are SeqID:3 and Seq ID:4.

[0127] The biological pore can be embedded in a membrane, such as an amphiphilic layer, for example, a lipid bilayer. The amphiphilic layer is a layer formed by amphiphilic molecules having hydrophilic and lipophilic properties, such as phospholipids. The amphiphilic layer can be a monolayer or a bilayer. The amphiphilic layer can be a co-block polymer as disclosed by (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450) or as disclosed by PCT / GB2013 / 052767 as WO2014 / 064444. Alternatively, the biological pore can be inserted into a solid state layer.

[0128] The pore 32 is an example of a nanopore. More generally, the sensor device 2 can have any form that includes at least one sensor element 30 that is operable to acquire measurements of the polymer during translocation of the polymer through the nanopore.

[0129] Nanopores are typically pores with nanoscale dimensions that permit polymers to pass through. Measurements can be made depending on the characteristics of the polymer units translocating through the pore. The characteristics can be related to the interaction between the polymer and the nanopore. The interaction of the polymer can occur in the constricted region of the nanopore. The biochemical analysis system 1 measures the characteristics, generating measurements that depend on the polymer units of the polymer.

[0130] Alternatively, the nanopore can be a solid-state pore that comprises a pore formed in a solid-state layer. In this case, it can have the following characteristics.

[0131] This solid-state layer typically does not have a biological origin. In other words, the solid-state layer is generally not derived from or isolated from a biological environment, such as an organism or cell, or a synthetically manufactured form of a biologically available structure. The solid-state layer can be formed from both organic and inorganic materials, which include, but are not limited to, microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamides, plastics such as or elastomers such as two-component addition-cured silicone rubber, and glass. The solid-state layer can be formed from graphene. Suitable graphene is disclosed in WO-2009 / 035647 and WO-2011 / 046706.

[0132] When the solid-state pore is a cavity in a solid-state layer, the cavity can be chemically or otherwise modified to enhance its characteristics as a nanopore.

[0133] The solid-state pore can be used in conjunction with additional components, where the additional element provides alternative or additional measurements of the polymer, such as tunneling electrodes (Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1):279-85), or field effect transistor (FET) devices (WO-2005 / 124888). The solid-state pores can be formed by known methods, including, for example, those described in WO 00 / 79257.

[0134] In such as Figure 1In an example of the biochemical analysis system 1 shown, the measured value is an electrical measured value, in particular a current measured value of an ionic current flowing through the pore 32. In general, these and other electrical measurements can be made using, for example, the standard single-channel recording devices described in Stoddart D et al., Proc Natl Acad Sci, 12; 106(19):7702-7, Lieberman KR et al., J Am Chem Soc. 2010; 132(50):17961-72 and WO-2000 / 28312. Alternatively, the electrical measurements can be made using, for example, the multi-channel systems described in WO-2009 / 077734 and WO-2011 / 067559.

[0135] To allow for the acquisition of measured values as the polymer translocates through the nanopore 32, the translocation rate can be controlled by a polymer-binding moiety. Typically, with or in response to an applied field, this moiety can cause the polymer to translocate through the pore 32. The moiety can be a molecular motor that uses, for example, enzymatic activity in the case where the moiety is an enzyme, or can act as a molecular brake. In the case where the polymer is a polynucleotide, a variety of methods have been proposed to control the translocation rate, including the use of polynucleotide-binding enzymes. Suitable enzymes for controlling the translocation rate of polynucleotides include, but are not limited to, polymerases, helicases, exonucleases, single-stranded and double-stranded binding proteins, and topoisomerases such as gyrase. For other polymer types, a moiety that interacts with that polymer type can be used. The polymer-interacting moiety can be any of those disclosed in WO-2010 / 086603, WO-2012 / 107778, and Lieberman KR et al., J Am Chem Soc. 2010; 132(50):17961-72) and any of those disclosed for voltage-gated schemes (Luan B et al., Phys Rev Lett. 2010; 104(23):238103).

[0136] The polymer-binding moiety can be used in a variety of ways to control polymer movement. With or in response to an applied field, this moiety can cause the polymer to translocate through the pore 32. The moiety can act as a molecular motor that uses, for example, enzymatic activity in the case where the moiety is an enzyme, or as a molecular brake. The translocation of the polymer can be controlled by a molecular ratchet that controls the translocation of the polymer through the pore. The molecular ratchet can be a polymer-binding protein.

[0137] For polynucleotides, the polynucleotide-binding protein is preferably a polynucleotide-processing enzyme. A polynucleotide-processing enzyme is a polypeptide that is capable of interacting with a polynucleotide and modifying at least one property of the polynucleotide. The enzyme can modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as dinucleotides or trinucleotides. The enzyme can modify the polynucleotide by directing it or translocating it to a specific location. The polynucleotide-processing enzyme does not need to display enzymatic activity as long as it can bind to the target polynucleotide and control its translocation through the pore. For example, the enzyme can be modified to remove its enzymatic activity, or it can be used under conditions that prevent it from acting as an enzyme. Such conditions are discussed in more detail below.

[0138] The polynucleotide-processing enzyme can be derived from a nucleolytic enzyme. The polynucleotide-processing enzyme used to construct the enzyme is more preferably derived from a member of any one of enzyme classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31. The enzyme can be any one of those disclosed in WO-2010 / 086603.

[0139] Preferred enzymes are polymerases, exonucleases, helicases, and topoisomerases, such as gyrase. Suitable enzymes include, but are not limited to, exonuclease I from Escherichia coli (Seq ID:5), exonuclease III enzyme from Escherichia coli (SeqID:6), RecJ from Thermus thermophilus (Seq ID:7), and bacteriophage λ exonuclease (Seq ID:8) and variants thereof. Three subunits containing the sequence shown in Seq ID:8 or variants thereof interact to form a trimeric exonuclease. The enzyme is preferably derived from Phi29 DNA polymerase. Enzymes derived from Phi29 polymerase include the sequences shown in Seq ID:9 or variants thereof.

[0140] Variants of Seq IDs:5, 6, 7, 8, or 9 are enzymes that have an amino acid sequence that varies from the amino acid sequences in Seq IDs:5, 6, 7, 8, or 9 and retain the polynucleotide-binding ability. The variant can include modifications that facilitate polynucleotide binding and / or enhance its activity at high salt concentrations and / or at room temperature.

[0141] For the entire length of the amino acid sequences of Seq IDs: 5, 6, 7, 8, or 9, the variant will preferably be at least 50% homologous to the above sequences based on amino acid identity. More preferably, for the entire sequence, based on amino acid identity, the variant polypeptide can be at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, and more preferably at least 95%, 97%, or 99% homologous to the amino acid sequences of Seq IDs: 5, 6, 7, 8, or 9. For a stretch of 200 or more, such as 230, 250, 270, or 280 or more adjacent amino acids, there can be at least 80%, such as at least 85%, 90%, or 95% amino acid identity (“hard homology”). Homology is determined as described above. The variant can differ from the wild-type sequence in any of the ways discussed above with reference to SEQ ID NO: 2. As discussed above, the enzyme can be covalently attached to the pore.

[0142] A suitable strategy for single-stranded DNA sequencing is to utilize or target the applied potential to translocate DNA through pore 32, cis-to-trans and trans-to-cis. The most favorable mechanism for strand sequencing is the controlled translocation of single-stranded DNA through pore 32 under an applied potential. An exonuclease acting gradually or continuously on double-stranded DNA can be used on the cis side of the pore to feed the remaining single strand through under an applied potential or on the trans side under a reverse potential. Similarly, a helicase that unwinds double-stranded DNA can also be used in a similar manner. It is also possible for sequencing applications that require strand translocation in response to an applied potential, but the DNA must first be “captured” by the enzyme in the reverse or no potential. After binding, by switching back the potential, the strand will translocate cis-to-trans through the pore and be held in an extended conformation by the current. A single-stranded DNA exonuclease or a single-stranded DNA-dependent polymerase can act as a molecular motor to pull back the most recently translocated single strand through the pore in a controlled stepwise manner, trans-to-cis, in response to an applied potential. Alternatively, a single-stranded DNA-dependent polymerase can act as a molecular brake that slows down the translocation of the polynucleotide through the pore. Any of the parts, techniques, or enzymes described in WO-2012 / 107778 or WO-2012 / 033524 can be used to control polymer movement.

[0143] Generally, when the measurement is a current measurement of the ion flow flowing through pore 32, the ion flow can typically be a DC ion flow, but in principle, an AC current can alternatively be used (i.e., the amplitude of the AC current flowing under an applied AC voltage).

[0144] The biochemical analysis system 1 can collect electrical measurements of types other than the current measurement of the ion flow through the nanopore described above.

[0145] Other possible electrical measurement values include: current measurement values, impedance measurement values, tunneling measurement values (e.g., as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1):279-85), and field effect transistor (FET) measurement values (e.g., as disclosed in WO2005 / 124888).

[0146] As an alternative to electrical measurement values, the biochemical analysis system 1 can acquire optical measurement values. Appropriate optical methods are disclosed in J. Am. Chem. Soc. 2009, 131 1652-1653, which relate to the measurement of fluorescence.

[0147] The measurement system 8 can acquire electrical measurement values of types other than the current measurement value of the ion current through the above-mentioned nanopore. Possible electrical measurement values include: current measurement values, impedance measurement values, tunneling measurement values (e.g., as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1):279-85), and field effect transistor (FET) measurement values (e.g., as disclosed in WO2005 / 124888).

[0148] Optical measurement can be combined with electrical measurement (Soni GV et al., Rev Sci Instrum. 2010 Jan; 81(1):014301).

[0149] The biochemical analysis system 1 can acquire measurement values of different natures simultaneously. The measurement values can have different natures because they are measurement values of different physical characteristics, which can be any of those mentioned above. Alternatively, the measurement values can have different natures because they are measurement values of the same physical characteristic under different conditions, such as electrical measurement values like current measurement values under different bias voltages.

[0150] The typical form of the signal output as a series of raw measurement values 11 by a variety of types of sensor devices 2 is a "noisy staircase wave", but is not limited to this type of signal. For the case of ion current measurement values obtained using a class of measurement systems 8 containing nanopores, Figure 4 An example of a series of raw measurement values 11 having this form is shown.

[0151] Typically, each measurement collected by the biochemical analysis system 1 depends on a k-mer, which is a k-polymer unit of the respective sequences of polymer units, where k is a positive integer. Although ideally the measurement would depend on a single polymer unit (i.e., where k is 1), for many typical types of biochemical analysis systems 1, each measurement depends on a k-mer of multiple polymer units (i.e., where k is a plural number). That is, each measurement depends on the sequence of each polymer unit in the k-mer, where k is a plural number.

[0152] In a series of measurements collected by the biochemical analysis system 1, consecutive groups of multiple measurements depend on the same k-mer. The multiple measurements in each group have a constant value, undergo some variations discussed below, and thus form a "level" in the series of raw measurements. Such a level can typically be formed by measurements that depend on the same k-mer (or consecutive k-mers of the same type) and thus corresponds to the normal state of the biochemical analysis system 1.

[0153] The signal moves between a set of levels (which can be a larger set). Given the sampling rate of the instrument and the noise on the signal, it can be considered that the transition between levels is instantaneous, and thus the signal can be approximated by an idealized step trace.

[0154] The measurements corresponding to each state are constant on the time scale of the event, but for most types of biochemical analysis systems 1, they will undergo variations over a short time range. The variations can be due to measurement noise, such as that generated from circuits and signal processing, especially in the particular case of electrophysiology from amplifiers. This measurement noise is inevitable due to the small magnitude of the characteristics to be measured. The variations can also come from intrinsic variations or fluctuations in the underlying physical or biological system of the biochemical analysis system 1. Most types of biochemical analysis systems 1 will undergo such intrinsic variations to a greater or lesser extent. For any given type of biochemical analysis system 1, both sources of variation can act or one of these noise sources can be dominant.

[0155] Additionally, typically, there is no existing knowledge of the number of measurements in a group, which varies unpredictably.

[0156] The above two factors of variation and the lack of knowledge of the number of measurements can make it difficult to distinguish some groups, for example, in cases where the groups are short and / or the levels of the measurements of two consecutive groups are close to each other.

[0157] Due to the physical or biological processes occurring in the biochemical analysis system 1, a series of raw measurements can take this form. Thus, in some cases, each group of measurements can be referred to as a "state".

[0158] For example, in some types of biochemical analysis systems 1, an event consisting of the translocation of a polymer through a pore 32 can occur in a ratchet manner. During each step of the ratchet movement, at a given voltage across the pore 32, the ionic current flowing through the nanopore is constant and undergoes the variations discussed above. Thus, each set of measurements is associated with a step of the ratchet movement. Each step corresponds to a state in which the polymer is in a corresponding position relative to the nanopore 32. Although during a state, there may be some variations in terms of the exact position, between states, there is a large-scale translocation of the polymer. Depending on the nature of the biochemical analysis system 1, the state can occur due to a binding event in the nanopore.

[0159] The duration of a single state can depend on various factors, such as the potential applied across the pore, the type of enzyme used to ratchet the polymer, regardless of whether the polymer is pushed or pulled through the pore by the enzyme present, pH, salt concentration, and the type of nucleoside triphosphates. The duration of a state can typically vary between 0.5 ms and 3 s, depending on the biochemical analysis system 1, and for any given nanopore system, there is some random variation between states. For any given biochemical analysis system 1, the expected distribution of the duration can be determined experimentally.

[0160] It can be experimentally examined to what extent a given biochemical analysis system 1 provides measurements that depend on the k-mer and the size of the k-mer. Possible approaches for this are disclosed in WO-2013 / 041878.

[0161] Returning to the biochemical analysis system 1, it can acquire types of electrical measurements other than current measurements of the ionic current through the nanopore described above.

[0162] Other possible electrical measurements include: current measurements, impedance measurements, tunneling effect measurements (e.g., as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12; 11(1):279-85), and field effect transistor (FET) measurements (e.g., as disclosed in WO2005 / 124888).

[0163] Returning to Figure 1 , the arrangement of the electronic circuit 4 will now be discussed. The electronic circuit 4 is connected to the sensor electrodes 22 regarding each sensor element 30 and is connected to the common electrode 25. The electronic circuit 4 can have an overall arrangement as described in WO 2011 / 067559. The electronic circuit 4 is arranged as follows to control the application of the bias voltage across each sensor element 3 and to acquire measurements by each sensor element 3.

[0164] In Figure 5Shows a first arrangement for an electronic circuit 4, which shows components with respect to a single sensor element 30, and this assembly is replicated for each sensor element 30. In this first arrangement, the electronic circuit 4 includes a detection channel 40 each connected to a sensor electrode 22 of the sensor element 30 and a bias control circuit 41.

[0165] The detection channel 40 acquires measurement values from the sensor electrode 22. The detection channel 40 is arranged to amplify the electrical signal from the sensor electrode 22. Thus, the detection channel 40 is designed to amplify very small currents at a sufficient resolution to detect characteristic changes caused by the interaction of interest. Also, the detection channel 40 is designed to have a sufficiently high bandwidth to provide the time resolution required to detect each such interaction. These constraints require sensitive and thus expensive components. Specifically, the detection channel 40 can be arranged as detailed in WO-2010 / 122293 or WO 2011 / 067559 (each is referenced therein and incorporated herein by reference).

[0166] The bias control circuit 41 supplies a bias voltage to the sensor electrode 22 for biasing the sensor electrode 22 with respect to the input of the detection channel 40.

[0167] During normal operation, the bias voltage supplied through the bias control circuit 41 is selected such that the polymer can be shifted through the hole 32. This bias voltage can typically be up to a level of -200 mV.

[0168] The bias voltage supplied by the bias control circuit 41 can also be selected such that it is sufficient to expel the shift from the hole 32. By causing the bias control circuit 41 to supply such a bias voltage, the sensor element 30 can be operated to expel the polymer that is being shifted through the hole 32. To ensure reliable expulsion, the bias voltage is typically a reverse bias voltage, but this is not always necessary. When such a bias voltage is applied, the input to the detection circuit 40 is designed to be maintained at a constant bias potential, even when a negative current is present (having a magnitude similar to that of the normal current, typically in the range of -50 pA to -100 pA).

[0169] For Figure 5 the first arrangement of the electronic circuit 4 illustrated in, a separate detection channel 40 is required for each sensor element 30, which is expensive to implement. Figure 6 Shows a second arrangement for the electronic circuit 4 that reduces the number of detection channels 40.

[0170] In this arrangement, the number of sensor elements 30 in the array is greater than the number of detection channels 40, and the biochemical sensing system is operable to acquire measurement values of the polymer by selected sensor elements in a multiplexed manner, in particular in an electrical measurement multiplexed manner. This is achieved by providing a switching arrangement 42 between the sensor electrodes 23 of the sensor elements 30 and the detection channels 40. Figure 6 A simplified example showing four sensor cells 30 and two detection channels 40 is shown, but the number of sensor cells 30 and detection channels 40 can be greater, typically much greater. For example, for some applications, the sensor device 2 can include a total of 4096 sensor elements 30 and 1024 detection channels 40.

[0171] The switching arrangement 42 can be arranged as described in detail in WO-2010 / 122293. For example, the switching arrangement 42 can include one to N multiplexers each connected to a group of N sensor elements 30 and can include suitable hardware such as latches to select the state of the switches.

[0172] Thus, by switching the switching arrangement 42, the biochemical analysis system 1 can be made operable to acquire measurement values of the polymer by selected sensor elements 30 in an electrical measurement multiplexed manner.

[0173] The switching arrangement 42 can be controlled in the manner described in WO-2010 / 122293 so as to selectively connect the detection channels 40 to respective sensor elements 30 which, based on the amplified electrical signals output by the detection channels 40, have performance of acceptable quality, but additionally, the switching arrangement is controlled as further described below.

[0174] As in the first arrangement, this second arrangement also includes a bias control circuit 41 for each sensor element 30.

[0175] Although in this example the sensor elements 30 are selected in an electrical measurement multiplexed manner, other types of biochemical analysis systems 1 can be configured to switch between sensor elements in a spatial multiplexed manner, for example by movement of a probe for making electrical measurements, or by controlling an optical system for acquiring optical measurement values at different spatial positions by different sensor elements 30.

[0176] The data processor 5 connected to the electronic circuit 4 is arranged as follows. The data processor 5 can be a computer device running an appropriate program, which can be carried out by dedicated hardware devices, or can be carried out by any combination thereof. The computer device used can be any type of computer system, but typically has a conventional configuration. The computer program can be written in any suitable programming language. The computer program can be stored in a computer-readable storage medium, which can be of any type, such as: a recording medium that can be inserted into a drive of a computing system and that can store information magnetically, optically, or magneto-optically; a fixed recording medium of the computer system, such as a hard disk drive; or computer memory. The data processor 5 can include a circuit board inserted into a computer, such as a desktop or laptop computer. The data stored by the data processor 5 can be stored in their memory 10 in a conventional manner.

[0177] The data processor 5 controls the operation of the electronic circuit 3. Just as it controls the operation of the detection channel 41, the data processor controls the bias control circuit 41 and controls the switching of the switch arrangement 31. The data processor 5 also receives and processes a series of measurement values from each detection channel 40. As described further below, the data processor 5 stores and analyzes a series of measurement values.

[0178] The data processor 5 controls the bias control circuit 41 to apply a bias sufficient to cause the polymer to shift through the hole 32 of the sensor element 30. This operation of the biochemical sensor element 41 enables a series of measurement values to be collected by different sensor elements 30, which can be analyzed by the data processor 5 or by another data processing unit to estimate the sequence of polymer units in the polymer, for example using the techniques described in WO-2013 / 041878. The data from different sensor elements 30 can be collected and combined.

[0179] The data processor 5 receives and analyzes a series of measurement values 11 collected by the sensor device 2 and supplied by the electronic circuit 4. The data processor 5 can also provide control signals to the electronic circuit 5, for example to select the voltage applied across the biological pore 1 in the sensor device 2. The series of raw measurement values 11 can be applied in any suitable connection, such as a direct connection when the data processor 5 and the sensor device 2 are physically located together, or any type of network connection when the data processor 5 and the sensor device 2 are physically remote from each other.

[0180] Now will be described Figure 7A method of controlling a biochemical analysis system 1 to analyze polymers. The method is carried out in a manner that increases the analysis speed by excluding polymers that do not require further analysis according to the first aspect of the present invention. The method is implemented in a data processor 5. The method is carried out in parallel for each sensor element 30 for collecting a series of measurement values, i.e., each sensor element 30 in the first arrangement for the electronic circuit 4, and each sensor element 30 connected to the detection channel 40 through a switch arrangement 42 in the second arrangement for the electronic circuit 4.

[0181] In step C1, the biochemical analysis system 1 is operated by controlling the bias control circuit 30 to apply a bias across the pores 32 of the sensor element 30 sufficient to enable the polymer to shift. Based on the output signal from the detection channel 40, the shift is detected and the acquisition of measurement values is started. A series of measurement values is acquired over time.

[0182] In some cases, the following steps are operated on a series of raw measurement values 11 collected by the sensor device 2, i.e., a series of measurement values of the above type, which includes consecutive groups of multiple measurement values depending on the same k-mer without prior knowledge of the number of measurements of any group.

[0183] In other cases, as Figure 8 shown, a state detection step SD is used to preprocess the raw measurement values 11 to derive a series of measurement values 12 to replace the raw measurement values for the following steps.

[0184] In this state detection step SD, a series of raw measurement values 11 is processed to identify consecutive groups of raw measurement values and to derive a series of measurement values 12, which consists of a predetermined number of measurement values for each identified group. Thus, a series of measurement values 12 is derived for each sequence of measured polymer units. The purpose of the state detection step SD is to reduce a series of raw measurement values to a predetermined number of measurement values related to each k-mer to simplify subsequent analysis. For example, a noise staircase signal, as Figure 4 shown, can be reduced to a state where a single measurement value related to each state can be the average current. Such a state can be referred to as a level.

[0185] Figure 9 An example of such a state detection step SD for finding short-term elevations in the derivatives of a series of raw measurement values 11 is shown below.

[0186] In step SD-1, a series of raw measurement values 11 is differentiated to derive its derivative.

[0187] In step SD-2, the derivative of step SD-1 is subjected to low-pass filtering to suppress high-frequency noise, which tends to be amplified by the differentiation in step SD-1.

[0188] In step SD-3, the filtered derivatives from step SD-2 are thresholded to detect transition points between groups of measurement values, thereby identifying groups of the original measurement values.

[0189] In step SD-4, a predetermined number of measurement values are derived from each group of the original measurement values identified in step SD-3. The measurement values output from step SD-4 form a series of measurement values 12.

[0190] The predetermined number of measurement values can be one or more.

[0191] In the simplest approach, a single measurement value is derived from each group of the original measurement values, such as the mean, median, standard deviation, or number of the original measurement values in each determined group.

[0192] In other approaches, a predetermined plurality of measurement values of different natures are derived from each group, such as any two or more of the mean, median, standard deviation, or number of the original measurement values in each determined group. In this case, a predetermined plurality of measurement values of different natures are collected according to the same k-mer, since they are different metrics in the same group of the original measurement values.

[0193] The state detection step SD can use Figure 9 different methods among those shown. For example, Figure 9 a common simplification of the method shown is to use a sliding window analysis, which compares the means of two adjacent windows of data. Then the threshold can be directly set based on the mean difference, or the threshold can be set based on the variance of the data points in the two windows (e.g., by calculating the Student's t-statistic). The unique advantage of these methods is that they can be applied without imposing multiple assumptions about the data.

[0194] Other information related to the measurement level can be stored for later analysis. Such information can include but is not limited to: the change of the signal; the asymmetry information; the confidence of the observation; the length of the group.

[0195] For example, Figure 10 illustrate a series of original measurement values 11 reduced by a moving window t-test determined by experiment. In particular, Figure 10 a series of original measurement values 11 are shown as thin lines. The level after state detection is shown as a dark line overlapping.

[0196] When the polymer is partially shifted through the nanopore, step C2 is performed during the shift. At this time, a series of measurements collected by the polymer during the partial shift are used for analysis, which will be referred to as a "chunk" of the measurements herein. Step C2 can be performed after a predetermined number of measurements have been collected such that the chunk of measurements has a certain size, or alternatively, step C2 can be performed after a predetermined amount of time. In the former case, the size of the chunk of measurements can be defined by a parameter initialized at the start of the run, but is changed dynamically such that the size of the chunk of measurements changes.

[0197] In step C3, the chunk of measurements collected in step C2 is analyzed. This analysis uses reference data 50. As discussed in more detail below, reference data 50 is derived from at least one reference sequence of polymer units. The analysis performed in step C3 provides a measure of the similarity between (a) the sequence of polymer units of the polymer for which the measurements have been collected during the partial shift and (b) a reference sequence. A variety of techniques can be used to perform this analysis, and some examples of which are described below.

[0198] The measure of similarity can represent similarity to the whole of the reference sequence or to a part of the reference sequence, depending on the application. The technique applied in step C3 for deriving the measure of similarity can be selected accordingly, for example, global or local methods.

[0199] In addition, the measure of similarity can represent similarity in terms of a variety of different metrics, provided that it generally provides a measure of how similar the sequences are. Some examples of specific metrics of similarity that can be determined from the sequences in different ways are set out below.

[0200] In step C4, a determination is made in response to the measure of similarity determined in step C3 to (a) reject the polymer being measured, (b) require additional measurements to make a determination, or (c) continue collecting measurements until the polymer is finally measured.

[0201] If the determination made in step C4 is (a) to reject the polymer being measured, then the method proceeds to step C5, where the biochemical analysis system 1 is controlled to reject the polymer such that measurements can be collected by another polymer.

[0202] As follows, step C5 is performed differently between the first and second arrangements of the electronic circuit 4.

[0203] In the case of the first arrangement of the electronic circuit 4, then in step C5, the bias control circuit 30 is controlled to apply a bias voltage across the aperture 32 of the sensor element 30 sufficient to expel the currently shifted polymer. This expels the polymer and thereby makes the aperture 32 available for receiving additional polymer. After such expulsion in step C5, the method returns to step C1, so that the bias control circuit 30 is controlled to apply a bias voltage across the aperture 32 of the sensor element 30 sufficient to shift additional polymer through the aperture 32.

[0204] In the case of the second arrangement of the electronic circuit 4, then in step C5, by controlling the switch arrangement 42, the detection channel 40 currently connected to the sensor element 30 is disconnected and the detection channel 40 is selectively connected to a different sensor element 30, causing the biochemical analysis system 1 to stop acquiring measurement values from the currently selected sensor element 30. At the same time, in step C5, the bias control circuit 30 is controlled to apply a bias voltage across the aperture 32 of the sensor element 30 sufficient to expel the polymer currently shifted through the currently selected sensor element 30, such that the sensor element 30 is available for receiving further polymer in the future.

[0205] The method then returns to step C1, where it is applied to the newly selected sensor element 30 such that the biochemical analysis system 1 begins acquiring measurement values from it.

[0206] If the determination made in step C4 is (b) that additional measurement values are needed to make a determination, then the method returns to step C2. Thus, measurement values of the shifted polymer are continuously acquired until they are next controlled in step C2 and analyzed in step C3. When step C2 is performed again, the block of measurement values collected can be either just the new measurement values for isolated analysis or can be new measurement values combined with the previous block of measurement values.

[0207] If the determination made in step C4 is (c) to continue acquiring measurement values until the end of the polymer, then the method proceeds to step C6 without repeating steps C2 and C3 such that no additional blocks of data are analyzed. In step C6, the sensor element 1 continues to operate such that measurement values are continuously acquired until the end of the polymer. Thereafter, the method returns to step C1 such that additional polymer can be analyzed.

[0208] As represented by the measure of similarity, the degree of similarity, i.e., the basis for the determination in step C4, can vary depending on the nature of the applied and reference sequences. Thus, if the determination is in response to a measure of similarity, then generally there is no limit to the degree of similarity used to make different determinations.

[0209] Some examples of how the dependence on the measure of similarity might vary are as follows.

[0210] Where the reference sequence of the polymer units is an undesired sequence, and in the application where a determination to reject the polymer is made in step C4 in response to a measure of similarity indicating that the partially shifted polymer is an undesired sequence, a relatively high degree of similarity can be used as the basis for rejecting the polymer. Similarly, in the context of the application, the similarity level can vary according to the nature of the reference sequence. When aiming to distinguish similar sequences, a higher similarity level can be required to be used as the basis for rejection.

[0211] Conversely, where the reference sequence of the polymer units from which the reference data 50 is derived is the target and in the application where a determination to reject the polymer is made in step C4 in response to a measure of similarity indicating that the partially shifted polymer is not the target, a relatively low similarity level can be used as the basis for rejecting the polymer.

[0212] As another example, if the application is to determine whether a known gene from a known bacterium is present in a sample of multiple bacteria, the similarity level required to determine whether a polynucleotide has the same sequence as the target will be higher if the gene has a conserved sequence across different strains than if the sequence is not conserved.

[0213] Similarly, in some embodiments of the present invention, the measure of similarity will be equivalent to the degree of identity of the polymer with the target polymer, while in other embodiments, the measure of similarity will be equivalent to the probability that the polymer is the same as the target polymer.

[0214] The similarity level required as the basis for rejection can also vary according to the possible time savings, which in turn depends on the application described below. The acceptable false positive rate can depend on the time savings. For example, when the possible time savings in rejecting undesired polymers is relatively high, it is acceptable to reject an increased proportion of polymers that are the target, provided there is an overall time savings in rejecting undesired polymers.

[0215] Now returning to Figure 7 the method, if at any point during the acquisition of the measurements of the polymer, it is detected that no further measurements are being acquired, indicating that the end of the polymer has been reached, then the method immediately returns to step C1 so that another polymer can be analyzed. After acquiring the measurements of the entire polymer in this way, those measurements can be analyzed as disclosed in WO-2013 / 041878, for example, to derive an estimate of the sequence of the polymer units.

[0216] The source of the reference data 50 can vary according to the application. The reference data 50 can be generated from the reference sequence of the polymer units or from the measurements acquired from the reference sequence of the polymer units.

[0217] In some applications, the previously generated reference data 50 can be pre-stored. In other applications, the reference data 50 is generated when the method is being carried out.

[0218] Reference data 50 can be provided with respect to a single reference sequence of polymer units or multiple reference sequences of polymer units. In the latter case, any step C3 is performed with respect to each sequence or alternatively one of the multiple reference sequences is selected for step C3. In the latter case, the selection can be made according to an application, based on a plurality of criteria. For example, the reference data 50 can be applied to different types of biochemical analysis systems 1 (such as different nanopores) and / or external conditions, and in such a case, the reference model 70 described below is selected based on the type of the actually used biochemical analysis system 1 and / or the actual external conditions.

[0219] Figure 7 The method shown can be changed according to an application. For example, in some variants, the determination in step C4 will never be (c) continue to collect measurement values until the end of the polymer, such that the method repeats the collection and analysis of chunks of measurement values until the end of the polymer.

[0220] In another variant, in step C3, instead of using the reference data 50 and determining a measure of similarity, the determination to reject the polymer in step C4 can be based on other analyses of a series of measurement values, generally any analysis based on chunks of measurement values.

[0221] In one possibility, step C3 can analyze whether a chunk of measurement values is of insufficient quality, such as having a noise level exceeding a threshold, having an error ratio, or a characteristic of the polymer being damaged.

[0222] The determination in step C4 is made based on this analysis, so as to reject the polymer based on an internal quality control check. This still involves making a determination to reject the polymer based on chunks of measurement values, i.e., a series of measurement values collected by the polymer during a partial shift. Therefore, contrary to the case of ejecting the polymer causing a blockage, in the case of rejecting the polymer, the polymer no longer shifts, so no k-mer-dependent measurement values are collected.

[0223] In another possibility, wherein the method is according to the second aspect of the present invention, the method is modified as Figure 11 shown. The method is the same as Figure 7The method is the same, except that step C3 is modified. In step C3, instead of using reference data 50 derived from at least one reference sequence of polymer units and determining a measure of similarity, measurements are processed as observations of a series of k-mer states of different possible types, and include the following general model 60: a transition weight 61, for each transition between consecutive k-mer states in the series of k-mer states, for possible transitions between possible types of k-mer states; and an emission weight 62, for each type of k-mer state, which represents the probability of observing a measurement of a given k-mer. Step C3 is modified to include a measure of fitting the reference model 60.

[0224] The general model 60 can be of the type described in WO-2013 / 041878. Details of the model refer to WO-2013 / 041878. Refer Figure 13 , and the general model 60 is further described below. Derive a measure of fitting, such as the likelihood of the measurement observed by the most similar sequence of k-mer states. This measure of fitting represents the quality of the measurement.

[0225] When step C3 is modified in this way, a decision in step C4 is made based on the measure of fitting, thereby rejecting the polymer based on an internal quality control check.

[0226] Thus, if the similarity to the reference sequence of the polymer unit indicates that the polymer does not require further analysis or if the measurements collected from the polymer have poor quality determined by the model such that additional shifts and measurements are not approved, the method causes the polymer to be rejected. The degree to which the model represents that the data is not good enough depends on the complexity of the model itself. For example, a more complex model can have parameters that can resolve some conditions that can cause rejection.

[0227] Conditions that can cause rejection can include, for example: flowing into an unacceptable signal; high noise; non-model behavior; irregular systematic errors such as temperature fluctuations; and / or errors due to an electro-physical system.

[0228] For example, one possibility is that a polymer or other debris has been accommodated in the nanopore, producing a slowly changing, rather static current. The model generally expects well-separated (piecewise constant in time) steps in the data, so such measurements will have a poor fit to the measure of the model.

[0229] A second possibility is transient noise, such as large variations in current between steps of another closely grouped set. If this noise occurs frequently at high frequencies, the data may have little use for practical purposes. Due to the high frequency of unwanted measurements, the measure of fitting to the model will be low.

[0230] These "errors" can occur in a non-instantaneous manner. Indeed, it is often observed that for adjacent portions, the measurement portions exhibit an offset in their average current. A possible explanation for this is the change in the morphology of the pores and the polymer molecules. Whatever the cause, this behavior is not captured by the model, so the data has little use for practical purposes.

[0231] The impact of such errors can be mitigated to some extent by increasing the complexity of the model. However, this is not desirable and can lead to an increase in the computational cost of modeling the data and decoding the polymer sequence.

[0232] By excluding such polymer chains, only those polymer sequences with strong homology in the derived model transition weights and emission weights give measurements with a good fit to the measure of the model.

[0233] After completing the acquisition of the measurements of the entire polymer, the measurements can be analyzed as disclosed in WO-2013 / 041878 to derive, for example, an estimate of the sequence of polymer units.

[0234] Alternative methods of Figure 7 and Figure 11 can be applied independently or in combination. In such a case, they can be applied simultaneously (e.g., performing step C3 of the two methods in parallel and performing other steps jointly) or sequentially (e.g., performing the method of Figure 7 before the method of Figure 11 ).

[0235] Now a method of controlling the biochemical analysis system 1 shown in Figure 12 to classify polymers will be described. This method is according to the third aspect of the present invention. In this case, the sample chamber 24 contains a sample containing polymers that can be of different types, and the recess 21 serves as a collection chamber for collecting the classified polymers.

[0236] This method is implemented in the data processor 5. This method is performed in parallel for a plurality of sensor elements 30, e.g., each sensor element 30 in a first arrangement for the electronic circuit 4, and each sensor element 30 connected to the detection channel 40 through a switching arrangement 42 in a second arrangement for the electronic circuit 4.

[0237] In step D1, the biochemical analysis system 1 is operated by controlling the bias control circuit 30 to apply a bias across the pores 32 of the sensor element 30 sufficient to enable the polymer to shift. This causes the polymer to start shifting through the nanopore and the following steps are performed during the shift. Based on the output signal from the detection channel 40, the shift is detected and the acquisition of measurements is started. A series of measurements of the polymer are acquired over time by the sensor element 30.

[0238] In some cases, the following steps operate on a series of raw measurements 11 acquired by the sensor device 2, i.e., a series of measurements of the above type, which includes consecutive groups of multiple measurements depending on the same k-mer without any prior knowledge of the number of measurements in any group.

[0239] In other cases, the state detection step SD is used to preprocess the raw measurements 11 to derive a series of measurements 12 for the following steps in place of the raw measurements. This can be done in the same way as the state detection state SD in step C1 described above. Figure 8 and Figure 9 , in the same way as in step C1 described above.

[0240] When the polymer portion moves through the nanopore, i.e., during the movement, step D2 is performed. At this time, a series of measurements acquired by the polymer during the partial movement are collected for analysis, which is referred to herein as a "chunk" of the measurements. Step D2 can be performed after a predetermined number of measurements have been acquired such that the chunk of measurements has a certain size, or alternatively, step D2 can be performed after a predetermined amount of time. In the former case, the size of the chunk of measurements can be defined by a parameter initialized at the start of the run, but is changed dynamically such that the size of the chunk of measurements changes.

[0241] In step D3, the chunk of measurements collected in step D2 is analyzed. This analysis uses the reference data 50. As discussed in more detail below, the reference data 50 is derived from at least one reference sequence of the polymer unit. The analysis performed in step D3 provides a measure of the similarity between (a) the sequence of the polymer unit of the polymer for which the measurements have been acquired during the partial movement and (b) a reference sequence. A variety of techniques can be used to perform this analysis, and some examples of which are described below.

[0242] The measure of similarity can represent similarity to the whole of the reference sequence or to a part of the reference sequence, depending on the application. The technique applied in step D3 for deriving the measure of similarity can be selected accordingly, e.g., global or local methods.

[0243] In addition, the measure of similarity can represent similarity in a variety of different measures, provided that it generally provides a measure of how similar the sequences are. Some examples of specific measures of similarity that can be determined from the sequences in different ways are set forth below.

[0244] In step D4, any determination is made based on the measure of similarity determined in step D3, (a) additional measurements are required to make the determination, (b) the displacement of the polymer into the recess 21 is completed, or (c) the measured polymer is discharged back to the sample chamber 24. If the determination made in step D4 is (a) additional measurements are required to make the determination, then the method returns to step D2. Thus, measurements of the displaced polymer are continuously acquired until the next block of measurements is collected in step D2 and analyzed in step D3. The block of measurements collected when step D2 is performed again can be either just new measurements for isolated analysis or can be new measurements combined with the previous block of measurements.

[0245] If the determination made in step D4 is (b) the displacement of the polymer into the recess 21 is completed, then the method proceeds to step D6 without repeating steps D2 and D3, such that no further analysis of the measurements is performed.

[0246] In step D6, the displacement of the polymer into the recess 21 is completed. As a result, the polymer is collected in the recess 21.

[0247] Step D6 can be performed by applying the same bias across the aperture 32 of the sensor element 30 that enables the displacement of the polymer.

[0248] Alternatively, in step D6, the bias can be changed to effect the remaining displacement of the polymer at an increased rate to reduce the time taken for the displacement. This is advantageous as it increases the overall rate of the sorting process. Increasing the displacement rate is acceptable as no further analysis of the polymer is required. Typically, the change in bias can be an increase. In a typical system, the increase can be significant. For example, in one embodiment, the displacement speed can be increased from about 30 bases / second to about 10,000 bases / second. The possibility of changing the displacement speed can depend on the configuration of the sensor element. For example, when a polymer-binding moiety such as an enzyme is used to control the displacement, it can depend on the polymer-binding moiety used. Advantageously, the polymer-binding moiety can be selected from those that can control the rate.

[0249] During step D6, the sensor element 1 can continue to be operated such that measurements are continuously acquired until the end of the polymer, but this is optional as the remaining sequence does not need to be determined.

[0250] After step D6, the method returns to step D1 such that additional polymers can be displaced.

[0251] If the determination made in step D4 is (c) to discharge the polymer, then the method proceeds to step D5, where the biochemical analysis system 1 is controlled to discharge the measured polymer back to the sample chamber 24 such that measurements of additional polymers can be acquired.

[0252] In step D5, the bias control circuit 30 is controlled to apply a bias across the pores 32 of the sensor element 30 sufficient to expel the currently shifted polymer. This expels the polymer and thereby makes the pores 32 available for receiving additional polymer. After this expulsion in step D5, the method returns to step D1, so that the bias control circuit 30 is controlled to apply a bias across the pores 32 of the sensor element 30 sufficient to shift additional polymer through the pores 32.

[0253] When returning to step D1, the method repeats. The repetitive nature of the method causes the polymer to be continuously shifted and processed from the sample chamber 24.

[0254] Thus, the method utilizes a measure of similarity provided by analyzing a series of measurements taken from the polymer during partial shifting as a basis for whether to collect successive polymers in the recess 21. In this way, the polymers in the sample in the sample chamber 24 are classified and the desired polymers are selectively collected in the recess 21.

[0255] The collected polymer can be recovered. By removing the sample from the sample chamber 24 and then recovering the polymer in the recess 21, this can be done after repeating the method. Alternatively, for example by providing a biochemical analysis system 1 with a fluid system for extracting the polymer from the recess 21, this can be done during the shifting of the polymer of the sample.

[0256] The method can be applied to a wide variety of applications. For example, the method can be applied to polymers that are polynucleotides, such as viral genomes or plasmids. Viral genomes typically have a length on the order of 10 - 15 kB (kilobases) and plasmids typically have a length on the order of 4 kB. In such embodiments, it is not necessary to fragment the polynucleotide and it can be collected whole. The collected viral genome or plasmid can be used in any way, for example to transfect cells. Transfection is the process of introducing DNA into the cell nucleus and is an important tool for research into gene function and regulation of gene expression, thus facilitating progress in basic cell research, new drug discovery, and target validation. RNA and proteins can also be transfected.

[0257] As indicated by the measure of similarity, the similarity, which is used as the basis for the determination in step D4, can vary according to the application and the nature of the reference sequence. Thus, if the determination is dependent on the measure of similarity, then generally there is no limitation on the similarity used to make different determinations.

[0258] Some examples of how the dependence on the measure of similarity might vary are as follows.

[0259] In a variety of applications, the reference sequence of the polymer units of the derived reference data 50 is the desired sequence. In that case, in step D4, a determination of completion of the shift is made in response to a measure of similarity indicating that the partially shifted polymer is the desired sequence, and a relatively high similarity can be used as the basis for completion of the shift.

[0260] However, this is not necessary. In some applications, the reference sequence of the polymer units is an undesired sequence. In that case, in step D4, a determination of completion of the shift is made in response to a measure of similarity indicating that the partially shifted polymer is not the undesired sequence.

[0261] Similarly, in the context of the application, the similarity can be varied according to the nature of the reference sequence. When aiming to distinguish similar sequences, a higher similarity can be required to be used as the basis for rejection.

[0262] This method is performed for each sensor element 30 using the same reference data 50 and the same criteria in step D4. In that case, each groove 21 collects the same polymer in parallel.

[0263] Alternatively, a method can be performed to collect different polymers into different grooves 21. In this case, differential sorting is performed. In one of its instances, different reference data 50 are used for different sensor elements 30. In another of its instances, the same reference data 50 are used for different sensor elements 30, but step D4 is performed with different dependencies on the measure of similarity for different sensor elements.

[0264] It can be varied according to the application Figure 7 、 Figure 11 and Figure 12 the methods shown.

[0265] A variety of different types of reference sequences of polymer units can be used according to the application. Without limitation, when the polymer is a polynucleotide, the reference sequence of the polymer units can include one or more reference genomes for comparing its measured values, or regions of interest of one or more genomes.

[0266] The source of the reference data 50 can be varied according to the application. The reference data can be generated from the reference sequence of the polymer units or from the measured values collected from the reference sequence of the polymer units.

[0267] In some applications, the previously generated reference data 50 can be pre-stored. In other applications, the reference data 50 are generated when the method is performed.

[0268] Reference data 50 may be provided with respect to a single reference sequence of polymer units or multiple reference sequences of polymer units. In the latter case, any step D3 is performed with respect to each sequence, or alternatively one of the multiple reference sequences is selected for step D3. In the latter case, the selection may be made according to the application, based on a variety of criteria. For example, the reference data 50 may be applied to different types of biochemical analysis systems 1 (such as different nanopores) and / or external conditions, in which case the reference model 70 described below is selected based on the type of biochemical analysis system 1 actually used and / or the actual external conditions.

[0269] The above-described biochemical analysis system 1 is an example of a biochemical analysis system that includes an array of sensor elements each containing a nanopore. However, generally, the above method can be applied to any biochemical analysis system that is operable to collect continuous measurements of polymers that may not use nanopores.

[0270] An example of such a biochemical analysis system that does not contain a nanopore is a scanning probe microscope, which can be an atomic force microscope (AFM), a scanning tunneling microscope (STM), or another form of scanning microscope. In this case, the biochemical analysis system can be operable to collect continuous measurements of polymers selected in a spatially multiplexed manner. For example, the polymers can be disposed on a substrate at different spatial positions, and spatial multiplexing can be provided by the movement of the scanning probe microscope.

[0271] In the case where the reader is an AFM, the resolution of the AFM tip can be less fine compared to the size of a single polymer unit. Thus, the measurement can be a function of multiple polymer units. The AFM tip can be functionalized so as to interact with the polymer units in an alternative manner as if it were not functionalized. The AFM can be operated in contact mode, non-contact mode, tapping mode, or any other mode.

[0272] In the case where the reader is an STM, the resolution of the measurement can be less fine compared to the size of a single polymer unit, such that the measurement is a function of multiple polymer units. The STM can be operated conventionally or in any other mode or spectroscopic measurements (STS) can be performed.

[0273] Now the form of the reference data 50 for any of the above methods will be discussed. The reference data 50 can take a variety of forms derived from the reference sequences of polymer units in different ways. The analysis for providing a measure of similarity performed in step C4 or D4 depends on the form of the reference data 50. Some non-limiting examples will now be described.

[0274] In a first example, the reference data 50 represents the consensus of the polymer units of at least one reference sequence. In that case, step C4 or D4 includes the followingFigure 13 The process shown.

[0275] In step C4a-1, the chunk 63 of measurement values is analyzed to provide an estimate 64 of the consistency of the polymer units of the sequence of polymer units of the partially shifted polymer. In general, any method for analyzing measurement values acquired by a biochemical analysis system can be used for step C4a-1.

[0276] The method described in detail in WO-2013 / 041878 can be particularly used for step C4a-1, which is incorporated herein by reference. Refer to the details of the method in WO-2013 / 041878, but a summary is given below.

[0277] The method refers to a general model 60 including transition weights 61 and emission weights 62 regarding a series of k-mer states corresponding to the chunk 63 of measurement values.

[0278] For each transition between consecutive k-mer states in a series of k-mer states, a transition weight 61 is provided. Each transition can be regarded as from a starting k-mer state to an ending k-mer state. The transition weight 61 represents the relative weight of a possible transition between possible types of k-mer states, which is from any type of starting k-mer state to any type of ending k-mer state. In general, this includes the weight for a transition between two k-mer states of the same type.

[0279] For each type of k-mer state, an emission weight 62 is provided. The emission weight 62 is the weight for different measurements to be observed when the k-mer state is of that type. Conceptually, the emission weights 62 can be considered to represent the probability of a given value of the measurement for observing that k-mer state, but they do not need to be probabilities.

[0280] Conceptually, the transition weights 61 can be considered to represent the probability of possible transitions, although they do not need to be probabilities. Thus, the transition weights 61 take into account the probability of k-mer states in which the measurement values depend on their transitions between different k-mer states, which may be more or less likely depending on the types of the starting and ending k-mer states.

[0281] By way of example and not limitation, the model can be an HMM, where the transition weights 61 and the emission weights 62 are probabilities.

[0282] Step C4a-1 uses reference model 60 to derive an estimate 64 of the consistency of the polymer units of the sequence of polymer units of the partially shifted polymer. This can be done using known techniques applicable to the properties of reference model 60. Typically, such techniques derive estimate 64 of the sequence observations of k-mer states based on the likelihood of the measurements predicted by reference model 60. As described in WO-2013 / 041878, such techniques can be performed on a series of raw measurements 11 or a series of measurements 12.

[0283] This method can also provide a measure of the fit of the measurements to the model, such as a quality score representing the likelihood of the measurements predicted by reference model 60 for the most likely sequence observations of k-mer states. Typically these measures are derived because they are used to derive estimate 64.

[0284] As an example, in the case where the general model is an HMM, analytical techniques can be used to solve the known algorithms of the HMM, such as the Viterbi algorithm well-known in the art. In that case, estimate 64 is derived based on the likelihood generated by the total sequence of k-mer states predicted by the general model.

[0285] As another example, in the case where the general model 60 is an HMM, the analytical technique can be of the type disclosed in Fariselli et al., “The posterior-Viterbi: a new decoding algorithm for hidden Markov models”, submitted on January 4, 2005 and archived in the Biology Department of the University of Casadio at Cornell University. In this method, a posterior matrix (representing the probability of the measurements observed for each k-mer state) and a consensus path (where adjacent k-mer states bias towards overlapping paths) are obtained, rather than simply choosing the most likely k-mer for each event. Substantially, this enables the recovery of the same information that would be directly arrived at by the application of the Viterbi algorithm.

[0286] The above description given is in terms of general model 60, which is an HMM where the transition weights 61 and the emission weights 62 are probabilities, and the methods used refer to the probabilistic techniques of general model 60. However, alternatively it is possible that general model 60 uses a framework where the transition weights 61 and / or the emission weights 62 are not probabilities but represent the odds of transitions or measurements in some other way. In such a case, the methods can use analytical techniques rather than probabilistic techniques, which are based on the likelihood predicted by general model 60 for a series of measurements generated by the sequence of polymer units. The analytical techniques can explicitly use a likelihood function, but generally this is not necessary.

[0287] In step C4a-2, the estimated value 64 is compared with the reference data 50 to provide a measure 65 of similarity. This comparison can use any known technique for comparing two sequences of polymer units, typically an alignment algorithm that derives an alignment map between the polymer units, together with a score for the accuracy of the alignment map (and thus the measure 65 of similarity). Any number of available fast alignment algorithms can be used, such as the Smith-Waterman alignment algorithm, BLAST, or their derivatives, or k-mer counting techniques.

[0288] This embodiment of the reference data 50 in this form has the advantage of a rapid process for deriving the measure 65 of similarity, but other forms of reference data are possible.

[0289] In a second embodiment, the reference data 50 represents actual or simulated measurements collected by the biochemical analysis system 1. In that case, step C4 or D4 includes Figure 14 the process shown, which simply includes step C4b of comparing a chunk 63 of the measurements (collected in this case from a series of raw measurements 11) with the reference data 50 to derive a measure 65 of similarity. Any suitable comparison can be made, for example using a distance function to provide a measure of the distance between two series of measurements as the measure 65 of similarity.

[0290] In a third embodiment, the reference data 50 represents a feature vector of temporal characteristics, which represents the characteristics of the measurements collected by the biochemical analysis system 1. Such a feature vector can be derived as detailed in WO-2013 / 121224, which is incorporated herein by reference. In that case, step C4 or D4 includes the following Figure 15 process shown.

[0291] In step C4c-1, a chunk 63 of the measurements, which in this case is collected from a series of raw measurements 11, is analyzed to derive a feature vector 66 of temporal characteristics representing the characteristics of the measurements.

[0292] In step C4c-2, the feature vector 66 is compared with the reference data 50 to derive a measure 65 of similarity. The comparison can be made using the method detailed in WO-2013 / 121224.

[0293] In a fourth embodiment, the reference data 50 represents a reference model 70. In that case, step C4 or D4 includes Figure 16 the process shown, which includes step C4d of fitting a model to a chunk 63 of a series of measurements to provide a measure 65 of similarity as the fit of the reference model 70 to the chunk 63 of the measurements. The chunk 63 of the measurements can be a series of raw measurements 11 or a series of measurements 12.

[0294] The C4d step can be carried out as follows.

[0295] The reference model 70 is a model of the reference sequence of polymer units in the biochemical analysis system 1. The reference model 70 processes the measured values as observations of the k-mer states of a reference series corresponding to the reference sequence of the polymer units. The k-mer state of the reference model 70 can model the measured values depending on the actual k-mers thereof, but this is not necessary mathematically, so the k-mer state can be an abstract concept of the actual k-mers. Therefore, different types of k-mer states can correspond to different types of k-mers present in the reference sequence of the polymer units.

[0296] The reference model 70 can be considered as an adaptation of the general model 60 of the type described above and in WO-2013 / 041878 to model the measured values specifically obtained when measuring the reference sequence. Therefore, the reference model 70 processes the measured values as observations of the k-mer state 73 of a reference series corresponding to the reference sequence of the polymer units. Thus, the reference model 70 has the same form as the general model 60, in particular including the transition weights 71 and emission weights 72 which will now be described.

[0297] The transition weights 71 represent the transitions between the k-mer states 73 of the reference series. Those k-mer states 73 correspond to the reference sequence of the polymer units. Therefore, consecutive k-mer states 73 in the reference series correspond to consecutive overlapping groups of k-polymer units. Thus, there is an inherent mapping between the k-mer states 73 present in the reference series and the polymer units of the reference sequence. Similarly, each k-mer state 73 has a type corresponding to a different type of combination of each polymer unit in the group of k-polymer units.

[0298] Reference Figure 17 to its state diagram illustrates it, Figure 17 showing an example of three consecutive k-mer states 73 in the reference series of the estimated k-mer states 73. In this example, k is 3, and the reference sequence of the polymer units contains consecutive polymer units labeled A, A, C, G, T (but of course those specific types of k-mer states 73 are not restricted). Therefore, the consecutive k-mer states 73 of the reference series corresponding to those polymer units are of the types AAC, ACG, CGT, which correspond to the measurement sequence AACGT of the polymer units.

[0299] Figure 18The state diagram of shows the transitions between the reference series of k-mer states 73 represented by the transition weights 71. In this embodiment, the state may only allow the reference series of k-mer states 73 to travel through forward (but in general may also allow backward travel). Three different types of transitions 74, 75 and 76 are shown below.

[0300] From each given k-mer state 73 in the reference series, a transition 74 to the next k-mer state 73 is allowed. This models the likelihood of a series of consecutive measurements 12 collected from consecutive k-mers of polymer units of the reference sequence. In the case of preprocessing blocks of measurements 63 to identify consecutive groups of measurements, and deriving a series of process measurements consisting of a predetermined number of measurements relative to each identified group for further analysis, the transition weight 71 represents such a transition 74 with a relatively high likelihood.

[0301] By each given k-mer state 73 in the reference series, a transition 75 to the same k-mer state is allowed. This models the likelihood of a series of measurements 12 collected from the same k-mer of the polymer unit of the reference sequence. It can be called a "stay". In the case of pre-processing a block 63 of measurements to identify a continuous group of measurements and deriving a series of process measurements consisting of a predetermined number of measurements (about each identified group) for further analysis, the transition weight 71 represents such a transition 75 with a relatively high likelihood compared to the transition 74.

[0302] For each given k-mer state 73 in the reference series, a transition 76 to a subsequent k-mer state 73 is allowed to skip the next k-mer state 73. This models the likelihood of no measurements collected from the next k-mer state, so that consecutive measurements in a series of measurements 12 collected by the k-mers of the reference sequence of polymer units are separated. It can be called a "skip". In the case of preprocessing a block 63 of measurement values ​​to identify consecutive groups of measurement values ​​and deriving a series of process measurements consisting of a predetermined number of measurement values ​​(about each identified group) for further analysis, the transition weight 71 represents such a transition 76 with a relatively high likelihood compared to the transition 74.

[0303] The levels of transitions 75 and 76 representing transitions for ... 75 and 76 may be obtained in the same manner as the transition weights 61 for transitions for transitions for transitions for transitions for transitions for transitions for transitions for transitions 71 in the general model 31 described above.

[0304] In an alternative embodiment, instead of preprocessing the chunks 63 of measurement values to identify contiguous groups of measurement values and derive a series of processed measurement values such that additional analysis is performed on the chunks 63 of measurement values themselves, the transition weights 71 are similar, but are rewritten to increase the likelihood of a transition 75 representing a jump to represent the likelihood of contiguous measurement values collected by the same k-mer. The level of the transition weights 71 for the transition 75 depends on the number of expected measurement values collected by any given k-mer and can be determined experimentally for the particular biochemical analysis system 1 being used.

[0305] For each k-mer state, emission weights 72 are provided. The emission weights 72 are the weights for the different measurement values observed when a k-mer state is observed. The emission weights 72 thus depend on the type of k-mer state under discussion. In particular, the emission weights 72 for any given type of k-mer state are the same as those for the types of k-mer states in the general model 60 described above.

[0306] Except for replacing the general model 60 with the reference model 70, step C4d is performed using the same techniques as described above Figure 13 to fit the model to the chunks 63 of a series of measurement values to provide a measure 65 of the similarity of the fit of the reference model 70 to the chunks 63 of measurement values.

[0307] Due to the form of the reference model 70, in particular the representation of the transitions between the reference series of k-mer states 73, applying the model inherently derives an estimate of the alignment mapping between the chunks 63 of measurement values and the reference series of k-mer states 73. The understanding of this can be as follows. Since the general model 60 represents the transitions between the possible types of k-mer states, applying the model provides an estimate of the type of k-mer state for which each measurement value is observed in particular. Since the reference model 70 represents the transitions between the reference series of k-mer states 73, applying the reference model 70 instead estimates the k-mer state 73 of the reference sequence for which each measurement value is observed by it, which is the alignment mapping between the series of measurement values and the reference series of k-mer states 73.

[0308] In addition, the algorithm derives a score for the accuracy of the alignment mapping, such as representing the likelihood that the estimate of the alignment mapping is correct, such as because the algorithm derives the alignment mapping based on such scores for the different paths in the model. Thus, such a score for the accuracy of the alignment mapping is thus a measure 65 of similarity.

[0309] As an example, in the case where the reference model 70 is an HMM and the analysis technique applied is the Viterbi algorithm described above, then the score is simply the likelihood related to the derived estimate of the alignment mapping predicted by the reference model 70.

[0310] As another example, in the case where the general model 60 is an HMM, the analysis technique can be of the type disclosed by Fariselli et al. as described above. Its re-derivation is a score that is a measure of similarity 65.

[0311] A reference model 70 can be generated as follows from a reference sequence of polymer units or from measurements taken from a reference sequence of polymer units.

[0312] It can be done as follows by Figure 19 The process shown generates a reference model 70 from a reference sequence of polymer units 80. This can be used in applications where the reference sequence is known from a database or earlier experiments. The input data representing the reference sequence of polymer units 80 can already be stored in the data processor 5 or can be input thereto.

[0313] The process uses stored emission weights 81, which include emission weights e1 to en for a set of possible types of k-mer states type-1 to type-n. Advantageously, this allows a reference model for any reference sequence of polymer units 80 to be generated based only on the emission weights 81 for the possible types of k-mer states.

[0314] The process proceeds as follows.

[0315] In step P1, a reference sequence of polymer units 80 is received and a reference sequence of k-mer states 73 is generated therefrom. This is a simple process of establishing the types of those k-mer states 73 for each k-mer state in the reference sequence, based on the combination 73 of the types of polymer units 80 to which the k-mer states 73 correspond.

[0316] In step P2, a reference model is generated as follows.

[0317] Transition weights 71 are derived for transitions between the reference series of k-mer states 73 derived in step P1. The transition weights 71 take the form defined above for the reference series of k-mer states 73.

[0318] In step P1, emission weights 72 are derived for each k-mer state 73 in a series of k-mer states 73 by selecting the stored emission weights 81 according to the type of k-mer state 73. For example, if a given k-mer state 73 is type type-4, then emission weight e4 is selected.

[0319] It is done as follows by Figure 20The process shown generates a reference model 70 from a series of reference measurements 93 taken from a reference sequence of polymer units. This can be used, for example, in applications where a reference sequence of polymer units is measured simultaneously with a target polymer. In particular, in this embodiment, it is not required that the polymer units of the reference sequence are themselves known. A series of reference measurements 93 can be taken from a polymer containing polymer units with a reference sequence by a biochemical analysis system 1.

[0320] The process uses an additional model 90 which processes a series of reference measurements as observations of a further series of possible similar k-mer states. Such an additional model 90 is a model of the biochemical analysis system 1 for taking a series of reference measurements 93 and can be the same as the general model 60 of the type disclosed, for example, in WO-2013 / 041878. Thus, the additional model includes a transition weight 91 for each transition between consecutive k-mer states in a further series of k-mer states, which is a transition weight 91 for possible transitions between possible types of k-mer states; and an emission weight 92 for each type of k-mer state, which is an emission weight 92 for different measurements to be observed when the k-mer state is of that type.

[0321] The process is carried out as follows.

[0322] In step Q1, the additional model 90 is applied to a series of reference measurements 93 to estimate the k-mer states 73 of the reference series as discretely estimated k-mer states. This can be done using the techniques described above.

[0323] In step Q2, the reference model 70 is generated as follows.

[0324] Transition weights 71 are derived for transitions between the reference series of k-mer states 73 derived in step D1. The transition weights 71 take the form defined above for the reference series of k-mer states 73.

[0325] In step Q1, emission weights 72 are derived for each k-mer state 73 in a series of k-mer states 73 by selecting the emission weights from the weighting of the additional model 50 according to the type of the k-mer state 73. Thus, the emission weight for each type of k-mer state 73 in the reference model is the same as the emission weight for the k-mer state 73 of that type in the further model 50.

[0326] Now will be described Figure 7Embodiments of the methods shown, and more generally of various applications in accordance with the first aspect of the present invention, explain the nature of the reference sequences of polymer units, the basis for the determination in step C4, and the representation of possible time savings. In the following embodiments, the polymer is a polynucleotide and it is assumed that measuring after the first 250 nucleotides and comparing with the reference sequence will be sufficient to determine (a) whether it relates to the reference sequence and (b) its position with respect to the total sequence. However, it can be more or less than that number. The number of polymer units required for determination will not have to be fixed. Typically, measurements will be made continuously on a continuous basis until such a determination is made.

[0327] For each of the application types, there may be Figure 7 slightly different uses of the methods shown. Mixtures of application types can also be used. It is also possible to dynamically adjust the analysis carried out in step C3 and / or the basis for the determination in step C4 as the run progresses. For example, there may be no initial application determination logic, and then the logic can be applied to the run after sufficient data has been established to make a determination. Alternatively, the determination logic can change during the run.

[0328] In a first class of applications, the reference sequence of the polymer units from which the reference data 50 is derived is an undesired sequence, and in step C4, a determination to reject the polymer is made in response to a measure of similarity indicating that the polymer with a partial shift is an undesired sequence.

[0329] This first class of applications has a variety of possible uses. For example, such an application can be used for an incompletely sequenced part of the genome of an organism. If a part defines the genome of an organism but the sequence is incomplete, the method of the present invention can be used to determine the incomplete part of the sequence. In such an embodiment, the reference sequence can be the sequence of the complete part of the genome. The polymer can be a fragment of a polynucleotide from the organism. If the measure of similarity indicates that the polymer is the reference sequence (i.e., the sequence of the already defined part of the genome), then the polymer is rejected and a new polymer can be received through the nanopore. This can be repeated until a polymer that is not similar to the reference sequence is partially shifted through the nanopore, and such a polymer will correspond to a previously undefined part of the genome and can be retained in the nanopore and fully sequenced. The method allows for the rapid sequencing of undefined parts of the genome.

[0330] Applications of the first type can also be advantageously used to sequence polymers from a polymer sample containing human DNA. Sequencing of human DNA has ethical issues associated with it. Thus, it is useful to be able to sequence a sample of a polymer and disregard the sequence of human DNA (such as bacterial identification in a sample extracted from a human patient). In this case, the reference sequence (the undesired sequence) can be the human genome. Any polymer having a measure of similarity indicating a portion corresponding to the human genome can be excluded, while a polymer having a measure of similarity indicating that it does not correspond to the human genome can be retained in the nanopore and sequenced in its entirety. Thus, this is an example of a method where the measure of similarity indicates similarity to a portion of the reference sequence. In the present application, the method avoids sequencing human DNA but allows sequencing of bacterial DNA. If the bacteria are in a sample from the human gut, we assume that the bacterial DNA (which is the DNA we want to sequence or "target" DNA) is about 5% of the DNA and 95% of the DNA in the sample is human DNA ("off-target DNA"). If we assume that a sequence of about 250 bp (base pairs) per fragment will be sufficient to provide the required measure of similarity, and that the polymer can translocate through the pore at a rate of 25 bases / second, then a polymer that is not the target DNA (i.e., DNA similar to the human DNA reference sequence ("off-target" polymer)) will translocate through the nanopore for about 10 seconds before being ejected. Thus, the relative amount of time during which the nanopore contains the off-target polymer can be considered to be 95% x 10 = 9.5. On the other hand, assuming the DNA is fragmented into 10 kB fragments, the amount of time taken to sequence one fragment of the target DNA will be 10,000 / 25, which is 400 seconds. Thus, the relative amount of time during which the nanopore contains the target polymer can be considered to be 5% x 400, which is 20 seconds. So the proportion of the time during which the nanopore contains the target strand can be considered to be the time during which the nanopore contains the target strand / (the time during which the nanopore contains the off-target strand + the time during which the nanopore contains the target strand), which is 20 / 29.5. On the other hand, if the off-target strands need to be sequenced in their entirety, the relative amount of time during which the nanopore contains the off-target strands will be 95% x 400, which is 380, and so the proportion of the time during which the nanopore contains the target strand can be considered to be 20 / 380. This represents an efficiency of about 13.6 times.

[0331] The first class of applications can also be advantageously used for sequencing contaminants in a sample. In such an embodiment, the reference sequence will be the sequence of a known component present in the sample. For example, it can be used to detect contaminants in foods such as meat products like beef. In this case, the reference sequence will be the sequence of a polynucleotide from an organism from which the food is derived (e.g., the genome of the organism). The reference sequence can be the sequence of the genome of a cow. Any polymers in the sample having a measure of similarity indicating that they correspond to the cow genome can be excluded, while polymers having a measure of similarity indicating that they do not correspond to the cow genome can be retained in the nanopore and fully sequenced. This will allow for the rapid and simple definition of the nature of the contaminant without the need to know the nature of the contaminant. This is advantageous compared to prior art methods such as quantitative PCR that require knowledge of the suspected contaminant. Assuming 99% of the DNA is off-target (meat DNA) and 1% of the DNA is the target (e.g., the contaminant), then the method of the present invention will be approximately 29 times more efficient than if the nanopore could not exclude the unwanted polymers.

[0332] In a second class of applications, the reference sequence of the polymer units from which the reference data 50 is derived is the target, and in step C4, a determination to exclude the polymer is made in response to a measure of similarity indicating that the partially shifted polymer is not the target.

[0333] This second type of application can be advantageously used for sequencing a gene of interest from a DNA sample. In this application, the reference sequence is the target, which can be part of a polynucleotide such as the gene of interest, and the polymers can comprise fragments of polynucleotides from the sample such as DNA. Any polymers in the sample having a measure of similarity indicating that they are not similar to the target (gene of interest) can be excluded. The remaining polymers can be retained and sequenced. This allows rapid sequencing of the gene of interest and is advantageous over the prior art, which requires isolation of the target gene of interest (e.g., by hybridizing the gene of interest to a probe attached to a solid surface) prior to sequencing. Such isolation techniques are time-consuming and not required when using the method of the present invention. An example of such an application would be sequencing the human genome. The human genome contains 50 Mb (megabases) of coding sequence. It would be desirable to be able to sequence that 50 Mb rather than the remaining 3,000 Mb. Thus, the amount of DNA that is "off-target" (to be excluded) is 3,000 Mb. The DNA will be fragmented into fragments of approximately 10 kB in length, and thus 3,000 Mb will represent approximately 300,000 fragments. Assuming that a sequence of approximately 250 bp per fragment will be sufficient to provide the required measure of similarity and that the polymers can translocate through the pore at a rate of 25 bases per second, then polymers that are not similar to the target polymer ("off-target" human DNA) will translocate through the nanopore for approximately 10 seconds before being ejected. Since there are 300,000 off-target fragments, the off-target fragments will remain in the pore for approximately 3,000,000 seconds / nanopore (number of fragments multiplied by the time each fragment remains in the pore - approximately 10 seconds). The remaining 50 Mb ("target") that is similar to the target polymer will take 2,000 seconds (the time it will take at 25 bases per second is equal to 50,000,000 / 25 or 2,000,000 seconds). The total time to sequence the described 50 Mb target polymer is the sum of the amount of time it takes to sequence the off-target polymers and the amount of time it takes to sequence the target polymers, which is 3,000,000 + 2,000,000 or 5,000,000 seconds / nanopore. On the other hand, if the entirety of each of the 300,000 off-target fragments is sequenced, then this will take 3,000,000,000 / 25 (sequencing 3,000 Mb at a rate of 25 base pairs per second) + 2,000,000 (the time it takes to sequence the target polymer), which is 122,000,000 seconds / pore (more than 50 times longer) to sequence the genome once.

[0334] This second type of application can also be advantageously used to identify whether bacteria in a sample (e.g., from an in-patient) are antibiotic-resistant. Here, the reference sequence will be the target, which can be a polynucleotide corresponding to a specific antibiotic-resistant gene. Any polymer in the sample with a measure of similarity indicating similarity to the target antibiotic-resistant gene can be excluded. If no polymer is detected with a measure of similarity indicating their similarity to the antibiotic-resistant gene, this will indicate that the bacteria are losing a specific antibiotic-resistant gene. Alternatively, if polymers are detected that do have a measure of similarity indicating their similarity to the antibiotic-resistant gene, they can be retained and sequenced, and the sequence used to determine whether the antibiotic-resistant gene is functional. In this case, the off-target polymer (the bacterial genome) will be approximately 5000 kB, and the target polymer (the region of the sensing zone) will be approximately 5 kB. Making the same assumptions as above, it means that the method of the present invention will sequence DNA approximately 40 times faster than if the nanopore could not eject unwanted polymers.

[0335] This second type of application can also be advantageously used to sequence total bacterial mRNA. In this case, it is desirable to be able to sequence the mRNA, but be able to ignore the sequences of rRNA or tRNA. Here, the reference sequence can be a target sequence such as an annotated version of the bacterial genome. The polymers can comprise RNA from a sample of bacteria. Any polymer in the sample with a measure of similarity indicating their dissimilarity to the target bacterial genome will be related to rRNA or tRNA and can be excluded. The remaining polymers will correspond to mRNA and can be sequenced to provide the sequence of total bacterial mRNA. In this case, the target polymer will be mRNA (which is approximately 5% of the total RNA), and the off-target polymers will be tRNA and rRNA, which are approximately 95% of the total RNA. Using the same assumptions as defined above, we expect an approximately 8.4-fold increase in sequencing efficiency.

[0336] This second type of application can also be advantageously used to identify strains for phenotypic or SNP (single nucleotide polymorphism) detection, where the strain of the bacteria is not known. For example, in this case, the polymers can be fragments of polynucleotides from a sample of bacteria. Initially, the polymers are not excluded (no reference sequence is used) and the polymers that have translocated through the pore are sequenced, but when enough sequence information has been obtained to allow the user to determine the strain of the bacteria, then a reference sequence is selected. The reference sequence will correspond to the target region of interest and will depend on the type of bacteria that has been defined. Once the reference sequence has been defined, any polymer (the target portion of interest) that has partially translocated through the pore and has a measure of similarity indicating their similarity to the reference sequence is retained and fully sequenced, while other polymers can be excluded. This will allow the detection of the presence of a phenotype or SNP.

[0337] Similarly, this second type of application will be useful for the phenotypes of cancer. In this application, the polymers can be fragments of polynucleotides obtained from cancer patients. Initially, the reference sequence can be the target sequence. These target sequences can be polynucleotides such as the sequences of genes associated with different classes of cancer. Any polymers that retain a measure of similarity to these target sequences will be retained, and other polymers will be excluded. However, once the class of cancer has been identified, the reference sequence can be refined such that the reference sequence now includes targets with sequences of polynucleotides associated with a subclass of the cancer.

[0338] In a third type of application, the reference sequence for the polymer units from which the reference data 50 is derived is the sequence of the polymer units that have been measured, and in step C4, a determination to exclude a polymer is made in response to a measure of similarity indicating that the partially shifted polymer is the sequence of the polymer units that have been measured.

[0339] This type of application can be used to enable accurate genome sequencing. Determining the sequence of a genome requires sequencing multiple strands of DNA, and for accuracy, the consensus sequence of that portion of DNA will be determined. Therefore, the polymers corresponding to the same portion of the sequence should be sequenced enough times to be able to define an accurate consensus sequence. For this purpose, the method of the present invention can be used to rapidly and accurately sequence a genome. For example, the polymers can contain DNA from a sample of DNA from an organism that will define the genome. The reference sequence is a portion of the DNA for which sufficient measurements have been taken (in this case, sufficient sequence data has been obtained to provide an accurate consensus sequence). Initially, no sequences are excluded. However, once it has been calculated that sufficient sequence has been obtained for a portion of the genome to allow calculation of an accurate consensus sequence, then that consensus sequence becomes the target (reference sequence). Any polymers that partially shift through the pore and have a measure of similarity indicating that they are similar to the reference sequence (the portion of DNA for which the accurate consensus sequence has been defined) can be excluded, releasing the nanopore to sequence other portions of the genome for which sufficient information has not yet been collected.

[0340] In a fourth type of application, the reference sequence for the polymer units from which the reference data 50 is derived contains multiple targets, and in step C4, a determination to exclude a polymer is made in response to a measure of similarity indicating that the partially shifted polymer is one of the targets.

[0341] This is a counting method that can be used to quantify the proportion of each target polymer in a sample of target polymers. For example, the targets can represent different polymers. When a polymer portion translocates through a nanopore, any polymer having a measure of similarity indicating its similarity to a reference sequence can be assigned to a "bin" and the number of polymers belonging to each "bin" can be quantified. In this embodiment, once sufficient information about the polymer is obtained to determine whether it has a measure of similarity indicating its similarity to one of the reference sequences, the polymer is excluded. An example of the use of this technique is to quantify contaminants. For example, the polymer can be a sample of food such as a beef product. In this case, the reference sequence can include targets having sequences found in cow DNA and targets having sequences found in horse DNA. The method can be used to calculate the proportion of polymers similar to the cow DNA target and the proportion of polymers similar to the horse DNA, and this will indicate the level of contamination of the beef product with horse meat.

[0342] Similarly, if the reference sequence used includes targets having sequences found in different bacteria, then this technique can be used to determine the proportion of different bacteria present in a sample such as a sample from an infected patient.

[0343] Figure 16 The method shown results in the generation of an alignment map. The method can be applied more generally as follows.

[0344] Figure 21 A method of estimating an alignment map between (a) a series of measurements of a polymer containing polymer units and (b) a reference sequence of the polymer units is shown. The method is carried out as follows.

[0345] As Figure 21 shown, input to the method can be a series of raw measurements of the sequences of the polymer units acquired by the biochemical analysis system 1 and a series of measurements 12 derived by subjecting them to preprocessing as described above. Alternatively, input to the method can be a series of raw measurements 11.

[0346] The method uses a reference model 70 of the reference sequence of the polymer units, and the reference model 70 is stored in the memory 10 of the data processor 5. The reference model 70 takes the same form as described above and processes the measurements as observations of the reference sequence corresponding to the k-mer states of the reference sequence of the polymer units.

[0347] The reference model 70 is used in the alignment step S1. In particular, in the alignment step S1, the reference model 70 is applied to a series of measurements 12. The alignment step S1 is carried out in the same manner as step C4d above. In other words, except that the reference model 70 replaces the general model 60, by using the same reference as above Figure 13Perform step C4d with the same technology as described, fitting the model to chunks of a series of measurements 63 to provide a measure of the similarity 65 of the fit of the reference model 70 to the chunks of measurements 63 for alignment step S1.

[0348] Due to the form of the reference model 70, particularly the representation of the transitions between reference series of k-mer states 73, applying the model inherently derives an estimate of the alignment mapping between a series of measurements and the reference series of k-mer states 73. The understanding of this can be as follows. Since the general model 60 represents the transitions between possible types of k-mer states, applying the model provides an estimate of the type of k-mer state from which each measurement is observed, i.e., an estimate of the initial series of k-mer states 34 and the discretely estimated k-mer states 35, for each estimate of the type of k-mer state from which each measurement is observed. Since the reference model 70 represents the transitions between reference series of k-mer states 73, applying the reference model 70 instead estimates the k-mer states 73 of the reference sequence particularly from which each measurement is observed, which is the alignment mapping between a series of measurements and the k-mer states 73 of the reference series.

[0349] Due to the inherent mapping between the k-mer states 73 of the reference series and the polymer units of the reference sequence, the alignment mapping between a series of measurements of k-mer states 73 and the reference series also provides an alignment mapping between a series of measurements of polymer units and the reference sequence.

[0350] Figure 22 An embodiment of the alignment mapping is shown to illustrate its nature. In particular, Figure 22 An alignment mapping between the polymer units p0 to p7 of the reference sequence, the k-mer states k1 to k6 of the reference series, and the measurements m1 to m7 is shown. By way of example, in this embodiment, k is three. The horizontal lines represent the alignment between the k-mer states and the measurements, or the alignment of gaps in other series in the case of dashes. Thus, inherently, the polymer units p0 to p7 of the reference sequence as illustrated pair to its k-mer states k1 to k6 of the reference series. The k-mer state k1 corresponds to and maps to the polymer units p1 to p3 and so on. As for the mapping between the k-mer states k1 to k6 of the reference series and the measurements m1 to m7: the k-mer state k1 maps to the measurement m1, the k-mer state k2 maps to the measurement m2, the k-mer state k3 maps to a gap in the series of measurements, the k-mer state k4 maps to the measurement m3, and the measurements m4 and m5 map to a gap in the series of k-mer states.

[0351] Depending on the method applied, the form of the estimate 13 of the alignment mapping can be changed as follows.

[0352] As described above, the analysis techniques applied in the alignment step S1 can take various forms suitable for the form of the reference model 70. For example, in the case where the reference model 70 is an HMM, the analysis technique can be a known algorithm for solving the HMM, such as the well-known Forward-Backward algorithm or Viterbi algorithm in the art. Generally, such algorithms can avoid brute-force calculation of the likelihood (probability) of all possible paths through the sequence of states, but instead use a simplified likelihood-based method to determine the state sequence.

[0353] By some of the techniques applied in the alignment step S1, the derived estimate 13 of the alignment mapping includes a weighting of different k-mer states 73 in the reference series with respect to each measurement value 12 in the series. For example, this can be done by M i,j represent such an alignment mapping, where the index i labels the measurement value and the index j labels the k-mer states in the reference series, so that in the presence of K k-mer states, M i,1 to M i,K values represent the weighting for the i-th measurement value with respect to each k-mer state 73 in the reference series of k-mer states 73. In this case, the estimate 13 does not represent a single k-mer state 73 because it is mapped to each measurement value, but instead provides a weighting of the different possible k-mer states 73 so mapped to each measurement value.

[0354] As an example in the case where the reference model 70 is an HMM, when the analysis technique applied is the above-mentioned Forward-Backward algorithm, the derived estimate can be of this type. In the Forward-Backward algorithm, the transition and emission weightings are used to cycle through the forward and backward directions to calculate the total likelihood of all sequences ending in a given k-mer state. Combining these forward and backward probabilities and calculating together with the total likelihood of the data, the probability of each measurement from a given k-mer state. This probability matrix, called the posterior matrix, is the estimate 13 of the alignment mapping.

[0355] In this case, in the subsequent scoring step S2 (which is optional), there is a score 14 representing the likelihood that the estimate 13 of the alignment mapping is correct. This can be derived from the estimate 13 of the alignment mapping using simple probability techniques, or alternatively, it can be derived as an inherent part of the alignment step S1.

[0356] By other techniques applied in the alignment step S1, the derived estimate 13 of the alignment mapping includes a discrete estimate of the k-mer states in the reference series of k-mer states for each measurement value in the series. For example, such an alignment mapping can be represented by Mi denotes, where the index i denotes the measured value and M i Values 1 to K representing the K k-mer states can be adopted. In this case, the estimated value 13 represents a single k-mer state 73 mapped to each measured value.

[0357] As an example where the reference model 70 is an HMM, when the applied analysis technique is the Viterbi algorithm described above, the derived estimated value can be of this type, where the analysis technique estimates the sequence of k-mers based on the likelihood expected by a model of a series of measured values generated by a reference series of k-mer states.

[0358] In the case where the derived estimated value 13 of the alignment mapping includes discrete estimated values of k-mer states, the algorithm inherently derives a score 14 representing the likelihood that the estimated value of the alignment mapping is correct, because the algorithm derives the alignment mapping based on scores for different paths through the model. Thus, in this case, no separate scoring step S2 is performed. As an example, in the case where the reference model 70 is an HMM and the applied analysis technique is the Viterbi algorithm described above, then the score is simply the likelihood related to the derived estimated value 13 of the alignment mapping predicted by the reference model 70.

[0359] Figure 21 The method shown has wide applications where it is desired to estimate the alignment mapping between a series of measured values of a polymer and a reference sequence of polymer units and / or a score representing the likelihood that the alignment mapping is accurate. Evaluation of such alignment mappings can be used in various applications, such as comparing references to provide identification or detection of the presence, absence, or degree of a polymer in a sample, for example to provide a diagnosis. The specific applications within the possible range are numerous and can be applied to detecting any analyte with a DNA sequence.

[0360] The above embodiments relate to a single reference model 70. In many applications, multiple reference models 70 can be used. As Figure 21 shown, the method can be applied to use each reference model 70, or one of the reference models 70 can be selected. Depending on the application, the selection can be based on various criteria. For example, the reference model 70 can be applied to different types of sensor devices 2 (such as different nanopores) and / or external conditions, in which case the following-described reference model 8 is selected based on the type of sensor device 2 actually used and / or the actual external conditions. In another embodiment, the selection can be made based on the analyte to be detected, such as specific G / C enrichment or whether specific epigenetic information is experimentally determined.

[0361] Accordingly, in a fourth aspect of the present invention, there is provided a method of estimating an alignment mapping between: (a) a series of measurements of a polymer comprising polymer units, wherein the measurements depend on k-mers, which are k polymer units of the polymer, where k is an integer, and (b) a reference sequence of polymer units;

[0362] The method uses a reference model that processes reference data that is a series of reference k-mer states corresponding to the reference sequence of polymer units, wherein the reference model includes:

[0363] Transition weights for transitions between k-mer states in the reference series of k-mer states; and

[0364] For each k-mer state, emission weights for different measurements when observing the k-mer state; and

[0365] The method includes applying the reference model to the series of measurements to derive an estimate of the alignment mapping between the series of measurements and a reference series of k-mer states corresponding to the reference sequence of polymer units.

[0366] The following features may optionally be applied to the fourth aspect of the present invention in any combination:

[0367] For each measurement in the series, the estimated alignment mapping may include a discrete estimate of the mapped k-mer state in the reference series of k-mer states.

[0368] For each measurement in the series, the estimated alignment mapping may include weights for different mapped k-mer states in the reference series of k-mer states.

[0369] The method may further include deriving a score representing the likelihood that the estimated alignment mapping is correct.

[0370] The method may further include generating a reference model from a reference sequence of polymer units by a process including using a stored set of possible types of emission weights for k-mer states:

[0371] Deriving a series of k-mer states corresponding to the reference sequence of the received polymer;

[0372] Generating the reference model by generating transition weights for transitions between k-mer states in the derived series of k-mer states and by selecting emission weights for each k-mer state in the derived series according to the type of k-mer state from the stored emission weights.

[0373] The method may further include generating a reference model from a series of reference measurements of a polymer comprising a reference sequence of polymer units.

[0374] The step of generating a reference model can use an additional model that processes a series of reference measurements as observations of a further series of k-mer states of different possible types, where the additional model includes:

[0375] For each transition between consecutive k-mer states in the further series of k-mer states, a transition weight for possible transitions between possible types of k-mer states; and

[0376] For each type of k-mer state, an emission weight for different measurements when the k-mer state is of that type.

[0377] The step of generating a reference model includes:

[0378] Generating an estimate of a reference series of k-mer states by applying the additional model to a series of reference measurements; and

[0379] Generating a reference model by generating a transition weight for transitions between k-mer states in the estimate of the reference series of k-mer states and by selecting an emission weight for each k-mer state in the estimate of the reference series to be generated according to the type of k-mer state by weighting of a further model.

[0380] The reference model can be pre-stored.

[0381] One or both of the transition weight and the emission weight can be a probability.

[0382] The model can be a hidden Markov model.

[0383] The integer k can be a complex number.

[0384] The measurements can be measurements acquired during the translocation of the polymer through the nanopore.

[0385] The translocation of the polymer through the nanopore can occur in a ratchet-like manner.

[0386] The nanopore can be a biological pore.

[0387] The polymer can be a polynucleotide, and the polymer unit can be a nucleotide.

[0388] A single measurement can depend on a k-mer, or a predetermined plurality of measurements of different properties can depend on the same k-mer.

[0389] The measurements can include one or more of current measurements, impedance measurements, tunneling measurements, field effect transistor measurements, and optical measurements.

[0390] The reference model can be stored in a memory.

[0391] Before the step of applying the reference model to a series of measurement values, the method can further include deriving a series of said measurement values by:

[0392] In the case of the number of measurement values in a previously unknown group, receiving, by a polymer, a series of raw measurement values, wherein a series of raw measurement value groups of the plurality of raw measurement values depends on the same k-mer, and

[0393] Processing the series of raw measurement values to identify consecutive groups of measurement values and deriving, for each identified group, different types of single measurement values or multiple measurement values to form the series of measurement values.

[0394] The method can further include acquiring, by a polymer, a series of raw measurement values.

[0395] In each of the plurality of series of measurement values, in the case of the number of measurement values in an unknown group, the groups of the plurality of measurement values can depend on the same k-mer.

[0396] The method can further include acquiring, by a polymer, the series of measurement values.

[0397] Sequence Listing

[0398] Seq ID 1: MS-(B1)8 = MS-(D90N / D91N / D93N / D118R / D134R / E139K)8

[0399]

[0400] Seq ID 2: MS-(B1)8 = MS-(D90N / D91N / D93N / D118R / D134R / E139K)8

[0401]

[0402] Seq ID 3: MS-(B2)8 = MS-(L88N / D90N / D91N / D93N / D118R / D134R / E13:9K)8

[0403] Seq ID 4: MS-(B2)8 = MS-(L88N / D90N / D91N / D93N / D118R / D134R / E13:9K)8

[0404] Seq ID: 5 (WT EcoExo I):

[0405]

[0406]

[0407] Seq ID: 6 (E. coli exonuclease II):

[0408]

[0409] Seq ID: 7 (Thermus thermophilus RecJ):

[0410]

[0411] Seq ID: 8 (λ exonuclease):

[0412]

[0413] seq ID: 9 (Phi29 DNA polymerase):

[0414]

Claims

1. A method of classifying polymers, each of the polymers comprising a sequence of polymer units, the method using a system comprising: a sample chamber containing a polymer-containing sample, a collection chamber isolated from the sample chamber, and a sensor element comprising a nanopore that communicates between the sample chamber and the collection chamber, The method comprises causing successive polymers to shift from the sample chamber through the nanopore, and during the shift of each polymer: acquiring successive measurements of the polymer from the sensor element; using reference data derived from at least one reference sequence of polymer units to analyze a series of the measurements taken from the polymer during a partial shift of the polymer to provide a measure of the similarity between the sequence of polymer units of the partially shifted polymer and the at least one reference sequence; selectively completing the shift of the polymer to the collection chamber or ejecting the polymer back into the sample chamber according to the measure of similarity.

2. The method according to claim 1, wherein, The system comprises a plurality of collection chambers, and for each collection chamber, a sensor element comprising a nanopore that provides communication between the sample chamber and each collection chamber, and the method is carried out with respect to a plurality of sensor elements in parallel.

3. The method according to claim 2, wherein, the method is carried out using different reference data for different nanopores or the method is carried out using the same reference data for different nanopores, and the step of selectively completing the shift of the polymer to the collection chamber or alternatively ejecting the polymer back into the sample chamber is carried out with different dependencies on the measures of similarity for the different nanopores.

4. The method according to any one of claims 1 to 3, wherein, The step of causing successive polymers to shift from the sample chamber through the nanopore comprises applying a bias voltage sufficient to initiate the shift, and the step of ejecting the polymer back into the sample chamber comprises applying an ejection bias voltage sufficient to eject the polymer.

5. The method according to any one of claims 1 to 3, wherein After the step of analyzing a series of the measurements taken from the polymer during the partial shift to provide a measure of similarity, no further analysis of the measurements is carried out if the step of completing the shift of the polymer to the collection chamber is carried out.

6. The method according to any one of claims 1 to 3, wherein After the step of analyzing a series of the measurements taken from the polymer during the partial shift to provide a measure of similarity, the shift is carried out at an increasing rate if the step of completing the shift of the polymer to the collection chamber is carried out.

7. The method according to any one of claims 1 to 3, wherein The at least one reference sequence of polymer units is a desired sequence, the reference data is derived from the at least one reference sequence of polymer units, and the step of selectively completing the shift of the polymer to the collection chamber is carried out in response to the measure of similarity, the measure of similarity indicating that the partially shifted polymer is the desired sequence.

8. The method according to any one of claims 1 to 3, wherein, The reference data from at least one reference sequence of polymer units represents actual or simulated measurement values acquired by a biochemical analysis system, and the step of analyzing a series of the measurement values of the polymer acquired during a partial shift includes: comparing the series of the measurement values with the reference data.

9. The method according to any one of claims 1 to 3, wherein, The reference data from at least one reference sequence of polymer units represents a feature vector of temporal characteristics, and the feature vector of the temporal characteristics represents characteristics of the measurement values acquired by the biochemical analysis system, and the step of analyzing a series of the measurement values acquired from the polymer during a partial shift includes: deriving a feature vector of temporal characteristics representing characteristics of the measurement values from the series of the measurement values, and comparing the derived feature vector with the reference data.

10. The method according to any one of claims 1 to 3, wherein the reference data from at least one reference sequence of polymer units represents the consistency of the polymer units of the at least one reference sequence, and the step of analyzing a series of the measurement values acquired from the polymer during a partial shift includes: analyzing a series of the measurement values to provide an estimated value of the consistency of the polymer units of the sequence of polymer units of the polymer that has been partially shifted, and comparing the estimated value with the reference data to provide a measure of the similarity.

11. The method according to any one of claims 1 to 3, wherein the measurement values depend on k-mers, where a k-mer is k polymer units of a polymer and k is an integer; the reference data represents a reference model that processes the measurement values as observations of a reference series corresponding to the k-mer states of the reference sequence of polymer units, wherein the reference model includes: a transition weight for transitions between the k-mer states in the reference series of k-mer states; and for each k-mer state, an emission weight for different measurement values when observing the k-mer state, and the step of analyzing a series of the measurement values acquired from the polymer during the partial shift includes fitting the model to the series of the measurement values to provide a measure of the similarity as a fit of the model to the series of the measurement values.

12. The method according to any one of claims 1-3, wherein, The nanopore is a biological pore, and / or the polymer is a polynucleotide and the polymer unit is a nucleotide.

13. The method according to any one of claims 1-3, wherein, The translocation of the polymer through the nanopore occurs in a ratchet manner.

14. The method according to any one of claims 1-3, wherein, The measurement values include electrical measurement values.

15. A system for classifying polymers, each polymer comprising a sequence of polymer units, the system comprising: a sample chamber for containing a sample that contains the polymer; a collection chamber isolated from the sample chamber; and a sensor element including a nanopore that communicates between the sample chamber and the collection chamber, wherein the system is arranged to cause successive polymers to start translocating from the sample chamber through the nanopore, and during the translocation of each polymer: The system is arranged to collect successive measurements of the polymer from the sensor element; The system is arranged to analyze a series of measurements taken from the polymer during a partial displacement of the polymer using reference data from at least one reference sequence of polymer units to provide a measure of the similarity between the sequence of polymer units of the partially displaced polymer and the at least one reference sequence, and depending on the measure of similarity, the system is arranged to selectively complete the displacement of the polymer into the collection chamber or alternatively discharge the polymer back into the sample chamber.

Citation Information

Patent Citations

  • Electrical device with detachable components

    GB201418512D0

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Molecular and atomic scale evaluation of biopolymers

    WO2000079257A1

  • Suspended carbon nanotube field effect transistor

    WO2005124888A1

  • High-resolution molecular graphene sensor comprising an aperture in the graphene layer

    WO2009035647A1

Cited By

  • Analysis of polymers

    CN120666007A

  • Analysis of polymers

    CN120666008A