Biochemical analysis system and method for controlling biochemical analysis system

By analyzing the real-time measurement results of polymers in a nanopore system and using a machine learning classifier to classify based on the modification characteristics and sequence information of the polymer, the problems of slow analysis speed and difficulty in effectively classification in the prior art are solved, and efficient enrichment and depletion of the polymer of interest are achieved.

CN120202306APending Publication Date: 2025-06-24OXFORD NANOPORE TECH LTD
View PDF 26 Cites 0 Cited by

Patent Information

Application Number
CN202380072557.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-01
Filing Date
2023-10-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing biochemical analysis systems using nanopores are difficult to improve the analysis speed when analyzing polymer sequences, and it is difficult to effectively classify and select the polymer units of interest.

Method used

By analyzing the real-time measurement results of the polymer in nanopores, it is determined whether to continue sequencing or reject the polymer based on the modification characteristics of the polymer. Using a machine learning classifier to combine modification information and sequence information, polymers are classified into different categories and a biochemical analysis system is operated based on the classification results.

Benefits of technology

Improves the enrichment and depletion of polymers of interest by nanopore systems, increases analysis speed and sequencing yield, and reduces sequencing time for polymers of interest.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202306A_ABST
    Figure CN120202306A_ABST
Patent Text Reader

Abstract

A method is provided that controls a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units. The system is operable to obtain continuous measurements of a polymer from a sensor element during displacement of the polymer relative to a nanopore of the sensor element. The method comprises analyzing measurements obtained from a polymer during its partial displacement to determine modification information about a portion of the sequence of polymer units when the polymer has partially displaced through the nanopore. Based on the modification information, the polymer is classified as belonging to one of class groups; and operating the system to reject the polymer or continue to acquire measurements from the polymer based on the category to which the polymer units are classified.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to identifying units of a polymer chain, and more particularly, to controlling an analysis system to identify a desired polymer chain.

[0002] There are many types of biochemical analysis systems that provide measurement results of polymer units for the purpose of determining a sequence. By way of example and not limitation, one type of measurement system uses a nanopore. Biochemical analysis systems using nanopores have been the subject of a great deal of recent development. Typically, during translocation of a polymer through a nanopore, continuous measurement results of the polymer are obtained from a sensor element that includes the nanopore. Some characteristics of the system depend on the polymer units in the nanopore, and measurement results of that characteristic are obtained. This type of measurement system using a nanopore has significant promise, particularly in the field of sequencing polynucleotides such as DNA or RNA.

[0003] Such biochemical analysis systems using nanopores can provide long continuous reads of a polymer, for example, in the case of a polynucleotide ranging from hundreds to tens of thousands (and possibly more) of nucleotides. Data collected in this way includes measurement results such as measurement results of an ionic current, where each translocation of the sequence through the sensitive portion of the nanopore results in a slight change in the measured characteristic. While such biochemical analysis systems using nanopores can provide significant advantages, there is still a desire to increase the analysis speed.

[0004] According to an aspect of the present invention, there is provided a method of controlling a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, wherein the biochemical analysis system includes at least one sensor element including a nanopore, and the biochemical analysis system is operable to obtain continuous measurement results of the polymer from the sensor element during translocation of the polymer relative to the nanopore of the sensor element, the method comprising: when the polymer has partially translocated through the nanopore, analyzing the measurement results obtained from the polymer during its translocation to determine modification information regarding a portion of the sequence of the polymer units, the modification information representing an estimated sequence of the modification state of the subject polymer unit regarding a modified portion of at least one classical type of polymer unit; classifying the polymer as belonging to one of a set of categories based on the modification information; and operating the biochemical analysis system to reject the polymer or continue to obtain measurement results from the polymer based on the category to which the polymer unit is classified as belonging.

[0005] The present invention analyzes the results of real-time measurements of polymer translocation in a nanopore (e.g., translocation of DNA / RNA molecules), and decides whether to continue sequencing the currently translocating polymer or to signal rejection of the relevant polymer from the nanopore based on the modification characteristics of the polymer. Rejecting polymers of no interest allows the nanopore to spend (the overall limited) polymer-translocation time on other input molecules. This results in the nanopore translocating (e.g., sequencing) a greater subset of the desired input molecules based on the enrichment / depletion of a particular subset of the entire set of molecules loaded onto the nanopore-based biochemical analysis system. The modification state of a polymer unit can indicate a range of properties of interest (such as the biological origin of the polymer), and is thus a useful classification for enabling the rejection of polymers.

[0006] In some embodiments, the estimation of the modification state includes a score regarding the subject polymer unit. The use of the score allows the method to take into account the level of certainty in the estimation of the modification state.

[0007] In some embodiments, the method further includes determining sequence information that represents an estimate of the identity of polymer units that are part of the series of polymer units. Identifying the polymer units can provide further information to inform the determination of the modification state of the polymer units.

[0008] In some embodiments, the estimate of the polymer unit identity includes a score for each of a set of types of polymer units. The use of the score allows the method to take into account the level of certainty in the estimate of the polymer unit identity.

[0009] In some embodiments, the estimate of the polymer unit identity is an estimate of a group that includes classical types of polymer units, and determining the modification information includes analyzing the sequence information to determine the modification information. The sequence information is first determined as classical types of polymer units, and then the sequence information is used to determine the modification information, which reduces the number of categories into which the polymer units must be classified at each stage. This can improve the overall accuracy of identifying polymer units, for example when using machine learning techniques.

[0010] In some embodiments, the estimate of the polymer unit identity is an estimate of a group that includes classical types of polymer units and at least one classical type of polymer unit in one or more modified forms, wherein the modification information includes an estimate of the at least one classical type of polymer unit in the modified form. Determining the complete identity of the polymer unit in a single step reduces the complexity of the classification process, which may be beneficial for increasing the classification speed.

[0011] In some embodiments, the polymer is classified as belonging to one of the category groups based on the modification information and the sequence information. Using these two types of information can increase the amount of data available for classification, thereby improving its accuracy.

[0012] In some embodiments, the polymer is classified as belonging to one of the category groups based only on the modification information. Using only the modification information reduces the complexity of the classification process, which may advantageously provide increased speed and reduced memory requirements.

[0013] In some embodiments, the polymer is derived from an organism, and the categories of the category group are taxonomic domains or kingdoms. In some embodiments, the first category of the category group is a bacterial organism or bacterial organism type, and the second category of the category group is a eukaryotic organism or eukaryotic organism type. These divisions allow the method to distinguish common categories of polymer sources that may be of interest for sequencing from categories that may not be of interest.

[0014] In some embodiments, at least one of the categories in the category group represents a target sequence, and the step of operating the biochemical analysis system to reject the polymer or continue to obtain measurement results from the polymer includes operating the biochemical analysis system to continue to obtain measurement results when the machine learning classifier classifies the polymer unit as belonging to the category that is the at least one category representing the target sequence. This allows the system to continue sequencing the polymer when it determines that the polymer is of an interesting type.

[0015] In some embodiments, at least one of the categories in the category group represents a background sequence, and the step of operating the biochemical analysis system to reject the polymer or continue to obtain measurement results from the polymer includes operating the biochemical analysis system to reject the polymer when the machine learning classifier classifies the polymer unit as belonging to the category that is the at least one category representing the background sequence. This allows the system to reject the polymer when it determines that the polymer is of an uninteresting type or represents a background contaminant (thereby freeing the nanopore to sequence other polymers).

[0016] In some embodiments, the polymer is a polynucleotide, and the polymer unit is a nucleotide. Sequencing of polynucleotides is particularly relevant for many biological applications.

[0017] In some embodiments, the modification state is a methylation state. In some embodiments, the subject polymer unit is a cytosine nucleotide, and the methylation state is a state of being methylated to at least one of 5-methyl-cytosine or 5-hydroxymethyl-cytosine; and / or the subject polymer unit is an adenosine nucleotide, and the methylation state is a state of being methylated to 6-methyl-adenine. The methylation state (especially the methylation state of cytosine and / or adenosine) is a common modification of nucleotides and can indicate various biological factors, such as diseases or biological origins, and is thus particularly meaningful for classification.

[0018] In some embodiments, the modification state is an oxidation state. The oxidation state can also indicate changes in the polymer, which may be meaningful for classification and selection of polymers to be sequenced.

[0019] In some embodiments, the subject polymer unit includes a polymer unit that forms part of a predetermined motif of the polymer unit. A particular combination of polymer units may typically occur together or be particularly useful as an indicator of a class. Thus, selecting the motif of the polymer unit as the subject polymer unit can improve the accuracy of classification by relying on the combination of these polymer units.

[0020] In some embodiments, the predetermined motif of the polymer unit is a cytosine nucleotide followed by a guanine nucleotide in a nucleotide sequence along the 5' → 3' direction. The modification state of this particular nucleotide combination can highly predict the class (such as biological origin), and is thus a useful choice of motif.

[0021] In some embodiments, the at least one sensor element is operable to eject a polymer that is positively translocating through the nanopore, wherein operating the biochemical analysis system to reject the polymer includes operating the sensor element to eject the polymer from the nanopore and accepting another polymer in the nanopore. Ejecting the polymer releases the nanopore for sequencing another molecule of interest, thereby accelerating the sequencing of polymers of the desired class.

[0022] In some embodiments, the at least one sensor element is operable to eject a polymer that is positively translocating through the nanopore by applying an ejection bias voltage sufficient to eject the polymer, wherein operating the sensor element to eject the polymer from the nanopore is performed by applying the ejection bias voltage. Since the polymer has a static charge, using a bias voltage can effectively eject the polymer. Bias voltages are also often used to drive polymers into the nanopore, so voltage inversion is a convenient way to provide an ejection function without extensive hardware modification.

[0023] In some embodiments, classifying the polymer as belonging to one of the category groups includes inputting the modification information into a machine learning classifier that classifies the polymer as belonging to one of the category groups based on the modification information. Compared to other techniques (such as comparing with a reference polymer sequence database), using a machine learning classifier allows for more accurate classification of previously unseen polymers or polymer variants. In some embodiments, the machine learning classifier includes a neural network. Neural networks have been found to be effective for this type of application.

[0024] In some embodiments, the machine learning classifier is trained using modification information for multiple categories. Using information from multiple categories increases the contrast of the training data available to the machine learning classifier, thereby improving its ability to distinguish between multiple categories and accurately identify the correct category of a particular polymer.

[0025] The present invention also provides a computer program comprising instructions that, when executed by a computer, cause the computer to perform the method according to the above aspects or any of its embodiments. A computer storage medium storing the computer program is also provided.

[0026] There is also provided a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, the biochemical analysis system comprising at least one sensor element including a nanopore, wherein the biochemical analysis system is operable to obtain continuous measurements of the polymer from the sensor element during translocation of the polymer relative to the nanopore of the sensor element; and wherein the biochemical analysis system is configured to perform the method according to the above aspects or any of its embodiments. The biochemical analysis system can be a portable biochemical analysis system.

[0027] For a better understanding, embodiments of the present invention will now be described by way of non-limiting examples with reference to the accompanying drawings, in which:

[0028] Figure 1 is a schematic diagram of a biochemical analysis system;

[0029] Figure 2 is a graph showing the variation of a typical measurement signal over time;

[0030] Figure 3 is a schematic diagram of a single sensor element including a nanopore;

[0031] Figure 4 is a flowchart of a method for controlling a biochemical analysis system;

[0032] Figure 5 is a flowchart of a method for determining modification information, which can be used in the Figure 4 analysis step;

[0033] Figure 6 is a flowchart showing further details of analyzing sequence information in the method of Figure 5 ;

[0034] Figure 7 shows generating sequence slices for determining modification information;

[0035] Figure 8 shows the use of the sequence slices by a neural network;

[0036] Figure 9 is a flowchart of an alternative method of the method of determining modification information Figure 5 ;

[0037] Figure 10 shows a comparison of the computational performance of the input data preprocessing steps for different read lengths and batch sizes when performed on a CPU and a GPU;

[0038] Figure 11 shows a comparison of the computational performance of classifying input data for different read lengths and batch sizes when performed on a CPU and a GPU;

[0039] Figure 12 shows an exemplary neural network structure for classifying polymers;

[0040] Figure 13 shows results demonstrating the performance of an exemplary classifier distinguishing DNA from plant and vertebrate sources based only on modification information;

[0041] Figure 14 shows results demonstrating the performance of an exemplary classifier distinguishing DNA from a range of sources based only on modification information;

[0042] Figure 15 shows results demonstrating the performance of an exemplary classifier distinguishing DNA from plant and vertebrate sources based on modification information and sequence information;

[0043] Figure 16 shows results demonstrating the performance of an exemplary classifier distinguishing DNA from a range of sources based on modification information and sequence information;

[0044] Figure 17 shows the difference in the read length distribution of DNA from human and bacterial sources in a first set of channels with and without using the method of the present invention; and

[0045] Figure 18 shows the difference in the read length distribution of DNA from human and bacterial sources in a second set of channels with and without using the method of the present invention.

[0046] The present invention relates to a method for controlling a biochemical analysis system. The biochemical analysis system can be a biochemical analysis system as described in WO2016 / 059427A1, which is incorporated herein by reference. The biochemical analysis system can be a portable biochemical analysis system.

[0047] Figure 1 is a schematic diagram of a biochemical analysis system 1 that can be controlled using the method of the present invention. The biochemical analysis system 1 is used for analyzing polymers and can also be used for sorting polymers. Returning to Figure 1 the biochemical analysis system 1 includes a sensor device 2 connected to an electronic circuit 4, which in turn is connected to a data processor 6.

[0048] The biochemical analysis system 1 includes at least one sensor element, as discussed further below. At least one sensor element can be included in the sensor device 2. The sensor element includes a nanopore, and the biochemical analysis system 1 is operable to obtain continuous measurements of the polymer from the sensor element during displacement of the polymer relative to the nanopore of the sensor element. The polymer contains a sequence (or series) of polymer units. The sensor device 2 derives a measurement signal 10 from the polymer, for example including the continuous measurements. The continuous measurements can hereinafter be referred to as the measurement signal 10. The data processor 5 analyzes the measurement signal 10 to derive information about the polymer.

[0049] In some preferred applications, the polymer is a polynucleotide (or nucleic acid), and the polymer unit is a nucleotide. However, generally, the polymer can be of any type, such as a polypeptide (such as a protein) or a polysaccharide. The polymer can be natural or synthetic. The polynucleotide can contain a homopolymer region. The homopolymer region can contain between 5 and 15 nucleotides.

[0050] In the case of a polynucleotide or nucleic acid, the polymer unit is a nucleotide. Nucleic acids are typically deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or synthetic nucleic acids known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleotide side chains. The PNA backbone consists of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone consists of repeating diol units linked by phosphodiester bonds. The TNA backbone consists of repeating threose linked together by phosphodiester bonds. LNA is formed from ribonucleotides having an additional bridge connecting the 2'-oxygen and 4'-carbon in the ribose moiety as discussed above. Nucleic acids can be single-stranded, double-stranded, or contain both single-stranded and double-stranded regions. Nucleic acids can contain one RNA strand hybridized to one DNA strand. Typically, cDNA, RNA, GNA, TNA, or LNA are single-stranded.

[0051] The polymer unit can be any type of nucleotide. Nucleotides can be naturally occurring or artificial. For example, the method can be used to verify the sequence of manufactured oligonucleotides. Nucleotides generally contain a nucleobase, a sugar, and at least one phosphate group. The nucleobase and the sugar form a nucleoside. The nucleobases are specifically adenine, guanine, thymine, uracil, and cytosine. The sugar is generally a pentose. Suitable sugars include, but are not limited to, ribose and deoxyribose. Nucleotides are generally ribonucleotides or deoxyribonucleotides. Nucleotides generally contain a monophosphate, diphosphate, or triphosphate.

[0052] The polymer unit can be a classical polymer unit. For example, in the case where the polymer is a DNA polynucleotide, the classical bases are adenine (A), cytosine (C), guanine (G), and thymine (T). In contrast, ribonucleic acid (RNA) contains the classical bases A, C, and G, where uracil (U) replaces thymine.

[0053] Nucleotides can be modified polymer units, such as damaged bases or epigenetic bases. For example, nucleotides can include pyrimidine dimers. Such dimers are generally associated with damage caused by ultraviolet light and are a major cause of cutaneous melanoma. Nucleotides can be labeled or modified to serve as labels with unique signals. This technique can be used to identify deletions of bases, such as abasic units or spacers in polynucleotides. The method can also be applied to any type of polymer.

[0054] In the case of a polypeptide, the polymer unit can be a naturally occurring or synthetic amino acid. In the case of a polysaccharide, the polymer unit can be a monosaccharide.

[0055] Particularly in the case where the sensor device 2 includes a nanopore and the polymer includes a polynucleotide, the length of the polynucleotide under study generally ranges from 500 nucleotides (500b) to a length greater than 2 Mb. However, shorter lengths of polynucleotides can be measured, with the lower limit estimated to be approximately 10 - 20 bases (depending on the length of the nanopore channel), which can include mRNA, tRNA, and cfDNA.

[0056] In the case where the polymer includes a polynucleotide, the polynucleotide is preferably a natural polynucleotide molecule, i.e., a polynucleotide molecule that has not been processed using amplification techniques such as polymerase chain reaction (PCR). This is because amplification processes such as PCR may remove or mask chemical modifications of nucleobases. However, if the amplification allows the modified state of the polymer unit to be retained, the method can be used for polymers that have undergone an amplification process.

[0057] The properties of the exemplary sensor device 2 and the resulting measurement signal 10 are as follows.

[0058] The sensor device 2 is a nanopore system comprising one or more nanopores. In a simple type, the sensor device 2 has only a single nanopore, but more practical measurement systems employ a number of nanopores (typically in an array) to provide parallel information collection. The measurement system may include at least 10 nanopores, optionally at least 100 nanopores, optionally at least 1000 nanopores.

[0059] The measurement signal 10 can be recorded during the translocation of the polymer relative to the nanopore (typically through the nanopore).

[0060] A nanopore is a pore that typically has nanoscale dimensions and can allow a polymer to pass therethrough. The nanopore can be a protein pore or a solid-state pore. The dimensions of the pore can be such that only one polymer can be translocated through the pore at a time.

[0061] When the nanopore is a protein pore, it can have the following characteristics.

[0062] The biological pore can be a transmembrane protein pore. The transmembrane protein pore for use according to the present invention can be derived from a β-barrel pore or an α-helix bundle pore. The β-barrel pore includes a barrel or channel formed by β-strands. Suitable β-barrel pores include but are not limited to β-toxins such as α-hemolysin, anthrax toxin, and leukocidin; and outer membrane proteins / porins of bacteria such as Mycobacterium smegmatis porin (Msp) (e.g., MspA, MspB, MspC, or MspD), pore-forming toxins, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP). The α-helix bundle pore includes a barrel or channel formed by α-helices. Suitable α-helix bundle pores include but are not limited to inner membrane proteins and outer membrane proteins such as WZA and ClyA toxins. The transmembrane pore can be derived from Msp or from α-hemolysin (α-HL). The transmembrane pore can be derived from lysine. Suitable pores derived from pore-forming toxins are disclosed in WO 2013 / 153359. Suitable pores derived from MspA are disclosed in WO-2012 / 107778. The pore can be derived from CsgG, such as disclosed in WO-2016 / 034591 and WO2019 / 002893, both of which are incorporated herein by reference in their entirety. The pore can be a DNA origami pore.

[0063] The protein pore can be a naturally occurring pore or can be a mutant pore.

[0064] Protein pores can be inserted into an amphiphilic layer, such as a biological membrane, e.g., a lipid bilayer. An amphiphilic layer is a layer formed by amphiphilic molecules having hydrophilic and lipophilic properties, such as phospholipids. The amphiphilic layer can be a monolayer or a bilayer. The amphiphilic layer can be a block copolymer, such as those disclosed in Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450; WO2014 / 064444 or US6723814, which are incorporated herein by reference in their entirety. Alternatively, protein pores can be inserted into orifices provided in a solid state layer, e.g., as disclosed in WO2012 / 005857.

[0065] A suitable device for providing a nanopore array is disclosed in WO-2014 / 064443. Nanopores can be provided across corresponding pores, where electrodes electrically connected to an ASIC are provided in each of the corresponding pores for measuring the current flowing through each nanopore. Suitable current measurement means can include a current sensing circuit as disclosed in WO-2016 / 181118.

[0066] The nanopores can include orifices formed in a solid state layer, which can be referred to as solid state pores. The orifices can be wells, gaps, channels, grooves or slits provided in the solid state layer through or into which an analyte can pass. Such a solid state layer is not of biological origin. In other words, the solid state layer is not derived from or isolated from a biological environment such as an organism or a cell, or a synthetically manufactured version of a biologically available structure. The solid state layer can be formed from both organic and inorganic materials, including but not limited to microelectronic materials, insulating materials (such as Si3N4, A1203 and SiO), organic and inorganic polymers (such as polyamides), plastics (such as Teflon®) or elastomers (such as two-component addition-curing silicone rubber) and glass. The solid state layer can be formed from graphene. Suitable graphene layers are disclosed in WO-2009 / 035647, WO-2011 / 046706 or WO-2012 / 138357. Suitable methods for preparing a solid state pore array are disclosed in WO-2016 / 187519.

[0067] Such solid state pores are typically orifices in a solid state layer. The orifices can be chemically or otherwise modified to enhance their properties as nanopores. The solid state pores can be used in combination with additional components that provide alternative or additional measurement results for polymers, such as tunneling electrodes (Ivanov AP et al., Nano Lett. January 12, 2011; 11(1): 279-85) or field effect transistor (FET) devices (as disclosed in, e.g., WO-2005 / 124888). The solid state pores can be formed by known methods, including, for example, those described in WO-00 / 79257.

[0068] The nanopore can be a hybrid of a solid-state pore and a protein pore.

[0069] The sensor device 2 obtains a series of measurement results of properties based on measurements that can be made on polymer units that are displaceable relative to the pore. This series of measurement results forms the measurement signal 10.

[0070] The property being measured can be related to the interaction between the polymer and the pore. This interaction can occur in the constriction region of the pore.

[0071] In one type of sensor device 2, the property being measured can be the ionic current flowing through the nanopore. These and other electrical properties can be measured using standard single-channel recording devices such as described in Stoddart D et al., Proc Natl Acad Sci, 12;106(19):7702-7; Lieberman KR et al., J Am Chem Soc.2010;132(50):17961-72 and WO-2000 / 28312. Alternatively, a multi-channel system (such as described in WO-2009 / 077734, WO-2011 / 067559 or WO-2014 / 064443) can be used to obtain measurement results of the electrical properties.

[0072] An ionic solution can be provided on either side of the membrane or solid-state layer, and this ionic solution can be present in corresponding compartments. A sample containing the polymer analyte of interest can be added to one side of the membrane and allowed to move relative to the nanopore, for example under a potential difference or a chemical gradient. The measurement signal 10 can be derived during the movement of the polymer relative to the pore, for example acquired during the translocation of the polymer through the nanopore. The polymer can partially translocate the nanopore.

[0073] To allow measurements to be taken as a polymer translocates through a nanopore, the translocation rate can be controlled by a polymer-binding moiety. Typically, this moiety moves the polymer through the nanopore under the influence of, or against, an applied field. The moiety can be a molecular motor, for example used in cases where the moiety is an enzyme, enzyme active, or used as a molecular gate. When the polymer is a polynucleotide, a variety of methods for controlling the translocation rate have been proposed, including the use of polynucleotide-binding enzymes. Suitable enzymes for controlling the translocation rate of polynucleotides include, but are not limited to, polymerases, helicases, exonucleases, single-stranded and double-stranded binding proteins, and topoisomerases such as gyrase. For other polymer types, a moiety that interacts with that polymer type can be used. The polymer-interacting moiety can be any binding moiety disclosed in WO-2010 / 086603, WO-2012 / 107778, and Lieberman KR et al., J Am Chem Soc. 2010;132(50):17961-72), as well as voltage gating schemes (Luan B et al., Phys Rev Lett. 2010;104(23):238103). The translocation rate of the polymer through the nanopore can be controlled by voltage control pulses to step the polymer through the nanopore, such as disclosed in WO2019 / 006214. The translocation of the polymer can be controlled by a molecular funnel, such as disclosed in WO2020 / 016573.

[0074] The polymer-binding moiety can be used to control polymer movement in a variety of ways. The moiety moves the polymer through the nanopore under the influence of, or against, an applied field. Polynucleotide-binding enzymes do not need to display enzyme activity as long as they can bind to the target polynucleotide and control its movement through the pore. For example, the enzyme can be modified to eliminate its enzyme activity, or it can be used under conditions that prevent it from acting as an enzyme. These conditions will be discussed in more detail below.

[0075] The polynucleotide-binding enzyme can be Dda helicase, such as disclosed in WO2015055981, which is hereby incorporated by reference in its entirety.

[0076] Polymer translocation through the nanopore can occur cis to trans or trans to cis, either under the influence of, or against, an applied electric potential applied through the electronic circuit 4 discussed below. The translocation can occur at an applied electric potential that can control the translocation. During the translocation of the polynucleotide through the nanopore under an applied electric potential, the binding enzyme typically docks at the cis or trans opening of the nanopore.

[0077] Exonucleases that act gradually or continuously on double-stranded DNA can be used at the applied electric potential on the cis side to pass the remaining single strand through the pore, or at the reverse electric potential for the trans side. Similarly, helicases that unwind double-stranded DNA can be used in a similar manner. There is also the possibility of sequencing applications that require strand translocation against the applied electric potential, but the DNA must first be "captured" by the enzyme at the reverse electric potential or zero potential. In the case where the electric potential is switched back after binding, the strand will pass from cis to trans through the pore and be held in the extended conformation by the current. A single-stranded DNA exonuclease or a single-stranded DNA-dependent polymerase can act as a molecular motor to pull the most recently translocated single strand back through the pore against the applied electric potential in a controlled stepwise manner (trans to cis). Alternatively, a single-stranded DNA-dependent polymerase can act as a molecular gate that slows down the movement of the polynucleotide through the pore. Any of the parts, techniques, or enzymes described in WO-2012 / 107778 or WO-2012 / 033524 can be used to control polymer movement.

[0078] However, the sensor device 2 can be an alternative type that includes one or more nanopores.

[0079] Similarly, the measured property may be of a type other than an ion current. Some examples of alternative types of properties include, but are not limited to: electrical properties and optical properties. Suitable optical methods involving fluorescence measurements are disclosed in J. Am. Chem. Soc. 2009, 131, 1652-1653. Possible electrical properties include: ion current, impedance, tunneling properties such as tunneling current (e.g., as disclosed in Ivanov AP et al., Nano Lett. January 12, 2011; 11(1):279-85), and FET (field effect transistor) voltage (e.g., as disclosed in WO2005 / 124888). One or more optical properties can be used, optionally in combination with electrical properties (Soni GV et al., Rev Sci Instrum. January 2010; 81(1):014301). The property can be a transmembrane current, such as an ion current flowing through a nanopore. The ion current can generally be a DC ion current, although in principle the alternative is to use an AC current (i.e., the magnitude of the AC current flowing under an applied AC voltage).

[0080] In some types of sensor devices 2, the measurement signal 10 can be characterized as including the measurement results from a sequence of events, where each event provides a set of measurement results. Figure 2A typical example of such a measurement signal 10 in the case of measuring current is shown. The groups of measurement results from each event have a similar level, although there will be some variations. This can be considered a noisy staircase wave, with each step corresponding to an event. These events may have biochemical significance, for example, generated by a given state or interaction of the sensor device 2. In some cases, this may be produced by a polymer shifting through a nanopore in a ratchet manner. However, this type of signal is not generated by all types of measurement systems, and the methods described herein do not depend on the signal type. For example, when the shifting rate is close to the measurement sampling rate, e.g., obtaining measurement results at 1 times, 2 times, 5 times, or 10 times the shifting rate of polymer units, compared to a slower sequencing speed or a faster sampling rate, the events may be less obvious or non-existent.

[0081] In addition, in the presence of events, there is typically no prior knowledge about the number of measurements in a group, which varies unpredictably. These factors of variation and the lack of knowledge about the number of measurements may make it difficult to distinguish some groups, for example, when the groups are short and / or the measurement levels of two consecutive groups are close to each other.

[0082] The groups of measurement results corresponding to each event typically have a level that is consistent on the time scale of the event, but for most types of sensor devices 2, variations occur on a short time scale. This variation may be due to measurement noise, for example, generated by the circuit and signal processing, especially by the amplifier in the case of electrophysiology. Due to the small magnitude of the characteristics of the measurement, this measurement noise is inevitable. This variation may also be due to variations or diffusion inherent in the underlying physical or biological system of the sensor device 2, such as variations in interactions, which may be caused by conformational changes of the polymer.

[0083] Most types of sensor devices 2 will experience this inherent variation more or less. For any given type of sensor device 2, both sources of variation may be at play, or one of these noise sources may dominate.

[0084] As the sequencing rate (i.e., the rate at which polymer units shift relative to the nanopore) increases, the events may become less obvious and thus more difficult to identify, or may disappear. Therefore, analytical methods that rely on detecting such events may become less efficient as the sequencing rate increases.

[0085] However, the methods disclosed herein do not rely on detecting such events. The methods described below are effective even at relatively high sequencing rates, including sequencing rates at which polymers shift at a rate of at least 10 polymer units per second, preferably 100 polymer units per second, more preferably 500 polymer units per second, or more preferably 1000 polymer units per second.

[0086] The sampling rate is the rate of measurement in the signal. Typically, the sampling rate is higher than the sequencing rate. For example, the sampling rate can be in the range of 100 Hz to 30 kHz, but this is not restrictive. In practice, the sampling rate may depend on the nature of the sensor device 2.

[0087] Returning to Figure 1 , the electronic circuit 4 will now be discussed. The electronic circuit can be the electronic circuit 4 disclosed in WO2016059427, which is incorporated herein by reference in its entirety. The electronic circuit 4 is arranged to control the application of a bias voltage across each sensor element of the sensor device 2. During normal operation, the bias voltage is selected such that the polymers can shift through the pores of the sensor elements. Such a bias voltage can typically be at a level of up to -200 mV.

[0088] The bias voltage provided by the electronic circuit 4 can also be selected such that it is sufficient to expel the shifted polymers from the pores. By causing the electronic circuit 4 to provide such a bias voltage, the sensor element is operable to expel the polymers that are positively shifting through the pores. To ensure reliable expulsion, the bias voltage is typically a reverse bias, although this is not always necessary.

[0089] An exemplary arrangement of the electronic circuit 4 is shown in Figure 3 . Figure 3 A single sensor element 230 of the sensor device 2 is shown, and an exemplary polymer 233 that is positively shifting through the pore 232 is shown. The sensor element 230 is made by forming a membrane 231 over the corresponding pore of the sensor device 2 and then inserting the pore 232 into the membrane 31. The membrane 231 seals the corresponding pore from the sample chamber of the sensor device 2.

[0090] The electronic circuit 4 is connected to electrodes 22, 25 that are connected on either side of the membrane 231. The electrode 25 can be a common electrode 25 that is shared by a plurality of sensor elements 230. The electrode 22 can be a sensor electrode 25 for the corresponding sensor element 230. The electronic circuit 4 controls the application of the bias voltage to create a bias between the electrodes 22, 25, thereby controlling the shift of the polymer 233 as described above. The sensor electrode 25 can be used to obtain an electrical measurement as the polymer 233 shifts through the pore 232, which can be used as the measurement signal 10 or processed to form the measurement signal 10.

[0091] Returning to Figure 1, a data processor 5 connected to the electronic circuit 4 is arranged as follows. The data processor 5 can be a computer device running an appropriate program, can be implemented by a dedicated hardware device, or can be implemented by any combination thereof. The computer device can be any type of computer system in use, but typically has a conventional structure. The computer program can be written in any suitable programming language. The computer program can be stored on a computer-readable storage medium, which can be of any type, such as: a recording medium that can be inserted into a drive of the computing system and can store information magnetically, optically, or magneto-optically; a fixed recording medium of the computer system, such as a hard disk; or computer memory. The data processor 5 can include a card that can be inserted into a computer, such as a desktop or laptop computer. The data used by the data processor 5 can be stored in its memory in a conventional manner.

[0092] The data processor 5 receives and processes continuous measurement results obtained using the sensor element 230 of the sensor device 2. That is, the data processor 5 receives the measurement signal 10. The data processor 5 stores and analyzes the continuous measurement results, as further described below.

[0093] Now will be described Figure 4 the method of controlling the biochemical analysis system 1 shown in

[0094] A biochemical analysis system (such as the biochemical analysis system 1 described above) can be configured to perform Figure 4 the method. Alternatively, a computer program can be provided, which includes instructions that, when the program is executed by a computer or a processor, cause the computer to perform Figure 4 the method. The computer program or instructions can be stored in the memory of a biochemical analysis system (such as the biochemical analysis system 1 described above). The processor can be the processor of the biochemical analysis system, such as the biochemical analysis system 1 described above. The computer program can be part of the biochemical analysis system firmware or interact with the biochemical analysis system firmware. Alternatively, the computer program can form a downstream component of the biochemical analysis system, which performs the analysis and classification steps and operates the biochemical analysis system by sending the required decisions back to the hardware / firmware of the biochemical analysis system.

[0095] Alternatively, a temporary or non-temporary computer-readable medium can be provided, which includes instructions that, when executed by the instruction processor, cause the processor to perform Figure 4 the method. The processor can be the processor of the biochemical analysis system, such as the biochemical analysis system 1 described above. In some instances, the method is implemented in the data processor 5 of the biochemical analysis system 1 described above. This method can be performed in parallel with respect to each sensor element 230 from which continuous measurement results of the polymer are obtained.

[0096] In step C1, the biochemical analysis system 1 is operated to apply a bias voltage sufficient to shift the polymer over the pores of the sensor element 230, for example using the electronic circuit 4. Based on the output signal detected, for example using the electronic circuit 4, the shift is detected and the acquisition of the measurement signal 10 is started. Over time, a series of consecutive measurement results are acquired.

[0097] In some cases, the following method steps operate on a series of raw measurement results acquired by the sensor device 2. In other cases, the raw measurement results are preprocessed to derive a series of measurement results, which are used in the following method steps instead of the raw measurement results.

[0098] When the polymer has been partially shifted through the nanopore, i.e., during the shift, method step C2 is performed. At this time, a sequence of measurements acquired from the polymer during the partial shift is collected for analysis, and the method includes analyzing the measurement results acquired from the polymer during its partial shift in C2 to determine modification information 20 about a part of the sequence of the polymer unit. The measurement results acquired from the polymer during the partial shift may be referred to herein as a measurement result "block".

[0099] Method step C2 may be performed after a predetermined number of measurement results have been acquired such that the measurement result block has a predefined size, for example corresponding to up to 100 polymer units, optionally up to 500 polymer units, optionally up to 1000 polymer units, optionally up to 5000 polymer units. Alternatively, method step C2 may be performed after a predetermined amount of time has elapsed after the polymer has shifted through the nanopore (e.g., up to 10 seconds, optionally up to 5 seconds, optionally up to 2 seconds, optionally up to 1 second, optionally up to 0.5 second, optionally up to 0.3 second, optionally up to 0.2 second, optionally up to 0.1 second after the polymer has shifted through the nanopore). In the former case, the size of the measurement result block may be defined by a parameter that is initialized at the start of the run (e.g., at the start of the process of sequencing the polymer in the sample) but that changes dynamically such that the size of the measurement result block changes. The size of the measurement result block may be selected based on any suitable factor, such as the size required to reliably classify the polymer, as described in more detail below. The size of the block may vary depending on the application. Some applications may require larger blocks such that a larger part of the polymer can be considered, for example, when a lower signal-to-noise ratio is required, or when the polymer has a smaller modification footprint.

[0100] The modification information 20 represents an estimated sequence of the modification state of the subject polymer unit of that part.

[0101] This portion is the portion that is partially shifted through the hole. The subject polymer unit can be a single polymer unit of a polymer. Alternatively, the subject polymer unit can include polymer units that form part of a predetermined motif of the polymer unit. For example, when the polymer is DNA, the predetermined motif of the polymer unit can be a cytosine nucleotide followed by a guanine nucleotide in the nucleotide sequence in the 5' → 3' direction.

[0102] The modification status is the modification status with respect to at least one classical type of polymer unit. The classical type of polymer unit can be, for example, an unmodified polymer unit.

[0103] Estimation can include an estimation of each of the subject polymer units that can include this portion, but this is not necessary. For example, the estimated sequence can include an estimation of each of the plurality of k-mers in this portion. The estimation of the modification status can include a score regarding the subject polymer unit. This score can represent the probability that the subject polymer unit is modified. This score can be normalized, for example, to fall between 0 and 1 or expressed as a percentage.

[0104] Preferably, the modification status is a methylation status. Methylation is a common modification that can be used to distinguish the origin of polymers (especially biopolymers such as DNA). Methylation can also indicate other states such as disease conditions.

[0105] By way of non-limiting example, in the case where the polymer is DNA and the polymer unit is a nucleotide, then the classical polymer unit can be cytosine or adenosine. In this example, the subject polymer unit is a cytosine nucleotide and the methylation status is the status of being methylated to at least one of 5-methyl-cytosine or 5-hydroxymethyl-cytosine, and / or the subject polymer unit is an adenosine nucleotide and the methylation status is the status of being methylated to 6-methyl-adenine.

[0106] From a more general perspective, the modified bases 5-methylcytosine (5mC) and 5-hydroxymethyl-cytosine (5hmC) are well-known epigenetic markers that regulate the transcription of the genome (the mechanism that turns DNA replication on and off into messenger RNA (mRNA), which is involved in protein synthesis). Thus, methylation is a type of modification that the modification information 20 can represent and is important because it is generally most relevant to biology. In the case where methylation is the desired modification, then the modification information 20 can be referred to as methylation information.

[0107] Although methylation is a modification of particular interest, the method is not limited to determining methylation status. The modification status can alternatively or additionally be an oxidation state. For example, methylated cytosine (5-mC) is oxidized to 5-hydroxymethyl-cytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), and adenine (A) is methylated to N6-methyladenine (6-mA), which have been identified as important epigenetic regulators.

[0108] In the case where the polymer is RNA, modifications are more prevalent and recent studies have shown that it plays a role in regulating mRNA stability. The stability of mRNA affects the control of gene expression and can impact various cellular and biological processes. To date, hundreds of RNA modifications have been characterized and can be represented by modification information 20. Non-limiting examples include N6-methyladenosine (m6A), inosine (I), N6,2'-O-dimethyladenosine (m6Am), 8-oxo-7,8-dihydroxyguanosine (8-oxoG), pseudouridine (Ψ), 5-methylcytidine (m5C), and N4-acetylcytidine (ac4C), which have been shown to regulate mRNA stability and function.

[0109] The modification information 20 can represent a modification status including a single chemical modification relative to a classical type of polymer unit, such as the 5mC modification of the cytosine DNA base as mentioned above. Alternatively, the modification information 20 can represent a combination of such modifications for a single polymer unit or multiple such polymer units. For example, for DNA, combinations of modifications can include 5mC, 5mC + 5hmC, 5mC + 6mA, 5mC + 5hmC + 6mA. This further applies to considerations of chemical modifications of classical polymer units present in one or more specific contexts (e.g., 5mC modification in a CG context), across all contexts (i.e., all positions of classical bases in the sequenced polymer), and any combination thereof.

[0110] The method can further include determining sequence information 30, which represents an estimate of the identity of the polymer units of a portion of the sequence of polymer units. Similar to the estimate of the modification status, the estimate of the polymer unit identity can include a score for each of a set of types of polymer units.

[0111] Figure 5 An exemplary method for determining modification information 20 (including determining sequence information) is shown. Figure 5 The method can be performed in Figure 4Used in step C2 of the method. In this example, the method uses an initial machine learning system in step C10 to determine sequence information 30 including polymer units of a classical type. Thus, the estimation of the polymer unit identity is an estimation regarding a group including polymer units of a classical type. As discussed above, polymer units of a classical type can be, for example, unmodified polymer units.

[0112] Then, a subsequent machine learning system is used in step C11 to determine modification information 20 from the derived sequence information 30. This means that, in this example, determining the modification information 20 includes analyzing the sequence information 30 to determine the modification information 20.

[0113] More specifically, in Figure 5In step C10, the measurement signal 10 is provided as an input to an initial machine learning system that is trained to provide an output as sequence information 30. In general, the initial machine learning system can take any suitable form, but is typically a neural network. For example, the initial machine learning system can be a neural network of the type disclosed in the following documents: Hochreiter, S. and Schmidhuber, J., 1997. Long short-term memory. Neural computation, 9(8), pp. 1735-1780; Cho, K., Van Merriënboer, B., Bahdanau, D. and Bengio, Y., 2014. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259; Kriman, S., Beliaev, S., Ginsburg, B., Huang, J., Kuchaiev, O., Lavrukhin, V., Leary, R., Li, J. and Zhang, Y., May 2020. Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6124-6128). IEEE; or Teng, H., Cao, M.D., Hall, M.B., Duarte, T., Wang, S. and Coin, L.J., 2018. Chiron: translating nanopore raw signal directly into nucleotide sequence using deep learning. GigaScience, 7(5), and standard training techniques are applied thereto.

[0114] The sequence information 30 determined in step C10 can be a class output. It can represent an estimate of the identity of polymer units in the sequence between classes that contain a set of predefined classical polymer units. For example, in the case where the polymer units are DNA polynucleotides, the classical nucleotides can be the four bases, adenine (A), cytosine (C), guanine (G), and thymine (T). Generally, such a class output can be implemented as a probability vector between classes. However, for use in subsequent methods, a difficult decision is made. That is, the most likely class (e.g., the most likely classical polymer unit) is selected and represented in the sequence information 30.

[0115] Optionally, the initial machine learning system can also output an initial mapping (also known as an input mapping) 13 between the measurement signal 10 and the sequence information 30. Typically, such an initial mapping 13 is inherently generated during the operation of a machine learning system (such as a neural network). In nanopore base identification documents and the prior art, it is commonly referred to as a "shift table". Generally, this initial mapping 13 is discarded because the typically desired output is simply a sequence estimate. However, when needed, the initial mapping 13 can usually be obtained and output from the initial machine learning system 11.

[0116] The initial mapping 13 simply describes the starting position of each polymer unit of the sequence information 30 and the corresponding sample of the measurement signal 10. The initial mapping 13 can be encoded in several equivalent forms. For example, an index array of the length of the sequence information 30 and having elements corresponding to the positions of the samples of the measurement signal 10 will fully represent this mapping. Equivalently, the length of each polymer unit of the sequence information 30 (expressed as the number of signal positions) will completely describe this mapping in a more compact manner.

[0117] It is assumed that the position of a polymer unit within the measurement signal 10 is not before the position of the polymer unit. In other words, a later polymer unit in the sequence information 30 may not be assigned to a position earlier in the measurement signal 10. It is also assumed that each input sequence polymer unit is assigned a starting position within the signal array, which means that many signal positions can be assigned to a single sequence base, and this is typically the case.

[0118] As an alternative to outputting the initial mapping 13 from the initial machine learning system 11, the initial mapping 13 can be derived from the measurement signal 10 and the sequence information 30 itself. Some methods are described in the prior art regarding the generation of such sequence-to-signal mappings, such as described in the following documents: Stoiber, M.H. et al. De novo Identification of DNA Modifications Enabled by Genome-Guided Nanopore Signal Processing. bioRxiv (2016); or Simpson, Jared T. et al. “Detecting DNA cytosine methylation using nanopore sequencing.” nature methods 14.4 (2017): 407-410. Such methods can be applied here.

[0119] After determining the sequence information, Figure 5 the method proceeds to step C11, which includes analyzing the sequence information 30 to determine the modification information 20.

[0120] Figure 6 FIG. shows the method for analyzing the sequence information in step C11, which uses a fragment machine learning system to derive the modification information 20 from the sequence information 30. In this example, there are three inputs, namely 1) the measurement signal 10, 2) the sequence information 30, and 3) the input mapping 13 between the measurement signal 10 and the sequence information 30.

[0121] In the derivation step S1, two fragments are derived, namely 1) a sequence fragment 51 and a signal fragment 52, which are input into the fragment machine learning system. The sequence fragment 51 is derived from a fragment of the sequence information 30 around the subject polymer unit in the sequence of polymer units. The signal fragment 52 is a fragment of the measurement signal 10. Importantly, the sequence fragment 51 and the signal fragment 52 are mapped to each other through the input mapping 13 between the measurement signal 10 and the sequence information 30.

[0122] Summarized at a high level, the method involves directly inputting the sequence fragment 51 (which is a classical sequence) (i.e., the sequence indicating each polymer unit of the classical type in the indication part) and the measurement fragment 52 of the measurement signal 10 (which is the original measurement signal) into the fragment machine learning system. This can be referred to as multi-directional input. In contrast, known classical base recognition systems typically rely on a one-way neural network because only one form of data is input into the neural network, namely the original nanopore signal. To achieve multi-directional input, the sequence fragment 51 and the signal fragment 52 are presented in the manner further described below.

[0123] This method can be applied to a single subject polymer unit in sequence information 30, or repeatedly applied to multiple subject polymers that are all or any subset of the polymer units in sequence information 30. For example, the method can be performed on a subject polymer unit that forms part of a predetermined motif comprising multiple classical polymer units. Generally, a motif (a short pattern of polymer units, such as nucleotides) may include uncertain positions, thereby allowing the use of multiple polymer units or polymer units of variable width to identify relevant subject polymer units. For example, the "CG" motif (also known as a CpG site) is the most common motif that undergoes methylation in most mammals and can form a motif as used herein.

[0124] Examples of the derivation of sequence fragment 51 and signal fragment 52 in derivation step S1 will now be described in more detail. As mentioned above, sequence fragment 51 is derived from a fragment of sequence information 30 surrounding the subject polymer unit, while signal fragment 52 is a fragment of measurement signal 10, and sequence fragment 51 and signal fragment 52 are mapped to each other by input mapping 13. There are various ways to achieve this, for example as follows.

[0125] Measurement signal 10, sequence estimate 30, and input mapping 13 can be provided as a complete sequencing read corresponding to the entire portion of the sequence obtained during a partial shift. However, this may be relatively long, depending on the configuration of the system, for example, for some types of sensor devices 2, consisting of dozens to thousands of individual polymer units. However, derivation step S1 provides sequence fragment 51 and signal fragment 52 with corresponding lengths that are selected to provide suitable accuracy for fragment machine learning system 41. This can be all or part of the polymer portion that has been shifted through the nanopore during the partial shift (so far).

[0126] In one method, signal fragment 52 is a predetermined length of measurement signal 10 around the position in measurement signal 10 that maps to the subject polymer unit. In this case, after identifying the subject polymer unit within sequence estimate 30, the position in measurement signal 10 that assigns the subject polymer unit from input mapping 13 is measured. The center of this extension of measurement signal 10 is defined as the center of the region of interest. From this position, a fixed signal width is extracted using a user-defined range before and after this position.

[0127] In this case, the predetermined length of measurement signal 10 can be, for example, in the range of 20 sample points to 1000 sample points, such as 100 sample points. A larger length of measurement signal 10 can be more than 1000 sample points. Signal fragment 52 can be arranged symmetrically or asymmetrically around the sample point that maps to the subject polymer unit.

[0128] In addition to extracting the signal segment 52 from this region, the sequence segment 51 is also selected as the polymer unit that extends the signal segment 52 mapped by the input mapping 13. Therefore, the length of the sequence segment 51 varies for different subject polymer units.

[0129] In another method, the sequence segment 51 is a predetermined length of the sequence estimate 30, i.e., a predetermined number of polymer units. In this case, after the sequence segment 51 is extracted, the signal segment 52 is derived as the part of the measurement signal 10 mapped to the sequence segment 51 by the input mapping 13. Therefore, the length of the signal segment 52 varies for different subject polymer units.

[0130] In this case, the predetermined number of polymer units can be in the range of 1 polymer unit to 100 polymer units. The range of polymer units to be considered may depend on the type of nanopore used.

[0131] Optionally, the sequence segment 51 can be selected to account for nanopore kinetics as follows. When the translocation rate of the polynucleotide through the nanopore is controlled by an enzyme - form molecular brake, for example, it is thought that modified bases can affect enzyme kinetics, such as the kinetics of certain helicases unwinding double - stranded polynucleotides. In the case of a helicase as a binding enzyme that can be used to unwind double - stranded DNA and control the passage of the resulting single - stranded DNA strand through the nanopore, considering those nucleotides within the enzyme - binding region can further provide information about the signal.

[0132] Therefore, it may be valuable to provide such information to the nanopore modified - base detection algorithm. This can be achieved by a sequence segment 51 derived in such a way that one or more nucleotides of the sequence segment 51 are within the region of the enzyme that acts as a molecular brake to control the translocation of the polymer.

[0133] This can improve accuracy compared to providing a signal of the same size but without the base of interest in the molecular brake. Note that this may provide improved performance compared to alternative nanopore modified - base detection algorithms that attempt to provide this information through the raw nanopore signal summary, as signal - to - sequence assignment / alignment algorithms are often error - prone. As described in other parts, passing the raw nanopore signal to a neural network can allow for improved performance, bypassing the problem of sequence - to - signal alignment.

[0134] It has been shown that signal variations may be mainly affected by the interaction of nucleotides with one or more constrictions of the nanopore, which are regions of narrower cross - section in the nanopore lumen. See, for example, Butler et al., Proceedings of the National Academy of Sciences 105 (52), 20647 - 20652Figure 1 , which shows that the MspA nanopore has an internal narrow constriction at the D90N / D91N region; and that of WO2016 / 034591 Figure 1 and Figure 2 , which shows the internal constriction region of the CsgG nanopore. However, the interaction with other regions of the nanopore affects the signal, and nucleotides outside the nanopore are also considered to have an impact on the measured signal. In use, during the translocation of a polynucleotide through the nanopore under an applied electrical potential, the binding enzyme typically leans against the cis or trans opening of the nanopore. Thus, nucleotides just outside the nanopore cavity typically lie within the region of the binding enzyme. For example, dDA helicase serves as a polynucleotide binding enzyme, and CsgG serves as a nanopore, and the distance between the enzyme and the constriction is estimated to be between 10 bases and 14 bases (or approximately 100 to 140 signal points). The signal point measurements depend on several factors and may vary significantly from these values for other pore chemistries.

[0135] Figure 7 shows a specific method for generating a sequence fragment 51 in a suitable form for input into a fragment machine learning system that maps to a signal fragment 52. This process aims to maximize the information presented to the fragment machine learning system.

[0136] First, a first signal fragment 61 is extracted as a fragment of the sequence estimate 30, which, for non-limiting, illustrative purposes, has a specific nucleotide sequence in Figure 7 , and these nucleotides are different classical nucleotides selected from the four bases A, C, G, or T. In Figure 7 , in graphical form, the input mapping 13 is represented by dashes. In particular, according to the input mapping 13, each element (whether a nucleotide or a dash) of the first sequence fragment 61 corresponds to a corresponding sample point in the corresponding signal fragment 52.

[0137] In step E1, the first sequence fragment 61 is encoded as a second sequence fragment 62 by replacing each polymer unit with a corresponding k-mer, such that the second sequence fragment 62 is a sequence of k-mers corresponding to the corresponding polymer units in the first input fragment 61. Thus, compared to the first sequence fragment 61, the second sequence fragment 62 has the same length but an increased dimension, such that each element of the second sequence fragment 62 is a k-dimensional vector (in Figure 7 , k is 3, by way of non-limiting example). Each k-mer in the second sequence fragment 62 includes a set of k polymer units (arranged vertically in Figure 7 ), where k is a complex integer. Each k-mer includes a) the corresponding polymer unit (along Figure 7in the intermediate dimension), and b) (k-1) polymer units adjacent to the corresponding polymer unit in the sequence estimate 30. Figure 7 The (k-1) adjacent polymer units in are symmetric about the corresponding polymer unit, but as an alternative, the (k-1) adjacent polymer units can be selected asymmetrically. It should be noted that this encoding requires a fixed number of polymer units before and after the first signal segment 61 in order to construct k-mers.

[0138] This change from polymer units to k-mers effectively provides additional context information for individual polymers. These k-mers can be considered to represent the portions of the polymer that physically interact with the nanopore at specific positions within the signal during partial translocation, although this is only conceptual and may not be a complete description of any particular sensor device 2. Nevertheless, in the case of polymer translocation through the nanopore, k may have a selected value such that the length of the k-mer is greater than the length of the nanopore cavity through which the polymer translocates.

[0139] Using k-mers in this way has been shown to improve the accuracy of the estimates made by the fragment machine learning system. In general, k can have any value that provides this improvement, noting that increasing k increases the size of the data without significantly increasing the computational cost. In some instances, k can have a value in the range from 3 to 50, but higher values are also possible.

[0140] As an alternative, step E1 can be omitted such that the following steps are performed on the first sequence segment 61, although this may reduce the accuracy of the estimates made by the fragment machine learning system.

[0141] In step E2, the second sequence segment 62 is amplified to a third sequence segment 63 such that it has the same length as the signal segment 52. In this instance, the amplification is performed by repeat padding, which is graphically shown in Figure 7 as replacing dashes with the k-mer that precedes them. This amplification allows for the efficient design of the fragment machine learning system, as described below.

[0142] In step E3, the third sequence segment 63 is binary encoded into a final sequence segment 64, which is used as the input sequence segment 51 for the fragment machine learning system. The binary encoding encodes each polymer unit in binary format, in this instance using one-hot encoding ("1000" for A; "0100" for C; "0010" for G; "0001" for T; and "0000" for unknown or missing bases). For each position in the third sequence segment 63, the k length-4 vectors of the k polymer units of the k-mer are concatenated to form a vector of length 4k.

[0143] Provide sequence segments 51 and signal segments 52 of equal length as two-way inputs to a segment machine learning system. The segment machine learning system has been trained to provide an output 80 that represents an estimate of the modification state of a subject polymer unit relative to at least one classical type of polymer unit. The output 80 is a categorical output. That is, the output 80 estimates the identity of the subject polymer unit among a set of categories (i.e., between a modified form and at least one classical type of polymer unit in an unmodified form). Such a categorical output 80 can be implemented as a probability vector among the categories. The segment machine learning system is trained to maximize the probability of the correct output category and minimize the probability of the incorrect output category. Although there are other loss functions that can be applied to such a categorical output 80, cross-entropy loss is typically used in the segment machine learning system further described below to optimize the categorical output type. The sequence of outputs 80 for multiple subject polymer units can be directly used as modification information 20. Alternatively, the sequence of outputs 80 can be further processed to derive the modification information 20. For example, the modification information 20 can represent only the modified subject polymer units in the polymer.

[0144] The categories represented by the output 80 are classical polymer units and at least one modified form of the classical polymer unit.

[0145] Generally speaking, a segment machine learning system can use a variety of different machine learning techniques. However, a particularly advantageous form of the segment machine learning system is as a neural network.

[0146] By way of illustration, Figure 8 An example of a segment machine learning system as a neural network 70 is shown. The features or components of the neural network 70 and the training method of such a neural network will now be described.

[0147] The neural network 70 includes a first input stage 71 that receives the sequence segment 51 and a second input stage 72 that inputs the signal segment 52.

[0148] The first input stage 71 includes at least one first input neural network layer. One or more input neural network layers of the first input stage 71 can be one or more convolutional neural network layers.

[0149] The second input stage 72 also includes at least one second input neural network layer. One or more input neural network layers of the second input stage 72 can be one or more convolutional neural network layers.

[0150] The outputs of the first input stage 71 and the second input stage 72 are provided to a connection layer 73, which connects these outputs to provide a connected output that is provided to the remaining layers, which also include at least one convolutional neural network layer. The connection is by feature such that the temporal (sequencing signal time direction) correspondence between the input and the connection layer 73 derived from the sequence segment 51 and the signal segment 52 is retained. Then, the output values from the connection layer 53 are further processed by the layers in the neural network 50 as a single input.

[0151] The additional layers are arranged as follows.

[0152] The connected output 74 is provided to a combined convolutional neural network stage 76 that includes at least one convolutional neural network layer.

[0153] The convolutional neural network layers of the first input stage 71 and the second input stage 72 and the combined convolutional neural network stage 76 can have a conventional structure. Such convolutional neural network layers are well known in the art, but generally they operate on a fixed-size moving window of the input data along the input data with a certain stride. At each window, the input features are multiplied by a set of weights in a matrix multiplication to produce the output of each layer.

[0154] Each of the first input stage 71 and the second input stage 72 and the combined convolutional neural network stage 76 can include any number of convolutional layers stacked together, where different hyperparameters are applied in each layer, including the window size, the stride, and the number of parameters / weights. Each convolutional layer can be followed by a batch normalization layer and an activation function (in this case the swish non-linearity) as well as other standard neural network components. The convolutional layers in the first input stage 71 and the second input stage 72 are designed to produce the same output size in terms of length and feature dimension. Note that the inputs of each of the first input stage 71 and the second input stage 72 have different feature dimension sizes.

[0155] No padding is used for any of the convolutional layers, as is common when using convolutional layers in some areas of machine learning.

[0156] The output of the combined convolutional neural network stage 76 is provided to an LSTM stage 77 that includes at least one LSTM (long short-term memory) layer, which is an instance of a recurrent neural network (RNN) layer and can have a conventional structure.

[0157] The LSTM stage 77 is optional and can be omitted.

[0158] The output of the LSTM stage 77 or, in the case of omitting the LSTM stage, the output of the convolutional neural network stage 76 combined is provided to a fully connected stage including at least one fully connected layer, which can also have a conventional structure.

[0159] A description of the recurrent neural network layer that can be applied to the LSTM stage 77 and the fully connected stage is given in Sak, H., Senior, A.W. and Beaufays, F., 2014. Long short - term memory recurrent neural network architectures for large scale acoustic modeling.

[0160] The neural network 70 processes the input in batches. The cross - entropy loss for each batch is calculated as described above. During training, backpropagation is performed using an optimizer. In one demonstration, the optimizer can be the AdamW optimizer. Backpropagation is performed in the standard manner described in the prior art, as described in the prior art (Loshchilov, I. and Hutter, F., 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101).

[0161] Use the output 80 as an estimate of the modified state of the subject polymer unit. A sequence of such outputs 80 for multiple polymer units can be used as Figure 4 the modification information 20 in the method. Alternatively, this sequence of the output 80 can be processed to generate the modification information 20, for example, to generate a dense representation of the modified units of the polymer.

[0162] Figure 9 Shows an alternative method of analyzing Figure 4 the measurement results in step C2 to determine the modification information 20 regarding a portion of the sequence of polymer units.

[0163] In Figure 9 the method, the estimate of the identity of the polymer units represented by the sequence information 30 is an estimate regarding a group of at least one classical type of polymer unit including classical types of polymer units and one or more modified forms. Then, the modification information 20 includes an estimate regarding at least one classical type of polymer unit in the modified form. For example, in the case where the polymer is DNA, Figure 5The groups in the method can include the classical nucleobases adenine (A), cytosine (C), guanine (G), and thymine (T), as well as the modified bases 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC). Then the sequence information 30 will include an estimate for each of these classical and modified bases. The estimates for the modified bases 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) will be regarded as the modification information 20.

[0164] In this method, the identity of the polymer units is directly estimated as being of the classical (i.e., unmodified) type or a modified form of the classical type. This simplifies the estimation because, compared to Figure 5 the two-stage method, only a single-stage process is required to obtain the modification information 20. However, in some cases, the accuracy of this method may be reduced compared to the two-stage method because it is more difficult to distinguish between a large number of different categories in a single step. For Figure 5 the method, machine learning algorithms such as neural networks can be used to estimate the identity of the polymer units.

[0165] Returning to Figure 4 the method, after analyzing the measurements in step C2 to determine the modification information 20, the method includes a step C3 of classifying the polymer.

[0166] In step C3, the modification information 20 collected in step C2 is used to classify the polymer as belonging to one of the category groups based on this modification information 20. Classification can be performed by any suitable method. For example, the modification information 20 can be passed to a machine learning classifier, as described in more detail below. However, the method (and in particular the step of classifying the polymer) does not include comparing the modification information 20 with a reference polymer sequence.

[0167] It is possible to classify the polymer by comparing the modification information 20 with a database of known reference sequences and, for example, determining the similarity to an exemplary reference sequence of a particular category. However, this requires a large database of reference polymer sequences, which may consume a large amount of memory and processor time for comparison. In addition, using reference polymer sequences reduces the ability of the method to classify uncommon polymers that do not match well with any reference polymer sequence. Therefore, in the present invention, classification is performed based on the modification information 20 (e.g., solely based on the modification information 20) without using a reference polymer sequence.

[0168] When the polymer is derived from an organism, the categories of the category group can be taxonomic domains or kingdoms. For example, the first category of the category group can be bacterial organisms or bacterial organism types, and the second category of the category group can be eukaryotic organisms or eukaryotic organism types.

[0169] At least one class of the class group may represent a target sequence or a class of the target sequence. For example, the target sequence may be a polymer from a eukaryotic organism such as a mammal or a human. At least one class of the class group may represent a background sequence or a class of the background sequence. For example, the background sequence may be a polymer from a bacterial organism.

[0170] While these are common examples, the classes of the class group are not limited to being based on the organism from which the polymer is derived. For some applications, the classes may represent the type or condition of the polymer, such as the modification profile of the polymer. In such cases, polymers of different classes may be derived from organisms within the same taxonomic domain or kingdom, or even from the same organism. For example, as will be further discussed below, another advantageous application of the methods of the present invention is the detection of artificial methylation regions for protein binding site identification and affinity analysis. For this application, the classes may represent whether the polymer has a specific modification profile, in this instance an artificially induced methylation profile.

[0171] The target sequence and the background sequence are not specific reference sequences (such as from a database) against which the modification information 20 is compared. They are more like the types of sequences of interest (target sequences) or sequences of no interest (background sequences) that are sequenced by the biochemical analysis system 1.

[0172] Figure 9 The method can use the techniques disclosed in WO-2020 / 109773, which is incorporated herein by reference.

[0173] As mentioned above, the polymer can be classified by any suitable method. Classifying the polymer as belonging to one of the class group may include inputting the modification information 20 into a machine learning classifier that classifies the polymer as belonging to one of the class group based on the modification information 20. The machine learning classifier may include a neural network.

[0174] Preferably, the machine learning classifier is trained using the modification information 20 for multiple classes. This will improve the accuracy and precision of the classification by providing additional information to distinguish between the multiple classes. However, this is not necessary, and in some cases, the machine learning classifier can be trained using only the modification information 20 for a subset of less than all of the classes in the class group. The machine learning classifier can be trained using only the modification information 20 for a single class of the class group (such as the class representing the target sequence or the class representing the background sequence).

[0175] The method classifies polymers using a pre-trained machine learning classifier. The machine learning classifier is trained on data external to the data being sequenced. For example, the modification information 20 for training can be obtained from the analysis of samples of known composition or can be derived from an external database of previously sequenced polymers.

[0176] In some embodiments, a polymer can be classified as belonging to one of a category group based only on the modification information 20. Specific examples of such embodiments, where the modification information 20 includes the 5mC methylation probability of CG motifs in DNA, are as follows.

[0177] The 5mC CG-environment-local methylation probabilities are aggregated over all CG motifs within a portion of the sequence of the polymer obtained during a partial shift. In this example, the sequence order of the methylation probabilities is not important and thus the classification does not depend on the order of the modification states of the subject polymer units. However, in other examples, the classification may depend on the order of the modification states of the subject polymer units.

[0178] The integer range of 1 - 255 is divided into N equally sized intervals (e.g., N = 10 results in intervals [1 - 25][26 - 51]…[231 - 255], and the frequency of methylation probability values within a particular interval is calculated. For example, since 1 <= 13 <= 25, a methylation probability of 13 triggers an increment of the counter for interval 1 - 25. The count value is then normalized (i.e., divided) by: the total number of CG motifs in the observed sequence. The resulting N-dimensional real-valued vector between 0 and 1 (the sum of the N values is 1.0) represents the tabulated frequency spectrum of 5mC CG methylation probabilities, simply referred to as the methylation spectrum. In addition to the methylation spectrum, the portion of the sequence is further characterized by the total length (an integer) of the portion in the polymer unit and the number (an integer) of CG motifs. These 2 + N integers and real numbers are referred to as the sequence list methylation spectrum. Thus, the classification of a polymer depends on one or more (preferably all) of the following: the length of the portion, the number of subject polymer units (in this example, the number of CG motifs), and the estimated modification state of each subject polymer unit.

[0179] Although the number of CG motifs and the length of the portions are used in the representation of the sequence list methylation profile, this does not necessarily require the use of classical sequence information. The method can be designed to estimate the modification state (e.g., 5mC methylation probability) of each motif polymer unit (e.g., CG motif) of a portion of the sequence from the measured signal (as described above) only. Then, without the need for classical base identification of the portion of the sequence of the polymer unit, the resulting modification information 20 (e.g., methylation probability vector) provides both an estimate of the number of motif polymer units and the modification. In addition, the length of the portion of the sequence can be estimated from the measured signal without the need for classical base identification.

[0180] After the above preprocessing to obtain the sequence list methylation profile, the sequence list methylation profile is provided to a pre-trained machine learning classifier (e.g., the regularized gradient boosting framework XGBoost). The classifier outputs the probability that the polymer is bacterial or non-bacterial (e.g., in the form of a real value between 0.0 and 1.0). A 0.5 probability threshold is used to convert the real value output to a binary classification result, assigning a bacterial state (<0.5) or a non-bacterial state (>=0.5) to the polymer, respectively.

[0181] The computational performance (in seconds) of the preprocessing steps of obtaining the sequence list methylation profile and classification inference of the machine learning classifier is shown in Figure 10 and Figure 11 . Figure 10 Shows a comparison of the computational performance of the input data preprocessing steps for different read lengths and batch sizes when performed on a) GPU and b) CPU. Figure 11 Shows a comparison of the computational performance of classifying the input data for different read lengths and batch sizes when performed on a) GPU and b) CPU.

[0182] As an alternative to the above embodiments, in some embodiments, the polymer is classified as belonging to one of the category groups based on the modification information 20 and the sequence information 30. Non-limiting examples of such embodiments (similar to the above example where the modification information 20 includes the 5mC methylation probability of CG motifs in DNA) are as follows.

[0183] The sequence information of the classical DNA sequence that identifies the bases of the portion of the sequence of length L is encoded into an Lx4 matrix using one-hot encoding, where each row (indexed by i) of the L rows is a 0|1 bit vector and one of the 4 columns (indexed by j) has a value of 1 and the remaining columns have a value of 0. The position where the value is 1 indicates the classical base, where j = 1 (for A), j = 2 (for C), j = 3 (for G), and j = 4 (for T).

[0184] Then, the 5mC CG-context-local methylation probability is represented as an Lx1 floating-point methylation modification vector. For each classical base in the input sequence of length L, the values of all bases except the cytosine base in the CG context are set to 0.0. In the case of the CG context, a real / floating-point value is set to the probability of methylation of the corresponding cytosine base (according to the output of the inference method).

[0185] Then, the methylation modification vector and the sequence one-hot encoding matrix are concatenated to form an Lx5 matrix, where the first 4 columns represent the one-hot encoding of the input bases to identify the classical DNA sequence (as described above), and the 5th column represents the 5mC Cg methylation profile of the part of the sequence. The total Lx5 matrix is called the input representation.

[0186] Figure 12 Shows how the input representation is processed via an exemplary trained neural network. Figure 12 The neural network consists of: 2 convolutional + max-pooling layers, followed by a recurrent neural network layer with long short-term memory (LSTM) modules, followed by two fully connected layers, and finally a softmax layer with 2 output neurons suitable for binary classification.

[0187] Any suitable method or framework (e.g., the PyTorch programming framework) can be used to implement the preprocessing to obtain the input representation and the neural network architecture encoding. Similarly, any suitable framework method (e.g., the skorch programming framework) can be used for neural network training.

[0188] Returning again to Figure 4 , after classifying the polymer in step C3, the method includes operating the biochemical analysis system 1 to reject the polymer or continue to obtain measurement results from the polymer based on the category to which the polymer unit is classified. This is done through steps C4, C5, and C6.

[0189] In step C4, a decision is made based on the response to the classification of the polymer in step C3: (a) reject the polymer being measured (i.e., eject the partially displaced polymer from the nanopore), or (b) continue to obtain measurement results (i.e., sequence the polymer) until the polymer ends. After making a decision in step C4, a feedback signal is sent to the biochemical analysis system 1 to control the operation to continue obtaining measurement results or reject the polymer. The feedback signal can be called an enrichment feedback signal (if the decision is to continue obtaining measurement results) or a depletion feedback signal (if the decision is to reject the polymer).

[0190] In some instances, a decision may additionally be made that additional measurement results are needed to make a decision. For example, this option may be taken if the certainty of polymer classification is lower than a predetermined threshold. In such a case, the method will not proceed to step C5 or C6, but will return to step C2 after obtaining additional measurement results. Then, the analysis of the measurement results and the classification of the polymer will be repeated, and the decision step C4 will be performed again.

[0191] In the case where at least one category in the category group represents a target sequence or a class of the target sequence, operating the biochemical analysis system 1 to reject the polymer or continue to obtain measurement results from the polymer may include operating the biochemical analysis system 1 to continue to obtain measurement results when the machine learning classifier classifies the polymer unit into a category that belongs to the at least one category representing the target sequence. This means that if the polymer is identified as belonging to the category of reads that should be enriched, sequencing continues.

[0192] In the case where at least one category in the category group represents a background sequence or a class of the background sequence, operating the biochemical analysis system 1 to reject the polymer or continue to obtain measurement results from the polymer may include operating the biochemical analysis system 1 to reject the polymer when the machine learning classifier classifies the polymer unit into a category that belongs to the at least one category representing the background sequence. This means that if the polymer is identified as belonging to the category of reads that should be depleted, sequencing stops.

[0193] Other criteria may also be used to determine whether the polymer should be rejected or whether measurement results should continue to be obtained. For example, in the case where at least one category in the category group represents a target sequence or a class of the target sequence, operating the biochemical analysis system 1 to reject the polymer or continue to obtain measurement results from the polymer may include operating the biochemical analysis system 1 to reject the polymer when the machine learning classifier classifies the polymer unit into a category that is not the at least one category representing the target sequence. In other words, if the polymer cannot be positively identified as belonging to the category of the target sequence, the polymer is rejected. This may include if it may not be possible to positively identify the polymer as belonging to any category, for example, if the uncertainty in the classification is higher than a predetermined threshold.

[0194] Similarly, in the case where at least one category in the category group represents a background sequence or a class of the background sequence, operating the biochemical analysis system 1 to reject the polymer or continue to obtain measurement results from the polymer may include operating the biochemical analysis system 1 to continue to obtain measurement results from the polymer when the machine learning classifier classifies the polymer unit into a category that is not the at least one category representing the background sequence.

[0195] If the decision made in step C4 is (a) to reject the polymer being measured, the method proceeds to step C5, where the biochemical analysis system 1 is operated to reject the polymer so that a measurement result can be obtained from another polymer.

[0196] At least one sensor element 230 is operable to eject the polymer that is being shifted through the nanopore 232. In this case, operating the biochemical analysis system 1 to reject the polymer includes operating the sensor element 230 to eject the polymer from the nanopore 232 and to receive another polymer in the nanopore 232.

[0197] Any suitable method can be used to eject the polymer. For example, at least one sensor element 230 is operable to eject the polymer that is being shifted through the nanopore 232 by applying an ejection bias voltage sufficient to eject the polymer. In this case, the sensor element 230 is operated to eject the polymer from the nanopore 232 by applying the ejection bias voltage.

[0198] For example, considering the above biochemical analysis system 1, the electronic circuit 4 can apply a bias voltage to the nanopore 232 of the sensor element 230 that is sufficient to eject the currently shifting polymer 233. This ejects the polymer 233, thereby making the pore 232 available to receive another polymer. After such ejection in step C5, the method can return to step C1, and thus the electronic circuit can apply a bias voltage to the pore 232 of the sensor element 230 that is sufficient to shift another polymer through the pore 232. This method is particularly convenient because it can utilize the same electronic circuit 4 for applying the bias voltage to facilitate the shifting of the polymer for measurement. This eliminates the need to provide additional hardware functionality in order to be able to use the method of the present invention, thereby allowing the method to be deployed in existing devices.

[0199] In some alternative instances, in step C5, the biochemical analysis system 1 is made to stop obtaining measurement results from the currently selected sensor element 230 and instead obtain measurement results from a different sensor element 230. At the same time, in step C5, the electronic circuit 4 is controlled to apply a bias voltage to the pore 232 of the sensor element 230 that is sufficient to eject the polymer 233 that is currently shifting through the currently selected sensor element 230, so that the sensor element 230 is available to receive another polymer in the future. Then, the method returns to step C1, which is applied to the newly selected sensor element 230, such that the biochemical analysis system 1 begins to obtain measurement results therefrom.

[0200] If the decision made in step C4 is (b) to continue acquiring measurement results until the polymer ends, the method proceeds to step C6 without repeating steps C2 and C3, such that no additional data blocks are analyzed. In step C6, sensor element 1 continues to operate such that measurement results continue to be acquired until the polymer ends. Thereafter, the method returns to step C1 such that additional polymers can be analyzed.

[0201] If the decision made in step C4 is that additional measurement results are needed to make a decision, the method returns to step C2. Thus, measurement results of the shifted polymer are continued to be acquired until the next measurement result block is collected in step C2 and analyzed in step C3. The measurement result block collected when step C2 is performed again may be just new measurement results to be analyzed separately, or may be a combination of new measurement results and a previous measurement result block. Using a larger block size may reduce the advantage in terms of enrichment / depletion of the target / background sequence, but can improve the accuracy of identifying polymers to be enriched / depleted. However, the maximum block size is ultimately limited by the maximum size that the underlying hardware / firmware of the biochemical analysis system can effectively eject a portion of the shifted polymer from nanopore 232.

[0202] The method of the present invention is advantageous for many applications, particularly in the field of environmental DNA sample sequencing. Most environmental DNA samples contain a collection of cellular and cell-free DNA from mixed sources. For example, fecal samples typically contain host DNA mixed with bacterial DNA (which is from the host's microbiome and the surrounding environment). However, experiments performed on the samples typically target only the analysis of full-length host DNA (rather than the mixed bacterial DNA).

[0203] Host DNA analysis would benefit from an increase in the sequencing yield of total host DNA. Thus, improving the sequencing yield of host DNA and reducing the sequencing yield of non-host DNA provides substantial benefits. There are various wet laboratory / library preparation protocols (e.g., https: / / www.nature.com / articles / s41598-018-20427-9) for enriching host DNA molecules and / or depleting bacterial DNA molecules in the library to be sequenced, which results in better host DNA yields in standard whole genome sequencing (WGS) experiments.

[0204] The method of the present invention can achieve a similar improvement by increasing the ratio of the sequencing yield of host DNA to the sequencing yield of non-host DNA. This is achieved by classifying polymers as belonging to one of the category groups according to the method of the present invention, where the group of categories includes bacterial organisms of a first category and eukaryotic organisms of a second category. This can be achieved using modification information 20, which may include an estimate of the modification status (such as the methylation status).

[0205] For example, there are well-characterized differences in the 5-methyl-cytosine methylation (referred to as 5mC methylation) profiles between bacterial genomes and eukaryotic genomes. The genomes of eukaryotic organisms contain strong 5mC methylation modifications of cytosine bases in the CG motif, while bacterial genomes lack 5mC methylation in the CG motif. Bacterial genomes also contain methylation in certain other sequence motifs (e.g., GGWCC, where W = (G|C)), which occur much less frequently.

[0206] The ability to classify partially shifted DNA molecules in real time as bacterial or non-bacterial enables the method of the present invention to increase the yield of host DNA during sequencing and reduce the non-host DNA sequencing output. As discussed above, various types of modifications can be used for this classification. A common example is 5mC methylation, which can be used either alone or in combination with other modifications such as methylation to 6-methyl-adenine (6mA methylation). 6mA methylation is abundant in bacterial genomes and absent in mammalian genomes. Importantly, this method does not require prior knowledge of the host DNA composition or non-host DNA composition of the sample under consideration, which is required when classifying polymers using a reference sequence. Instead, it relies on general biological facts about modification information, such as the 5mC and 6mA methylation patterns between the DNA of specific targets.

[0207] This method is different from the previously mentioned wet-lab protocols, such as methylation-based host DNA enrichment for WGS sequencing. In particular, this method (i) is entirely computational and thus does not require wet-lab sample manipulation before being processed by the biochemical analysis system 1, and (ii) allows for the analysis of background sequences, which is different from existing wet-lab enrichment strategies. The reason for this latter difference is that even if polymers classified as background sequences are rejected to stop sequencing of undesired polymer molecules, the modification information (and sequence information if determined) associated with the polymer portions measured during the partial shift prior to rejection can still be retained. This information can then be further analyzed, for example, for taxonomic classification and / or taxonomic abundance estimation of the origin of non-host DNA species in the original sample. Thus, compared to existing methods, this method has a twofold advantage in providing enrichment of polymer classes of interest while also providing a great deal of information about background polymer classes.

[0208] This method also has advantages compared to methods that classify by comparing a sequence portion of a polymer to a reference polymer sequence. This is because methods that use comparison to a reference sequence require prior knowledge of the target sequence to be enriched to allow alignment-based sequence comparison to the reference sequence. Additionally, if there are a large number of classes representing the target sequences (e.g., analysis of DNA in a water sample aimed at enriching DNA of all native fish species), the number of required reference sequences is very large. This may mean that the number of reference sequences is too large to be computationally compared effectively, or the reference sequences for some or all of the target sequences may be of poor quality or even non-existent.

[0209] Another advantageous application of this method is artificial methylation regions for protein binding site identification and affinity analysis. Academic researchers have previously described several protocols for mapping genome-wide protein-DNA interactions using single molecule long-read sequencing (e.g., DiMeLo-seq; https: / / www.nature.com / articles / s41592-022-01475-6). In the above protocols, 6mA methylation modification is artificially induced in the region of interest of the protein (where the protein interacts with the DNA molecule). Then, the DNA molecule is sequenced in a WGS manner using, for example, a nanopore sequencing platform without an adaptive sampling setting. Subsequently, 6mA methylation is inferred on the fully sequenced and base-called DNA molecule. Then, the sequenced molecules are aligned based on a reference, and the 6mA methylation frequency of each region of the reference genome sequence is considered. Due to the fact that, for example, human DNA does not have natural 6mA methylation modification, regions where 6mA methylation is observed are considered to interact with the protein.

[0210] Since the protein interaction regions typically do not contain the entire reference genome sequence, this method can improve the analytical power and precision of methods such as the WGS sequencing setting proposed in DiMeLo-seq. Based on the measurements made during the partial shift, this method can enrich DNA molecules with 6mA modification observed in the first part of the sequence. This provides a greater yield for the protein interaction regions while reducing the yield of off-target regions. This method will still allow the qualitative and quantitative types of analysis described in DiMeLo-seq (and similar) protocols while enhancing their capabilities.

[0211] The training and validation results of the method embodiments are described below. In particular, the results for the following are described: 1) embodiments in which the polymer is classified as belonging to one of a group of classes based only on the modification information 20 (hereinafter referred to as pure adaptive methylation sampling, or pAMS), and 2) embodiments in which the polymer is classified as belonging to one of a group of classes based on both the modification information 20 and the sequence information 30 (hereinafter referred to as enhanced adaptive methylation sampling, or aAMS). The results are based on the training and validation of machine learning classifiers using proprietary and publicly available nanopore sequencing DNA datasets.

[0212] The following datasets were used, as shown in Table 1. The internal dataset is a proprietary dataset, while the external dataset is a publicly available dataset. The internal dataset was generated using the commercially available ligation sequencing kit SQK-LSK110 for DNA library preparation, where sequencing was performed on an Oxford Nanopore Technologies device, specifically using the R9.4.1 flow cell of the MinION or PromethION device.

[0213]

[0214] Table 1

[0215] For each dataset, a random set of 100,000 sequence reads (in fast5 format) was considered. Each read was base-called using bonito v0.5 of the high-accuracy model. The 5mC methylation in the CG motif was inferred using Oxford Nanopore's integrated Remora methylation inference algorithm. Then, each fully base-called DNA sequence of the read was taxonomically classified using the kraken2 sequence taxonomy classification method (https: / / genomebiology.biomedcentral.com / articles / 10.1186 / s13059-019-1891-0) and the PlusPF database (https: / / benlangmead.github.io / aws-indexes / k2) or using the centrifuge sequence taxonomy classification method (paper: https: / / genome.cshlp.org / content / early / 2016 / 11 / 16 / gr.210641.116) and the NCBI database (publicly available here: https: / / benlangmead.github.io / aws-indexes / centrifuge).

[0216] For the bacterial target category dataset, only reads classified as "Bacteria" (taxid:2) were retained. For the non-bacterial target dataset, only reads classified as the same taxonomic id as the organism of the corresponding dataset were retained. A subset of 150,000 bacterial classification reads and 150,000 non-bacterial classification reads were further retained (randomly selected in equal amounts from the corresponding target dataset). Based on the taxonomic classification results, each read was assigned a category label (bacteria / vertebrates / plants). Each read was further assigned an organism label (bacteria / specific organism based on the taxonomic classification results and the organism of the target dataset).

[0217] For each retained read, consider the first N canonical base pairs of the sequence, where N is randomly drawn from a normal distribution with mean 500 and standard deviation 80 (representing the expected length of the portion of the DNA molecule translocated through the nanopore in the first 1 second of sequencing, assuming a translocation speed of approximately 450 bp / s).

[0218] For the first N canonical base pairs of each read, the 5mC methylation inference within the CG motif of the considered sequence was performed. The corresponding methylation profiles were preprocessed for the pAMS and aAMS classifiers as described above. For both pAMS and aAMS classifiers, the preprocessed methylation profiles of the considered 300,000 reads and the associated datasets for each read class and organism label were randomly split into a 90% / 10% split for training / validation purposes. Both pAMS and aAMS classifiers were trained using the recommended program settings, utilizing the aforementioned 90% (training portion) of the preprocessed 300,000 read inputs for the binary classification task of distinguishing bacterial reads from non-bacterial reads based on the input preprocessed methylation profiles.

[0219] Figure 13 , Figure 14 , Figure 15 and Figure 16 The performance of the classifiers in terms of precision, recall, and AUC metrics for a subset of the dataset at both the category label and organism level is shown.

[0220] Figure 13 and Figure 14 Results are shown for a pAMS embodiment in which the polymers are classified as belonging to one of a set of categories based solely on modification information 20 . Figure 13 The results of the category label classification are shown, and Figure 14 Results for the organism-level dataset are shown.

[0221] Figure 15 and Figure 16Shows the results of an aAMS embodiment, where polymers are classified as belonging to one of a group of categories based on modification information 20 and sequence information 30. Figure 15 Shows the results of the class label classification, and Figure 16 Shows the results of the organism-level dataset.

[0222] The pAMS embodiment was further tested to determine its effect on the read lengths of target and background categories and thus its ability to enrich / consume polymers from the target / background categories. This test included "playback simulation". For the playback simulation, the real-time measurement results of a nanopore sequencing experiment were "played back" via the nanopore sequencing control firmware. The playback experiment could be controlled via the pAMS embodiment outlined previously. In this test, a "recording" of a full-length WGS nanopore sequencing experiment was considered, where the input molecules contained approximately 95% yeast community DNA and approximately 5% human HG002 cell line DNA. The input DNA mixture was prepared using the ligation sequencing kit SQK-LSK110 according to the standard ONT ligation library preparation workflow and sequenced on a GridION device using a MinION flow cell (aperture version 9.4.1).

[0223] The main difference between the playback simulation and an actual real-time sequencing experiment is that during the playback simulation, a rejection signal sent to a particular nanopore does not cause the sequenced molecule to pop out of the pore (because another molecule will replace its position along the line). Instead, the firmware saves information about the sequenced fragment of the molecule in question before sending the rejection signal and considers any signals further measured in this pore as coming from a separate molecule.

[0224] The playback simulation allows the effectiveness of the method to be observed. The effectiveness is not measured by the change in the yield of on-target polymer molecules and off-target polymer molecules sequenced. Instead, the effectiveness is measured by the change in the length distribution of on-target molecules and off-target molecules and the number of on-target reads and off-target reads. This is because off-target reads are actually split into a collection of larger and shorter versions of their full length selves.

[0225] Playback simulations were performed with and without the method interacting with the sequencing process. During the first hour of sequencing (a time limit of the first hour of sequencing was chosen for convenience), two replicates of the sequencing playback were performed using different sets of 128 channels (1 to 128, 256 to 384) controlled by the method. In each instance, the reported sequencing results (fast5 files) were further base-called using the guppy 6.1.3 basecalling algorithm with a high-accuracy model and classified using the kraken2 taxonomic classification software to infer bacterial and human reads from the reported data.

[0226] Reads that were base-called and successfully classified were further split into bacterial and human subsets and further processed using the NanoPlot software (https: / / academic.oup.com / bioinformatics / article / 34 / 15 / 2666 / 4934939) to visualize read length distributions and other summary statistics. Figure 17 and Figure 18 The following table provides a comparison of bacterial and human reads in experiments with and without control using this method.

[0227] Figure 17 Tables 2 and 3 show the results for channels 1 to 128. Figure 17 a) shows the bacterial read lengths that consumed polymers classified as bacterial using this method, and Figure 17 b) shows the bacterial read lengths without using this method. The same results as Figure 17 a) and Figure 17 b) are shown in Table 2.

[0228] Channels 1 to 128

[0229] Bacterial reads

[0230]

[0231] Table 2

[0232] Figure 17 c) shows the human read lengths that consumed polymers classified as bacterial using this method, and Figure 17 d) shows the human read lengths without using this method. The same results as Figure 17 c) and Figure 17 d) are shown in Table 3.

[0233] Human reads

[0234]

[0235] Table 3

[0236] Figure 18 Tables 4 and 5 show the results for channels 256 to 384. Figure 18 a) shows the bacterial read lengths that consumed polymers classified as bacterial using this method, and Figure 18 b) shows the bacterial read lengths without using this method. The same results as Figure 18 a) and Figure 18 b) are shown in Table 4.

[0237] Channels 256 to 384

[0238] Bacterial reads

[0239]

[0240] Table 4

[0241] Figure 18 c) Show the human-readable read lengths that consume polymers classified as bacteria using this method, and Figure 18 d) Show the human-readable read lengths without using this method. The same results as Figure 18 c) and Figure 18 d) are shown in Table 5.

[0242] Human reads

[0243]

[0244] Table 5

[0245] In the playback simulation experiment conducted, the average time required for the AMS classifier to preprocess data and perform classification to continue / reject all input molecules / channels was: for 128 nanopore channel data inputs, approximately 0.1 seconds. It was also observed that the AMS calculation performance for 256 nanopore channel inputs was approximately 0.18 seconds, and the AMS calculation performance for 512 nanopore channel inputs was approximately 0.3 seconds.

Claims

1. A method of controlling a biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, wherein the biochemical analysis system comprises at least one sensor element comprising a nanopore, and the biochemical analysis system is operable to obtain successive measurements of the polymer from the sensor element during displacement of the polymer relative to the nanopore of the sensor element, the method comprising: When the polymer has been partially displaced through the nanopore, analyzing the measurements obtained from the polymer during its partial displacement to determine modification information regarding a portion of the sequence of the polymer units, the modification information representing an estimated sequence of the modification state of the subject polymer unit with respect to a modified portion of at least one classical type of polymer unit; Classifying the polymer as belonging to one of a set of categories based on the modification information; And Operating the biochemical analysis system to reject the polymer or continue obtaining measurements from the polymer based on the category to which the polymer unit is classified as belonging.

2. The method according to claim 1, wherein the estimation of the modification state comprises a score regarding the subject polymer unit.

3. The method according to claim 1 or claim 2, further comprising determining sequence information representing an estimate of the identity of the polymer units of the portion of the sequence of the polymer units.

4. The method according to claim 3, wherein the estimate of the identity of the polymer units comprises a score regarding each of a set of types of polymer units.

5. The method according to claim 3 or 4, wherein the estimate of the identity of the polymer units is an estimate regarding a group comprising classical types of polymer units, and determining the modification information comprises analyzing the sequence information to determine the modification information.

6. The method according to claim 3 or 4, wherein the estimate of the identity of the polymer units is an estimate regarding a group comprising classical types of polymer units and at least one classical type of the polymer units in one or more modified forms, wherein the modification information comprises an estimate regarding the at least one classical type of the polymer units in the modified forms.

7. The method according to any one of claims 3 to 6, wherein the polymer is classified as belonging to one of a group of categories based on the modification information and the sequence information.

8. The method according to any one of claims 1 to 6, wherein the polymer is classified as belonging to the group of categories based only on the modification information.

9. The method according to any of the preceding claims, wherein the polymer is derived from an organism, and wherein the categories of the group of categories are taxonomic domains or kingdoms.

10. The method according to claim 9, wherein a first category of the group of categories is a bacterial organism or bacterial organism type, and a second category of the group of categories is a eukaryotic organism or eukaryotic organism type.

11. The method according to any one of the preceding claims, wherein at least one category in the group of categories represents a target sequence, and the step of operating the biochemical analysis system to reject the polymer or continue to obtain measurement results from the polymer includes operating the biochemical analysis system to continue to obtain measurement results when the machine learning classifier classifies the polymer unit as belonging to the at least one category representing the target sequence.

12. The method according to any one of the preceding claims, wherein at least one category in the group of categories represents a background sequence, and the step of operating the biochemical analysis system to reject the polymer or continue to obtain measurement results from the polymer includes operating the biochemical analysis system to reject the polymer when the machine learning classifier classifies the polymer unit as belonging to the at least one category representing the background sequence.

13. The method according to any one of the preceding claims, wherein the polymer is a polynucleotide and the polymer unit is a nucleotide.

14. The method according to any one of the preceding claims, wherein the modification state is a methylation state.

15. The method according to claim 14, wherein: the subject polymer unit is a cytosine nucleotide, and the methylation state is a state of being methylated to at least one of 5-methyl-cytosine or 5-hydroxymethyl-cytosine; and / or the subject polymer unit is an adenosine nucleotide, and the methylation state is a state of being methylated to 6-methyl-adenine.

16. The method according to any one of the preceding claims, wherein the modification state is an oxidation state.

17. The method according to any one of the preceding claims, wherein the subject polymer unit includes a polymer unit that forms part of a predetermined motif of the polymer unit.

18. The method according to claim 17, wherein the predetermined motif of the polymer unit is a cytosine nucleotide followed by a guanine nucleotide in a nucleotide sequence along the 5' → 3' direction.

19. The method according to any one of the preceding claims, wherein the at least one sensor element is operable to eject a polymer that is positively translocating through the nanopore, and operating the biochemical analysis system to reject the polymer includes operating the sensor element to eject the polymer from the nanopore and accept another polymer in the nanopore.

20. The method according to claim 21, wherein the at least one sensor element is operable to eject a polymer that is positively translocating through the nanopore by applying an ejection bias voltage sufficient to eject the polymer, and operating the sensor element to eject the polymer from the nanopore is performed by applying the ejection bias voltage.

21. The method according to any one of the preceding claims, wherein classifying the polymer as belonging to one of the group of categories includes inputting the modification information into a machine learning classifier, and the machine learning classifier classifies the polymer as belonging to one of the group of categories based on the modification information.

22. The method according to claim 21, wherein the machine learning classifier comprises a neural network.

23. The method according to claim 21 or 22, wherein the machine learning classifier is trained using modification information regarding a plurality of classes.

24. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of the preceding claims.

25. A computer storage medium storing the computer program according to claim 24.

26. A biochemical analysis system for analyzing a polymer comprising a sequence of polymer units, the biochemical analysis system comprising at least one sensor element comprising a nanopore, wherein the biochemical analysis system is operable to obtain successive measurements of the polymer from the sensor element during translocation of the polymer relative to the nanopore of the sensor element; wherein the biochemical analysis system is configured to perform the method according to any one of claims 1 to 23.

27. The biochemical analysis system according to claim 24, wherein the biochemical analysis system is a portable biochemical analysis system.

Citation Information

Patent Citations

  • Amphiphilic copolymer planar membranes

    US6723814B2

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Molecular and atomic scale evaluation of biopolymers

    WO2000079257A1

  • Suspended carbon nanotube field effect transistor

    WO2005124888A1

  • High-resolution molecular graphene sensor comprising an aperture in the graphene layer

    WO2009035647A1