Augmented consensus and variant calling

By using two machine learning models to generate and process augmented sequence data with signal features, the method addresses information loss in nanopore-based sequencing, improving accuracy in consensus and variant calling, particularly in homopolymeric regions.

WO2026058009A1PCT designated stage Publication Date: 2026-03-19OXFORD NANOPORE TECH LTD
View PDF 38 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current consensus and variant calling methods using nanopore-based sequencing face limitations due to information loss, particularly in homopolymeric regions, as they rely solely on basecalled read sequences, leading to inaccuracies in determining polymer unit sequences and identifying genetic variants.

Method used

A method involving two machine learning models is employed to generate and process augmented sequence data, incorporating signal features like dwell time and latent space features, which are typically disregarded after basecalling, to enhance accuracy in consensus and variant calling.

Benefits of technology

This approach improves the accuracy of consensus and variant calling, especially in challenging regions like homopolymeric sequences, by leveraging rich internal representations and temporal dynamics of nucleotide translocation, resolving ambiguities and enhancing detection of indels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025052034_19032026_PF_FP_ABST
    Figure GB2025052034_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A method includes receiving a plurality of signals each comprising measurements of a respective polymer sample by a respective nanopore, processing the plurality of signals using a first machine learning model to generate respective sequence data for each 5 signal of the plurality of signals indicating, for each position in a respective sequence: an estimated polymer unit, and one or more signal features associated with a corresponding segment of the signal. The method includes aligning the respective sequence data for the plurality of signals to determine aligned sequence data including, for each position in a common sequence, a plurality of estimated polymer units and a respective one or more signal features associated with each of the plurality of estimated polymer units, and processing the aligned sequence data using a second machine learning model to determine a consensus estimate of a polymer unit at a given position in the common sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AUGMENTED CONSENSUS AND VARIANT CALLING

[0002] Technical Field

[0003] The present invention relates to determining a consensus estimate of a polymer unit of a polymer from measurements of multiple samples of the polymer using a nanopore. The invention has particular relevance to consensus calling and variant calling.

[0004] Background

[0005] Various biochemical analysis systems provide measurements of polymer samples for the purpose of determining the sequence of polymer units, or determining characteristics of individual polymer units or groups of polymer units within such a sequence. One such type of analysis system uses a nanopore. Biochemical analysis systems that use a nanopore have been the subject of much recent development. Typically, successive measurements of a polymer are taken from a sensor element comprising a nanopore during translocation of the polymer through the nanopore. Some property of the system depends on the polymer units in the nanopore, and measurements of that property are taken. This type of measurement system using a nanopore has considerable promise, particularly in the field of sequencing a polynucleotide such as Deoxyribonucleic acid (DNA) or Ribonucleic acid (RNA), sometimes referred to as “basecalling”.

[0006] Measurements by a nanopore are subject to various sources of noise and measurement errors. Although significant progress has been made in the accuracy of basecalling from noisy signals, for example using machine learning techniques, errors typically still occur. In view of this, the process of consensus calling is used to combine multiple read sequences determined by basecalling polymer samples with overlapping sequences, to determine a consensus sequence that is more reliable than any of the individual read sequences. Variant calling, similarly, uses read sequences for multiple overlapping samples to reliably identify variants with respect to a given reference sequence, such as genetic variants with respect to a reference genome.

[0007] Basecalling models typically generate rich internal representations of features of the measurement signals they process. Examples of suitable models include recurrent neural networks and transformer models, each of which generates internal representations reflecting the interdependence between parts of the measurement signal at different time steps.

[0008] Various approaches to consensus calling and variant calling are known, including those based on heuristic methods, Bayesian methods, machine learning methods, and graph-based methods. In all cases, the accuracy of an estimate produced by a consensus calling or variant calling algorithm is limited by the depth (number of polymer samples) at a given sequence position, and the fact that the information available to the algorithm is typically restricted to the individual read sequences, and possibly Q-scores (confidence scores), output by the basecalling models. However, current consensus and variant calling methods that rely solely on basecalled read sequences face significant limitations due to information loss, especially in homopolymeric regions where the basecaller must make a definitive call on the length of a sequence of identical bases, despite inherent uncertainties.

[0009] Summary

[0010] According to aspects of the present disclosure, there are provided a computer- implemented method, a data processing system comprising means for carrying out the method, and a computer program product (such as one or more non-transitory computer-readable storage media) comprising instructions which, when executed by a computer, cause the computer to carry out the method.

[0011] The method includes receiving a plurality of signals each comprising measurements of a respective polymer sample by a respective nanopore, processing the plurality of signals using a first machine learning model to generate respective sequence data for each signal of the plurality of signals indicating, for each position in a respective sequence: an estimated polymer unit, and one or more signal features associated with a segment of the signal corresponding to the estimated polymer unit. The method includes aligning the respective sequence data for the plurality of signals to determine aligned sequence data including, for each position in a common sequence, a plurality of estimated polymer units and a respective one or more signal features associated with each of the plurality of estimated polymer units, and processing the aligned sequence data using a second machine learning model to determine a consensus estimate of a polymer unit at a given position in the common sequence.

[0012] By including signal features in the respective sequence data, the second machine learning model can leverage information that is generated internally by the first machine learning model, but which is not reflected by the estimated polymer units (i.e., read sequences). This can result in a more accurate and reliable consensus estimate, particularly in homopolymeric regions or low complexity regions such as tandem repeats within DNA, where accurate sequencing can be challenging. For example, the second machine learning model may incorporate dwell time and / or latent space features from the first machine learning model into its downstream analysis, allowing consensus and variant calling algorithms to leverage rich internal representations and temporal dynamics of nucleotide translocation - information that is typically disregarded after basecalling and which may capture subtle variations in the raw signal indicative of the precise number of repeated bases, enhancing accuracy in detecting indels and resolving ambiguities in homopolymer lengths.

[0013] Further features and advantages of the invention will become apparent from the following description of preferred embodiments of the invention, given by way of example only, which is made with reference to the accompanying drawings.

[0014] Brief Description of the Drawings

[0015] Fig. 1 shows an example of apparatus for analysing polymer samples.

[0016] Fig. 2 shows a flow diagram representing a method of determine a consensus estimate according to the present disclosure.

[0017] Fig. 3 schematically shows an example of a machine learning architecture for generating augmented sequence data according to the present disclosure.

[0018] Fig. 4 shows schematically an example of a method of processing aligned augmented sequence data according to the present disclosure.

[0019] Fig. 5 illustrates an example of consensus calling using augmented sequence data according to the present disclosure.

[0020] Fig. 6 illustrates a further example of consensus calling using augmented sequence data according to the present disclosure. Fig. 7 illustrates an example of variant calling using augmented sequence data according to the present disclosure.

[0021] Detailed Description

[0022] Details of systems and methods according to examples will become apparent from the following description with reference to the figures. In this description, for the purposes of explanation, numerous specific details of certain examples are set forth. Reference in the specification to ‘an example’ or similar language means that a feature, structure, or characteristic described in connection with the example is included in at least that one example but not necessarily in other examples. It should be further noted that certain examples are described schematically with certain features omitted and / or necessarily simplified for the ease of explanation and understanding of the concepts underlying the examples.

[0023] Embodiments of the present disclosure relate to consensus calling and variant calling algorithms. In particular, embodiments address shortcomings of existing algorithms in which information from measurement signals generated by a nanopore sensor unit is effectively thrown away after basecalling.

[0024] Fig. 1 shows apparatus 100 for determining a consensus estimate of one or more polymer units from a set of polymer samples 102 (of which three polymer samples, 102a, 102b, 102c, are shown in the figure). In this example, the polymer samples 102 are provided in a single chamber or volume, for example within a flow cell, though in other examples different polymer samples may be provided in multiple chambers.

[0025] The polymer samples 102 may for example be samples of polynucleotide (or nucleic acid), and the polymer units may be nucleotides. The nucleic acid is typically deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or other synthetic polymers with nucleotide side chains. The nucleic acid may be single-stranded, double-stranded or comprise both single-stranded and double-stranded regions. The nucleic acid may comprise one strand of RNA hybridised to one strand of DNA. Typically cDNA, RNA, GNA, TNA or LNA are single stranded. The polymer units may be any type of nucleotide. The nucleotide can be naturally occurring or artificial. For instance, the method may be used to verify the sequence of a manufactured oligonucleotide.

[0026] The polymer units may be canonical polymer units. For example, in the case that the polymer sample 102 is a DNA polynucleotide, the canonical bases are adenine (A), cytosine (C), guanine (G), and thymine (T). By contrast ribonucleic acid (RNA) comprises the canonical bases A, C and G, with uracil (U) in place of thymine. A nucleotide may also lack a nucleobase and a sugar. The nucleotide may be modified such as 5mC, 5hmC and N1 -methylpseudouridine. Machine learning methods to detect modified bases are disclosed in WO23094806.

[0027] The polymer may be a nanostructured DNA polymer made by hybridization and origami, such as disclosed by Hannewald et al Angew. Chem. Int. Ed. 2021, 60, 6218- 6229. It may be used to create a nanostructured DNA-based barcode. In other examples, the polymer sample 102 may be a polypeptide such as a protein, or a polysaccharide. The polymer sample 102 may be natural or synthetic. In other examples the polymer sample 102 may be replaced with other types of molecules, including other types of biomolecules.

[0028] The polymer samples 102 may be polynucleotide samples from a common organic source, such as from a common chromosome of a particular individual, or a common part of a common chromosome of a particular individual, or a particular virus, etc. In some examples, additional polymer samples may be present as well, for example corresponding to different chromosomes, different parts of a common chromosome, or even from different individuals (for example where samples are difficult to separate). The polymer samples 102 may correspond to partially or fully overlapping sequences, so that in combination, the polymer samples 102 cover every position within a common sequence multiple times. The number of polymer samples covering a given position of the common sequence may be referred to as the depth at that position.

[0029] The apparatus 100 comprises a nanopore 104 situated in a membrane 106, and electronic circuitry comprising an electrical power supply 108 for driving a potential difference between two electrodes 110a, 110b, and a sensor 112 which in this example is arranged to measure a current through the nanopore. For example, the nanopore may be positioned within a conductive solution (not shown), resulting in an ionic current through the nanopore 104 under the influence of the applied potential difference. The ionic current may be transiently blocked during an event in which the polymer sample 102 translocates or passes through the nanopore 104. As will be explained in more detail hereinafter, the sensor 112 may generate a measurement signal by sampling the ionic current at a sequence of time steps during the event.

[0030] Although only a single nanopore 104 is shown in Fig. 1, the apparatus 100 may employ many nanopores, for example arranged in an array, to provide parallelised collection of information. The measurement system may comprise at least 10 nanopores, at least 100 nanopores, or at least 1,000 nanopores.

[0031] The nanopore 104 is a pore or hole, typically having a size of the order of nanometres, which may allow the passage of polymers therethrough. The nanopore may be a protein pore or a solid-state pore. The nanopore 104 may, for example, be a protein pore such as a polypeptide or a collection of polypeptides. Alternatively, the nanopore may be composed of any other such molecules that allow the nanopore 104 to function as an aperture in the membrane 106. The biological pore may be a transmembrane protein pore. Transmembrane protein pores for use in accordance with the invention can be derived from P-barrel pores or a-helix bundle pores. P-barrel pores comprise a barrel or channel that is formed from P-strands. Suitable P-barrel pores include, but are not limited to, P-toxins, such as a-hemolysin, anthrax toxin and leukocidins, and outer membrane proteins / porins of bacteria, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, lysenin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A and Neisseria autotransporter lipoprotein (NalP). a-helix bundle pores comprise a barrel or channel that is formed from a-helices. Suitable a-helix bundle pores include, but are not limited to, inner membrane proteins and a outer membrane proteins, such as WZA and ClyA toxin. The transmembrane pore may be derived from Msp or from a-hemolysin (a-HL). The transmembrane pore may be derived from lysenin. Suitable pores derived from lysenin are disclosed in WO 2013 / 153359. Suitable pores derived from MspA are disclosed in WO-2012 / 107778. The pore may be derived from CsgG, such as disclosed in WO-2016 / 034591 and W02019 / 002893, both herein incorporated by reference in their entirety. The pore may be a DNA origami pore. A protein pore may be a naturally occurring pore or may be a mutant pore. A protein pore may be inserted into an amphiphilic layer such as a biological membrane, for example a lipid bilayer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer may be a co-block polymer such as disclosed in Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450, WO2014 / 064444, or US6723814 herein incorporated by reference in its entirety. Alternatively, a protein pore may be inserted into an aperture provided in a solid-state layer, for example as disclosed in W02012 / 005857.

[0032] A suitable apparatus for providing an array of nanopores is disclosed in WO 2014 / 064443. The nanopores may be provided across respective wells wherein electrodes are provided in each respective well in electrical connection with an ASIC for measuring current flow through each nanopore. A suitable current measuring apparatus may comprise the current sensing circuit as disclosed in WO-2016 / 181118.

[0033] The nanopore 104 may comprise an aperture formed in a solid-state layer, which may be referred to as a solid-state pore. The aperture may be a well, gap, channel, trench or slit provided in the solid-state layer along or into which the analyte may pass. Methods of forming a pore of controllable size are disclosed in US 9,777,390. Such a solid-state layer is not of biological origin. In other words, a solid-state layer is not derived from or isolated from a biological environment such as an organism or cell, or a synthetically manufactured version of a biologically available structure. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si3N4, A1203, HfO2, SiO2 and multi-layered versions of these materials, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene or other 2D materials such as MoS2. Suitable graphene layers are disclosed in WO-2009 / 035647, WO-2011 / 046706 or WO-2012 / 138357. Suitable methods to prepare an array of solid-state pores are disclosed in WO-2016 / 187519.

[0034] Such a solid-state pore is typically an aperture in a solid-state layer. The aperture may be modified, chemically, or otherwise, to enhance its properties as a nanopore. A solid-state pore may be used in combination with additional components which provide an alternative or additional measurement of the polymer such as tunnelling electrodes (Ivanov AP et al., Nano Lett. 2011 Jan 12; 1 l(l):279-85), or a field effect transistor (FET) device (as disclosed for example in WO-2005 / 124888). Solid state pores may be formed by known processes including for example those described in WO-OO / 79257. The nanopore 104 may be a hybrid of a solid-state pore with a protein pore.

[0035] As mentioned above, the sensor 112 may takes a series or sequence of measurements of a property that depends on the polymer units of a polymer sample 102 being translocated with respect to the nanopore 104. The series of measurements may form a measurement signal. The property that is measured may be associated with an interaction between the polymer sample 102 and the nanopore 104. Such an interaction may occur at a constricted region of the pore. In the example of Fig. 1, the property that is measured is the ion current flowing through a nanopore. These and other electrical properties may be measured using standard single channel recording equipment as describe in Stoddart D et al., Proc Natl Acad Sci, 12;106(19):7702-7, Lieberman KR et al, J Am Chem Soc. 2010; 132(50): 17961-72, and WO-2000 / 28312. Alternatively, measurements of electrical properties may be made using a multi-channel system, for example as described in WO-2009 / 077734, WO-2011 / 067559 or WO- 2014 / 064443. Methods of integrating nanopore sensors within microfluidic channel arrays are disclosed in US 10,718,064.

[0036] The property that is measured by the sensor may not necessarily be ionic current. Some examples of alternative types of property include electrical properties and optical properties. A suitable optical method involving the measurement of fluorescence is disclosed by J. Am. Chem. Soc. 2009, 131 1652-1653. Possible electrical properties include: ionic current, impedance, a tunnelling property, for example tunnelling current (for example as disclosed in Ivanov AP et al., Nano Lett. 2011 Jan 12; 1 l(l):279-85), and a FET (field effect transistor) voltage (for example as disclosed in WO2005 / 124888). One or more optical properties may be used, optionally combined with electrical properties (Soni GV et al., Rev Sci Instrum. 2010 Jan;81(l):014301). The property may be a transmembrane current, such as ion current flow through a nanopore. The ion current may typically be the DC ion current, although in principle an alternative is to use the AC current flow (i.e. the magnitude of the AC current flowing under application of an AC voltage). As mentioned above, one or more conductive or ionic solutions may be provided on either side of the membrane 106 or solid-state layer, which ionic solutions may be present in respective compartments. A sample containing the polymer analyte of interest may be added to one side of the membrane 106 and allowed to move with respect to the nanopore, for example under a potential difference or chemical gradient. The measurement signal may be derived during the movement of the polymer sample 102 with respect to the nanopore 104, for example taken during translocation of the polymer sample 102 through the nanopore 104. The polymer may partially translocate the nanopore 104.

[0037] In order to allow measurements to be taken as the polymer sample 102 passes through the nanopore 104, the rate of translocation can be controlled by a polymer binding moiety. Typically, the moiety can move the polymer sample 102 through the nanopore with or against an applied field. The moiety can be a molecular motor using for example, in the case where the moiety is an enzyme, enzymatic activity, or as a molecular brake. Where the polymer is a polynucleotide there are a number of methods proposed for controlling the rate of translocation including use of polynucleotide binding enzymes. Suitable enzymes for controlling the rate of translocation of polynucleotides include, but are not limited to, polymerases, helicases, exonucleases, single stranded and double stranded binding proteins, and topoisomerases, such as gyrases. The helicase may be any of the helicases, modified helicases or helicase constructs disclosed in WO2013 / 057495, WO 2013 / 098562, WO2013098561, WO 2014 / 013259; WO 2014 / 013262 and WO 2014013260. The helicase may be added to the polynucleotide during sample preparation and stalled by one or more spacers as disclosed in WO2014135838. For other polymer types, moieties that interact with that polymer type can be used. The polymer interacting moiety may be any disclosed in WO-2010 / 086603, WO-2012 / 107778, and Lieberman KR et al, J Am Chem Soc. 2010; 132(50): 17961-72), and for voltage gated schemes (Luan B et al., Phys Rev Lett. 2010;104(23):238103). The rate of translocation of the polymer through the nanopore may be controlled by a voltage control pulse to step the polymer through the nanopore such as disclosed in W02019 / 006214. Translocation of the polymer may be controlled by a molecular hopper such as disclosed by W02020 / 016573.

[0038] The polynucleotide may comprise a polymer leader sequence which preferentially threads into the pore. The leader is preferably negatively charged and may be a polynucleotide, such as DNA or RNA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide. The leader sequence may form part of a Y adaptor typically comprising (a) a double stranded region and (b) a single stranded region or a region that is not complementary at the other end. Leader sequences and Y adaptors suitable for use are disclosed for example in WO2017149316.

[0039] The effective surface concentration at the membrane surface can be enhanced by coupling the polynucleotide to the membrane. Suitable coupling moieties are disclosed for example in WO2017149316 and WO12164270.

[0040] The polymer binding moiety can be used in a number of ways to control the polymer motion. The moiety can move the polymer through the nanopore with or against the applied field. The polynucleotide binding enzyme does not need to display enzymatic activity as long as it is capable of binding the target polynucleotide and controlling its movement through the pore. For instance, the enzyme may be modified to remove its enzymatic activity or may be used under conditions which prevent it from acting as an enzyme. Such conditions are discussed in more detail below. The polynucleotide binding enzyme may be a Dda helicase such as disclosed in WO2015055981, hereby incorporated by reference in its entirety.

[0041] Translocation of the polymer sample 102 through the nanopore may occur either with or against an applied potential, applied by an electronic circuit (not shown) as described hereafter. The translocation may occur under an applied potential which may control the translocation. The binding enzyme is typically held against the cis or trans opening of the nanopore during translocation of the polynucleotide through the nanopore under an applied potential.

[0042] It will be appreciated that the electrical circuitry in this example is illustrative, and may include different or additional components for generating and / or measuring a signal during translocation of a polymer sample 102 with respect to the nanopore 104. For example, the electronic circuit may be substantially as disclosed in WO20 16059427, incorporated herein by reference in its entirety. The electronic circuit may be arranged to control the application of bias voltages across sensor elements of the sensor. During normal operation, the bias voltage is selected to enable translocation of a polymer through the pore of a sensor element. Such a bias voltage may typically be of a level up to -200 mV. The bias voltage supplied by the electronic circuit may also be selected so that it is sufficient to eject the translocating polymer from the pore. By causing the electronic circuit to supply such a bias voltage, the sensor element is operable to eject a polymer that is translocating through the pore. To ensure reliable ejection, the bias voltage is typically a reverse bias, although that is not always essential. The electronic circuit may be connected to electrodes on either side of the membrane 106. The electronic circuit controls the application of bias voltages to generate a bias between the electrodes to control translocation of the polymer sample 102 as described above.

[0043] In some examples, exonucleases that act progressively or processively on double stranded DNA, or similar enzymatic motor structures that act on other types of polymers (for example, an unfoldase acting on a protein molecule, such as disclosed in WO2013123379), can be used on the cis side of the nanopore 104 to feed the remaining single strand through under an applied potential or the trans side under a reverse potential. Likewise, a helicase that unwinds the double stranded DNA can also be used in a similar manner. There are also possibilities for sequencing applications that require strand translocation against an applied potential, but the DNA must be first “caught” by the enzyme under a reverse or no potential. With the potential then switched back following binding the strand will pass cis to trans through the pore and be held in an extended conformation by the current flow. The single strand DNA exonucleases or single strand DNA dependent polymerases can act as molecular motors to pull the recently translocated single strand back through the pore in a controlled stepwise manner, trans to cis, against the applied potential. Alternatively, the single strand DNA dependent polymerases can act as a molecular brake slowing down the movement of a polynucleotide through the pore. Any moieties, techniques or enzymes described in WO-2012 / 107778 or WO-2012 / 033524 could be used to control polymer motion.

[0044] Measurement signals generated by the sensor 112 may be provided to a data processing system 114 for analysis. Variations of the property measured by the sensor 112 may be extremely small. For example, measurements of ionic current may be on the order of picoamps. Measurements of the property may optionally be amplified before being digitised to generate the measurement signal. The data processing system 114 may be a desktop computer, a laptop computer, a tablet computer, a server computer, or any combination thereof. In some examples, the data processing system 114 may be a dedicated device for sequencing polymers, such as a device produced by Oxford Nanopore Technologies (RTM).

[0045] The data processing system 114 includes one or more processors 116 and memory 118. Additionally, the data processing system 114 includes a power supply 120, which may include a mains power supply, a battery, a solar power supply, and / or the like, and the data processing system 114 may also include one or more interface devices 122. The interface devices 122 may include input devices such as a keyboards, a touch screen, a touch pad, a mouse, a microphone, etc. to enable a user to control the data processing device 114, along with output devices such as a display, a speaker, etc. The interface devices 122 may also include network interface devices for example to enable data or information generated by the data processing system 114 to be transmitted to other devices or systems.

[0046] The processor(s) 116 may include one or more of each of a central processing unit (CPUs), graphics processing unit (GPU), neural processing unit (NPU), neural network accelerator (NNA), tensor processing unit (TPU), application-specific integrated circuit (ASIC), application-specific standard product (ASSP), digital signal processor (DSP), field programmable gate array (FPGA), system-on-chip (SoC), any other suitable form of integrated circuit.

[0047] In the present disclosure, the term memory is used to encapsulate both volatile and non-volatile working memory, as well as non-volatile storage. The memory 118 is configured to store measurement signals generated by the sensor device 110, as well as program code for implementing the methods described hereinafter. The program code may include source code, object code, firmware, etc. in any suitable language. For example, the source code may be written in Python, C, C++, Rust, Julia, etc and may use specific development frameworks, libraries, or packages, including PyTorch, TensorFlow, Keras, CUDA, etc.

[0048] The program code may include instructions for carrying out one or more algorithms. In particular, the program code may include code for a basecalling algorithm configured to process measurement signals for individual polymer samples 102 for generate augmented sequence data. The augmented sequence data for a given polymer sample 102 may include a sequence of estimated polymer units for the polymer sample 102, and for each estimated polymer unit, one or more signal features associated with, or derived from, a corresponding segment of the measurement signal. The augmented sequence data may also optionally include confidence scores associated with the estimated polymer units (not shown). The augmented sequence data may be arranged in an augmented pileup format or augmented pileup image format as described in more detail hereinafter.

[0049] The program code may further include code for a consensus calling algorithm and / or a variant calling algorithm, configured to process the augmented signal data for multiple peptide samples 102 to determine consensus estimates of one or more polymer units common to at least some of the polymer samples 102. The basecalling algorithm and the variant / consensus calling algorithm may both be machine learning algorithms with associated machine learning models. For example, the basecalling algorithm may use a neural network model, such as a recurrent neural network (RNN) model or long short-term memory (LSTM) model as disclosed in WO 2023094806 and WO 2020109773, or any other suitable neural network model for processing sequential data, such as a transformer encoder or a full transformer-based neural network. The variant / consensus calling algorithm may also use any neural network model suitable for processing the augmented sequence data, such as an RNN or transformer if the augmented sequence data is provided in an augmented pileup format, or a convolutional neural network (CNN) or vision transformer (ViT) in the case that the augmented sequence data is provided in an augmented pileup image format. Alternatively, the neural network model may be a recurrent-convolutional neural network (R-CNN) if the augmented sequence data is provided in the form of a sequence of images. Other types of machine learning models may be used instead of neural network models, such as for example Gaussian process models or deep Gaussian process models.

[0050] Additional program code may be included that is not shown in Fig. 1, such as code for pre-processing measurement signals and / or for downstream tasks. Preprocessing measurement signals may include normalising and / or filtering the signals, for example low-pass filtering to remove high-frequency noise. The pre-processing may further include performing signal fitting to determining a piecewise constant approximation of the signal, for example using the Cumulative Sum (CUSUM) algorithm, or the Adaptive Time-Series Analysis (ADEPT) algorithm, or any other suitable fitting algorithm. In some examples, the accuracy of basecalling algorithms can be improved by providing piecewise constant approximations of a measurement signal as input, effectively filtering out sources of signal degradation such as measurement noise and finite response time of measurement electronics. In some examples, a signal fitting step may be considered part of the basecalling algorithm.

[0051] A method 200 of performing consensus of variant calling according to the present disclosure will now be described with reference to Fig. 2. The method 200 begins with receiving, at 202, a set of measurement signals each including measurements of a respective polymer sample by a respective nanopore. The polymer samples may have a common biological origin, for example being taken from a common virus or from a common chromosome of an individual. The polymer samples may be partially or completely overlapping. The polymer samples may be within a common physical sample and measured by a single sensor unit comprising one or more nanopores, or may reside in different samples, for example obtained at different times, and may be measured by different sensor units. The measurement signals may optionally be pre-processed, for example by normalisation, filtering, and / or signal fitting, before moving to the next stage. Some measurement signals in the set may be discarded based on one or more quality criteria.

[0052] The method 200 continues with processing, at 204, each of the measurement signals using a first machine learning model to generate augmented sequence data. The augmented sequence data may include a read sequence comprising an estimate or prediction of the identity of each polymer unit in the polymer sample. For each position in the read sequence, the augmented sequence data may further include one or more signal features corresponding to, or derived from, a corresponding segment of the signal

[0053] The first machine learning model may be a basecalling model, such as an RNN model or a transformer model, or any other suitable neural network model with a suitable network head or decoder. A suitable form of network head may be a connectionist temporal classification (CTC) head or a CTC-conditional random field (CTC-CRF) head, or a Hidden Markov Model (HMM) head. Alternatively, the model may include an autoregressive transformer decoder, or any other form of network head or decoder that can estimate or classify polymer units from a sequence of hidden states. The first machine learning model may have been trained using supervised learning with training data comprising a set of training signals including measurements of a target polymer sample, and a target (ground truth) sequence for the polymer sample. The training data may be obtained by measuring a polymer with a known sequence using the sensor element, or by measuring a polymer with an unknown sequence using the sensor element and using an alternative sequencing method to determine the target sequence.

[0054] During processing of a measurement signal, the first machine learning model may temporarily store one or more embeddings, internal states, hidden states, or latent states representing features of the measurement signal at different signal positions. For example, a neural network model may initially process measurement signals using a one or more convolutional layers to generate a sequence of feature vectors corresponding to overlapping segments of the signal. The feature vectors may be processed by one or more further layers such as RNN layers or transformer encoder layers (also known as blocks), each of which may generate a new sequence of hidden states. Each given segment of the signal may therefore be associated with a set of hidden states. One or more of these hidden states may be described as features in an abstract latent space of the first machine learning model. For example, a feature of the latent space may be given by, or derived from, the hidden state generated by the final layer of the network, as this may be the hidden state deemed most relevant to estimating the polymer unit. Alternatively, hidden states from multiple layers may be concatenated or otherwise combined to determine a corresponding feature in the latent space.

[0055] It will be appreciated that multiple hidden states may correspond to measurements of a single polymer unit, and that the number of hidden states associated with different polymer units may vary, depending on how long the polymer unit took to translocate the nanopore. In other words, multiple hidden states may correspond to a given sequence position in the determined read sequence. The network head or decoder may be arranged to estimate points in the signal corresponding to transitions between polymer units. For example, for a network head that uses the CTC topology, an additional blank symbol may be included as a candidate classification for a given signal position, and the head may be arranged to determine an intermediate symbol for each hidden state of the final hidden network layer, then collapse the resulting sequence of intermediate symbols by merging repeated symbols and then removing instances of the blank symbol. With this algorithm, it is possible to estimate which positions in the signal correspond to a given output symbol (i.e. polymer unit prediction), and also the “dwell time” taken for the polymer to pass through the nanopore. It will be appreciated other network heads or decoders can also enable multiple hidden states to be associated with a given sequence positions and for dwell time to be estimated. For example, using an HMM model head may enable transitions between polymer units to be identified based on evaluations of a likelihood function. Multiple latent space features corresponding to a given sequence position may be combined, for example by averaging, pooling, concatenating, or any other suitable mathematical operation.

[0056] Fig. 3 shows an example in which a measurement signal 302 is pre-processed using a signal fitting algorithm 304 before being processed by a basecalling model 306. The basecalling model 306 in this instance includes multiple hidden layers 308, of which three hidden layers 308a, 308b, 308c are shown, and an output layer 310. The hidden layers 308 may for example include a convolutional layer followed by a series of RNN layers or transformer layers, and the output layer may be a network head such as a CTC head, a CTC-CRF head, or an HMM network head. Each of the hidden layers 308 may output a sequence of hidden states corresponding to respective (possibly overlapping) segments of the signal 302. The output layer 310 in this example is arranged to determine, for each position in a sequence, an estimated polymer unit 312, a Q score (e.g. Phred quality score) or confidence score 314 for the estimate, and a dwell time 316 associated with the corresponding segment of the signal.

[0057] In the example of Fig. 3, hidden states are extracted from each of the hidden layers 308 and provided to an embedding model 318 (for example as separate inputs or after concatenation). The embedding model 318 processes the hidden states corresponding each position in the sequence to generate a latent space feature 320, which may be a relatively low-dimensional embedding of the hidden states. In some examples, the embedding model 318 may be a further machine learning model such as a fully connected neural network or any other suitable neural network. In other examples, the embedding model 318 may instead apply a manually coded mathematical function. In further examples still, the embedding model 318 may be omitted and one or more hidden states may be used directly as the latent space feature 320.

[0058] In this example, the estimated polymer units 312, Q scores 314, dwell times 316 and latent space features 320 form augmented sequence data 322 for the measurement signal 302. The dwell times 316 and the latent space features 320 may be described as signal features. In other examples, other signal features may also be included in addition to, or instead of, the dwell times 316 and / or the latent space features 320. For example, measurable characteristics of the signal such as mean signal value, variance, Fourier domain features, and the like. Furthermore, in some examples, the Q score 314 may be omitted from the augmented signal data 322.

[0059] The method 200 continues with aligning, at 206, the augmented sequence data from the different measurement signals to determine aligned sequence data. Aligning the sequence data may involve compiling data for the individual read sequences in a FASTQ file, and processing compiled data using appropriate software to generate a sequence alignment map (SAM), a binary alignment map (B M), and / or a Compressed Reference-oriented Alignment Map (CRAM). The aligning may further include generating augmented pileup data and / or an augmented pileup image representing the aligned sequence data. Augmented pileup data may for example add one or more further columns to the standard five or six columns in the pileup data format, where the further columns representing dwell time, latent space features (such as components of a low-dimensional feature space embedding), and / or any other signal feature present in the augmented signal data. An augmented pileup image may include additional channels with visual representations of the signal features, and / or may modify the appearance of existing channels (for example, adjusting pixel values in dependence on dwell time).

[0060] In the case of variant calling, the process of aligning the sequence data (i.e. read sequences) may include read alignment or read mapping in which the individual read sequences are mapped to a common reference sequence, such as a reference genome. Such mapping may be performed for example using the open-source minimap2 software licensed by the Dana-Farber Cancer Institute and Broad Institute, Inc.

[0061] In the case of consensus calling, or optionally in the case of variant calling, aligning sequence data may include assembly methods in which a draft genome is assembled without the need for a reference sequence. In another example, many-to- many read-to-read mapping (such as Multiple Sequence Alignment) may be performed, which uses neither a reference sequence nor a draft assembly. Suitable software for aligning sequence data includes the Telomere-to-telomere (T2T) software developed by Oxford Nanopore Technologies. In other examples, read mapping or assembly can be performed during the basecalling step 204, for example using the MinKNOW and Dorado software created by Oxford Nanopore Technologies.

[0062] In some examples, signal features from the augmented signal data may be used to improve the accuracy of read mapping or alignment. In that regard, read mapping or alignment algorithms may aim to optimize an objective that penalizes differences between symbols (e.g. for individual polymer units or k-mers) for different candidate mappings (for example by calculating cross-correlations). By including one or more additional terms in the objective to penalizes deviations between signal features, the additional information may be provided to the objective, which may for example enable resolution of ambiguities within multimodal objectives. For example, the additional term(s) may be dependent on an (absolute) difference between dwell times and / or a metric distance between latent space features.

[0063] In any event, the aligned sequence data determined at step 206 includes, for a given position in a common sequence (e.g. corresponding to a reference sequence or a draft sequence), multiple estimated polymer units (with the number of estimated polymer units being the depth at the given position), and one or more signal features associated with each of the estimated polymer units. In some examples, the alignment process may determine a mapping quality score for each estimated polymer unit indicative of how confident the alignment algorithm is in mapping the estimated polymer unit to the corresponding position in the common sequence. In such cases, these mapping quality scores may additionally be included in the aligned sequence data.

[0064] In some examples, local realignment techniques may be used to correct alignment errors from early in the alignment process and further improve alignment accuracy. Furthermore, the alignment process may result in different reads being grouped and aligned to different sequences, for example different draft genes or different reference genes within a common genome. The method 200 may then continue independently for the individual sequences. The method 200 continues with processing, at 208, the aligned sequence data using a second machine learning model to determine a consensus estimate of a polymer units at a given position in the common sequence. In the case of consensus calling, the second machine learning model may generate a consensus sequence comprising a consensus estimate for each position in the common sequence, whereas in the case of variant calling the second machine learning model may generate a consensus estimate of one or more polymer units corresponding to a variant, in which case the consensus estimate may indicate a homozygous variant that is present on both alleles of a given gene, or a heterozygous reference variant that is present on only one allele. In the latter case, the variant may be present at a given position in approximately half of the samples. A third case that may be identified is heterozygous variants on both alleles (i.e. both alleles differ from the reference and from one another).

[0065] In some examples, it may be possible to generate more than one consensus sequence for a given common sequence, for example in the case of a diploid genome, in which a consensus sequence may be estimated for a region of each allele that is covered by a single (haploid) common sequence. In other examples, the second machine learning model may also be trained to estimate correlations between variants at different positions, for example a set of variants that collectively belong to one allele or the other.

[0066] The second machine learning model may be arranged to process the aligned sequence data in any suitable format. For example, if the aligned sequence data is provided in the format of augmented pileup data or a sequential representation of the augmented pileup data, then the second machine learning model may have an architecture suitable for processing sequential data, such as an RNN, LSTM, or transformer architecture In the aligned sequence data is in the format of an augmented pileup image, then the second machine learning model may have an architecture suitable for processing multi-channelled image data, such as a CNN or ViT. In other examples, the augmented sequence data may be in the format of a sequence of images, in which case an R-CNN or ViT architecture may be used. In the case of variant calling, a reference sequence may be provided to the second machine learning model along with the aligned sequence data. Fig. 4 shows an example in which aligned sequence data 402 is provided in the format of an augmented pileup image with multiple channels. A first channel 404a of the image depicts an array comprising a symbolic or other representation of the aligned read sequences, with the position number in the common sequence increasing in one direction (e.g. the horizontal direction) and the sequence number increasing another direction (e.g. the vertical direction). In further examples, the first channel is actually multiple channels, such as three channels, corresponding to a colour image representation of the read sequence. In further examples still, different channels may be associated with each candidate polymer unit symbol. Different pileup image formats are used by different commercially available variant calling algorithms, such as DeepVariant developed by Google (RTM) and Clair developed by the University of Hong Kong, Department of Computer Science, and the augmented pileup image may be an extension of any of these formats. In the case of variant calling, the pileup image may include an additional row representing the reference sequence.

[0067] A second channel 404b of the augmented pileup image may represent a quality score such as an Phred quality score determined by the first machine learning model, or a mapping quality score determined by the alignment algorithm. In some cases, separate channels may be included for the two different scores. The second channel 404b may for example include pixels, or regions of pixels, with pixel values corresponding to the quality score.

[0068] A third channel 404c of the augmented pileup image may represent an estimated dwell time for each position in each read sequence. The third channel 404b may for example include pixels, or regions of pixels, with pixel values corresponding to the dwell time.

[0069] A fourth channel 404d of the augmented pileup image 402 may include a representation of a latent space feature. The fourth channel may for example include symbolic representations of the latent space feature (such as in the form of arrows or other two-dimensional representations, or the fourth channel 404d may include multiple sub-channels corresponding to respective dimensions of the latent space.

[0070] It will be appreciated that not all of the channels of the aligned sequence data 402 may be used in every example, and in some cases other channels may be included in addition to, or as alternatives to, the channels shown, for example representing other signal features.

[0071] The aligned sequence data 402 is processed using a variant calling or consensus calling model 406, to generate consensus data 408. In the present example, the model 406 is a CNN model comprising one or more CNN layers each having a learned kernel or filter. The CNN may be a fully convolutional network or may include other layers such as fully connected layers. The CNN may for example have an encoder-decoder architecture which downscales the augmented pileup image to a latent representation then upscales the latent representation to a sequential output. A suitable network head such as a classification head or a CRF head may be included as a final layer to generate the consensus data 408.

[0072] The consensus data 408 may include consensus estimates of polymer units at one or more positions in the common sequence. In the case of consensus calling, the consensus data 408 may be a consensus sequence, which may be longer than any of the read sequences included in the aligned sequence data. The consensus sequence may for example correspond to an entire genome and may constitute an estimated genotype for an individual. In the case of variant calling, the consensus data 408 may indicate positions at which variants have been identified relative to the reference sequence, along with an indication of the type of variant.

[0073] The second machine learning model may be trained using supervise learning with training data comprising aligned sequence data for known sequences, such as known genotypes, with known variants present in the case of variant calling. In some examples, the second machine learning model may be trained jointly with one or more other models in the consensus calling or variant calling pipeline. For example, the second (consensus / variant calling) machine learning model may be trained or finetuned jointly with the embedding model 318 of Fig. 3, so that the embedding model learns to generate embeddings that provide relevant information to the second machine learning model. Additionally, or alternatively, the second machine learning model may be trained jointly with the first (basecalling) machine learning model 306, encouraging the basecalling model to generate internal representations that are relevant to consensus calling or variant calling. In this way, the entire machine learning pipeline may be trained in an end-to-end manner (optionally after pre-training the individual models individually) by providing a consensus / variant calling objective that backpropagates through each of the machine learning models during training.

[0074] In the case of variant calling, the method 200 may continue with identifying, at 210, one or more variants with respect to the reference sequence, and optionally identifying the type of variant. The variant may for example be a single-nucleotide variant (SNV), a short insertion / deletion (InDei), a copy-number variant (CNV), or a structural variant (SV). Identifying the variant(s) may include distinguishing true variants from false positives (such as caused by measurement errors or anomalous samples) based on outputs (such as likelihood scores) of the second machine learning model. The identified variant(s) may include generating a file in a suitable format such as variant call format (VCF), genomic variant call format (gVCF) or binary variant call format (BCF).

[0075] As explained above, the methods disclosed herein may enable signal features to be used to increase the accuracy of variant calling or consensus calling . Figs. 5-7 show illustrative examples of signal features resulting in improved accuracy of consensus calling or variant calling.

[0076] Fig. 5 shows an example in which segments of augmented sequence data 502a- d are aligned and processed using a consensus calling model to determine a consensus sequence 504. In this example, each segment of augmented sequence data 502a-d is aligned to a common sequence, and includes a respective sequence of estimated bases (polymer units), Q scores, and signal embeddings (latent space features), represented here as two-dimensional arrows. It is observed that three different bases, A, G and C are estimated at the common sequence position 506a. The Q scores for the A and C estimates are higher than the Q score for the G estimate, which may enable the consensus calling model to determine that the G estimate is more likely to be a measurement error. However, the Q scores for the A and C estimates are identical, so the consensus calling model may be unable to resolve which is the more likely base. However, it is observed that the signal embeddings for all three estimates are much closer to the embedding that is typical for A (up arrow) than they are to the embedding that is typical for C (down arrow), which may cause the model to determine that A is a more likely candidate than C. A similar example is shown at position 506b, where the embeddings all seem to indicate A, enabling the model to be more confident that the T estimate is erroneous.

[0077] Fig. 6 shows a further example in which segments of augmented sequence data 602a-d are aligned and processed using a consensus calling model to determine a consensus sequence 604. In this example, each segment of augmented sequence data includes a respective sequence of estimated bases and dwell times. It is observed that three different bases, A, G and C are estimated at the common sequence position 606. The dwell times for all three estimates are identical, and equal to the dwell time typically observed for A (8 units), which may cause the model to determine that A is the most likely candidate.

[0078] Fig. 7 shows an example in which segments of augmented sequence data 702a- d are aligned and processed using a variant calling model to identify variants with respect to a reference sequence 704. In this example, each segment of augmented sequence data includes a respective sequence of estimated bases, Q scores, dwell times, and signal embeddings. At the sequence position 704, there are three estimates of A, and one estimate of T. The model has the task of determining whether the T is a true variant (such as a heterozygous variant) or an error or anomaly. It is observed that the dwell time for the A estimate in 702c is equal to the dwell time typically observed for T (8 units), while the embedding for this estimate is more similar to the embedding typically seen for T (left arrow) than to the embedding typically seen for A (up arrow). This information may cause the model to determine that the A estimate in 702c is erroneous and should be a T, consistent with the T being a heterozygous variant at the position 704.

[0079] It will be appreciated that the examples of Figs. 5-7 are highly idealised with the aim of illustrating simple cases where uncertainty can be resolved using signal features within augmented sequence data. In reality, signal features may be more complex, and the variant or consensus calling model may learn to determine rich, relevant information based on signal features and their interdependence within a given sequence.

[0080] The above embodiments are to be understood as illustrative examples of the invention. Further embodiments of the invention are envisaged. For example, additional algorithms or models may be provided to determine signal features, independently of the basecalling model, for use by the variant / consensus calling model. Furthermore, the method could be used to determine an extent to which there is correspondence between polymer samples from multiple sources. It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

CLAIMS1. A computer-implemented method comprising: receiving a plurality of signals each comprising measurements of a respective polymer sample by a respective nanopore; processing the plurality of signals using a first machine learning model to generate, for each signal of the plurality of signals, respective sequence data indicating, for each position in a respective sequence: an estimated polymer unit; and one or more signal features associated with a segment of the signal corresponding to the estimated polymer unit; and aligning the respective sequence data for the plurality of signals to determine aligned sequence data comprising, for each position in a common sequence, a plurality of estimated polymer units and a respective one or more signal features associated with each of the plurality of estimated polymer units; and processing, using a second machine learning model, the aligned sequence data to determine a consensus estimate of a polymer unit at a given position in the common sequence.

2. The computer-implemented method of claim 1, wherein the one or more signal features comprise an estimated dwell time of the estimated polymer unit within respective nanopore.

3. The computer-implemented method of any preceding claim, wherein the one or more signal features comprise a feature extracted from a latent space of the first machine learning model.

4. The computer-implemented method of claim 3, wherein the first machine learning model comprises a neural network and the latent space corresponds to an output of one or more hidden layers of the neural network.

5. The computer-implemented method of claim 4 or claim 5, comprising extracting the feature from the latent space by processing data generated by the first machine learning model using an embedding model.

6. The computer-implemented method of any preceding claim, wherein the respective sequence data further indicates, for each position in the respective sequence, a quality score associated with the estimated polymer unit.

7. The computer-implemented method of any preceding claim, wherein the aligned sequence data comprises a pileup image with one or more channels corresponding to each of the one or more signal features.

8. The computer-implemented method of claim 7, wherein the second machine learning model comprises a convolutional layer for processing the aligned sequence data.

9. The computer-implemented method of any preceding claim, wherein the aligning uses at least one of the one or more signal features.

10. The computer-implemented method of any preceding claim, wherein the processing using the second machine learning model determines a consensus sequence of polymer units.

11. The computer-implemented method of any of claims 1 to 9, comprising determining that the consensus estimate of the polymer unit is a variant with respect to a reference sequence.

12. The computer-implemented method of claim 11, comprising determining, based on a variation of the estimated polymer units at the given position in the common sequence, that the variant is a heterozygous variant.

13. The computer-implemented method of any preceding claim, wherein therespective polymer samples are respective polynucleotide samples, and the estimated polymer units are estimated nucleotides.

14. The computer-implemented method of any preceding claim, wherein each signal of the plurality of signals is indicative of ion flow through the respective nanopore during translocation of the respective polymer sample with respect to the respective nanopore.

15. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any preceding claim.

16. A data processing system comprising means for carrying out the method of any of claims 1 to 14.

17. Apparatus comprising: a sensor element comprising a nanopore; means for translocating a polymer with respect to the nanopore; and a data processing system comprising means for carrying out the method of any of claims 1 to 14 to analyse a signal comprising measurements of a polymer by the sensor element during translocation of the polymer with respect to the nanopore.

18. A method compri sing : translocating a polymer with respect to a nanopore; generating a signal comprising measurements of the polymer by a sensor element comprising the nanopore; and analysing the signal using the computer-implemented method of any of claims1 to 14.

Citation Information

Patent Citations

  • Integrating nanopore sensors within microfluidic channel arrays using controlled breakdown

    US10718064B2

  • Amphiphilic copolymer planar membranes

    US6723814B2

  • Method for controlling the size of solid-state nanopores

    US9777390B2

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Molecular and atomic scale evaluation of biopolymers

    WO2000079257A1