Peptide characterisation
By using a composite polymer with a polynucleotide connected to a peptide to segment the signal and train a classification model with polynucleotide barcodes, the method addresses the challenges of peptide characterization in nanopore sensors, achieving accurate and efficient peptide identification.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-03-19
AI Technical Summary
Existing methods for characterizing peptides using nanopore sensors face challenges due to the lack of controlled mechanisms for translocating peptides and the difficulty in obtaining accurate measurements of relatively short peptide lengths.
A method involving a composite polymer comprising a polynucleotide connected to a peptide, where the region connecting the polynucleotide to the peptide is identified to segment the signal accurately, enabling a classification algorithm to characterize the peptide, and a classification model is trained using polynucleotide barcodes as ground truth labels.
Facilitates accurate peptide characterization by segmenting the signal effectively and enables efficient training of the classification model, allowing for low-cost and precise peptide identification in mixed samples.
Smart Images

Figure GB2025052024_19032026_PF_FP_ABST
Abstract
Description
[0001] PEPTIDE CHARACTERISATION
[0002] Technical Field
[0003] The present invention relates to techniques for characterising peptides. The disclosure has particular, but not exclusive, relevance to training a classification model for characterising peptides.
[0004] Background
[0005] Many tasks in medicine, biochemistry and other fields include characterising molecules within a sample. Classifying or identifying biomolecules associated with an individual can help analyse and determine the individual’s biological function. As an example, a patient’s sample may contain an enzyme specific to the patient. By identifying the enzyme, an enzyme assay may be conducted for measuring the enzymatic activity. In other examples, there may be specific proteins present in an individual’s sample that may be characterised.
[0006] Peptides are chains of amino acids that constitute a protein. In order to characterise a peptide, interactions between the peptide and a nanopore sensor may be measured during translocation of the peptide through the nanopore. Such nanopore sensors are commercially available, such as the MinlON™ device sold by Oxford Nanopore Technologies Ltd, comprising an array of nanopores integrated with an electronic chip that may be used to obtain measurements. However, mechanisms for translocating an individual peptide through a nanopore in a controlled manner may not be available. Furthermore, obtaining accurate measurements of a peptide or peptide fragment in this manner can be challenging due to their relatively short lengths.
[0007] Summary
[0008] According to aspects of the present disclosure, there is provided a computer- implemented method, apparatus for carrying out the method, and a computer program product (such as one or more non-transitory storage media) comprising machine readable instructions which, when executed by a computer, cause the computer to carry out the method. The method includes receiving a signal indicative of measurements of a composite polymer by a sensor unit comprising a nanopore, the composite polymer comprising a polynucleotide connected to a peptide. The method further includes determining, by identifying a region of the composite polymer connecting the polynucleotide to the peptide, a segment of the signal comprising measurements of the peptide. The method also includes processing the determined segment of the signal using a classification algorithm to characterise the peptide. In some examples, the identified region may include measurements of the polynucleotide and / or measurements of one or more amino acids associated with the peptide.
[0009] By including the polynucleotide in the composite polymer, measurements of the peptide may be obtained by translocating the composite polymer with respect to the nanopore. By identifying the region of the composite polymer connecting the polynucleotide to the peptide, the signal may be accurately segmented to enable the classification algorithm to be focused on the peptide region, which may for example be considerably shorter than the polynucleotide.
[0010] The above method may be carried out in order to train a classification model used by the classification algorithm. In such cases, characterising the peptide may include estimating a characteristic of the peptide, and updating parameter values of the classification model based on a comparison between the estimated characteristic of the peptide and a ground truth characteristic of the peptide. The method may include determining the ground truth characteristic of the peptide being based on the identified polynucleotide, and determining the ground truth characteristic of the peptide being based on the identified polynucleotide. In some examples, the polynucleotide connected to the peptide may function as a polynucleotide barcode containing a known sequence of polynucleotides. In that regard, a sample may be prepared containing multiple composite polymers including different known peptides associated with respective different polynucleotide barcodes to serve as ground truth labels for the associated peptides. During training, a given polynucleotide barcode may be identified within a given signal, and used to determine the ground truth characteristic of the associated peptide. Parameter values of the classification model may then be updated based on a comparison between the estimated characteristic of the peptide and the ground truth characteristic of the peptide. In this manner, training of the classification model may be facilitated with a mixed sample of peptides, for example within a single flow cell, resulting in a relatively low cost and efficient training process.
[0011] Further features and advantages of the invention will become apparent from the following description of preferred embodiments of the invention, given by way of example only, which is made with reference to the accompanying drawings.
[0012] Brief Description of the Drawings
[0013] Figure 1 A shows a schematic diagram of apparatus for characterising a peptide in a composite polymer.
[0014] Figure IB shows an example measurement signal obtained by translocating a composite polymer through a nanopore.
[0015] Figure 2 shows a schematic diagram of a method for characterising a peptide in a composite polymer.
[0016] Figure 3 shows a schematic diagram of a further method for characterising a peptide in a composite polymer.
[0017] Figure 4 shows a schematic diagram of a method for segmenting a signal based on a hidden Markov model (HMM).
[0018] Figure 5 shows example results of evaluating the likelihood functions and the transition matrix for a signal segment.
[0019] Figure 6 shows a schematic diagram of a method of polypeptide identification based on an unsupervised machine learning model.
[0020] Figure 7 shows a flow diagram representing a method of characterising a peptide in a composite polymer.
[0021] Detailed Description
[0022] Details of systems and methods according to examples will become apparent from the following description with reference to the figures. In this description, for the purposes of explanation, numerous specific details of certain examples are set forth. Reference in the specification to ‘an example’ or similar language means that a feature, structure, or characteristic described in connection with the example is included in at least that one example but not necessarily in other examples. It should be further noted that certain examples are described schematically with certain features omitted and / or necessarily simplified for the ease of explanation and understanding of the concepts underlying the examples.
[0023] Fig. 1A shows a schematic of apparatus 100 for characterising a peptide. The system 100 comprises a nanopore 102 situated in a membrane 104; in this example, the nanopore 102 is a transmembrane nanopore. The nanopore 102 may function as an aperture in the membrane 104. For example, the nanopore may be a protein pore. The membrane 104 may be flanked by ionic solutions and the membrane material may preferably have a high resistance or resistivity so that the material volume is inconducive to the flow of ions. As an example, the membrane 104 may be constructed out of a lipid bilayer. In examples, the composite polymer 106 comprises a polynucleotide 108 such as DNA or RNA, or a section of a polynucleotide such as DNA / RNA. Such a composite polymer 106 may be moved, i.e. “translocated” through the nanopore 102. Translocation of the composite polymer through the nanopore may be carried out using a method such as described in WO2021 / 111125, WO2021 / 133168 or WO2023 / 118891.
[0024] The composite polymer 106 further comprises a peptide 110, which is connected to the polynucleotide 108. The peptide 110 may include a short sequence of amino acids such as, for example, a fragment of a protein. In some examples, the peptide 110 may comprise anywhere between 1 and 50 amino acids; in other examples, fewer or more amino acids may be included. In some examples, the peptide 110 may be connected directly to an end of the polynucleotide 108. Alternatively, the peptide 110 may be connected to the polynucleotide 108 by a bridging polymer 112. In the example shown in Fig. 1, the composite polymer 106 comprises a bridging polymer 112 that may itself be a polynucleotide comprising a sequence of nucleotide units. As an example, the bridging polymer 112 may be a polythymine or polythymidine molecule comprising a sequence of thymine DNA bases. In some examples, a polythymine or polythymidine with anywhere between 5 and 100 bases, some thymine bases, may be included; in other examples, the bridging polymer 112 may include fewer or more thymine DNA bases. In other examples, the bridging (first) polymer may be a polyadenosine or a polycytosine. The bridging polymer 112 may also include other nucleotide bases. For example, the bridging polymer 112 may comprise a sequence of adenine or guanine or cytosine or uracil DNA bases. The bridging polymer 112 may also comprise a sequence of a combination of two or more DNA bases. In general, one or more components of the composite polymer 106 may not be naturally occurring; for example, the polynucleotide 108, the first polymer 112 and the second polymer 114 may comprise artificially synthesised bases or molecules. The bridging polymer may comprise nonnucleotides. Generally speaking the bridging polymer may be chosen from one that provides a signal which is distinguished from the polynucleotide signal.
[0025] The composite polymer 106 may be translocated through the nanopore 102 to perform measurements indicative of interactions between the composite polymer 106 and the nanopore 102. In some examples, the composite polymer 106 may include a DNA sequence with the 5’ end being free and the 3’ end being connected to peptide 110. In the example shown in Fig. 1 A, the 3’ end is connected to an end of the bridging polymer 112. Further, the peptide 110 may be connected to the bridging polymer 112 at one of its terminii, for example, either the C-terminus or the N-terminus. The polynucleotide 108 may therefore be read from the 5’ end, and the peptide 110 may be read from the C-terminus when the composite polymer 106 is translocated through the nanopore 102. In examples, the polynucleotide 108 may comprise a motor enzyme to aid translocation of the composite polymer 106 through the nanopore 102. The motor enzyme may be a polynucleotide binding protein such as a helicase. In this manner, a signal comprising measurements of the composite polymer 106 (i.e. including measurements of the polynucleotide 108, and the peptide 110, and / or the bridging polymer 112) may be obtained.
[0026] As explained in more detail herein, the composite polymer comprises a polynucleotide connected, i.e. “conjugated”, to a target peptide. The target peptide can be conjugated to the polynucleotide at any suitable position. For example, the peptide can be conjugated to the polynucleotide at the N-terminus or the C-terminus of the peptide. The peptide can be conjugated to the polynucleotide via a side chain group of a residue (e.g. an amino acid residue) in the peptide.
[0027] In some embodiments the target peptide has a naturally occurring reactive functional group which can be used to facilitate conjugation to the polynucleotide. For example, a cysteine residue can be used to form a disulphide bond to the polynucleotide or to a modified group thereon. In some embodiments the target peptide is modified in order to facilitate its conjugation to the polynucleotide. For example, in some embodiments the peptide is modified by attaching a moiety comprising a reactive functional group for attaching to the polynucleotide. For example, in some embodiments the peptide can be extended at the N-terminus or the C-terminus by one or more residues (e.g. amino acid residues) comprising one or more reactive functional groups for reacting with a corresponding reactive functional group on the polynucleotide. For example, in some embodiments the peptide can be extended at the N-terminus and / or the C-terminus by one or more cysteine residues. Such residues can be used for attachment to the polynucleotide portion of the conjugate, e.g. by maleimide chemistry (e.g. by reaction of cysteine with an azido-maleimide compound such as azido-[Pol]-maleimide wherein [Pol] is typically a short chain polymer such as PEG, e.g. PEG2, PEG3, or PEG4; followed by coupling to appropriately functionalised polynucleotide e.g. polynucleotide carrying a BCN group for reaction with the azide). For avoidance of doubt, when the peptide comprises an appropriate naturally occurring residue at the N- and / or C-terminus (e.g. a naturally occurring cysteine residue at the N- and / or C-terminus) then such residue(s) can be used for attachment to the polynucleotide.
[0028] In some embodiments a residue in the target peptide is modified to facilitate attachment of the target peptide to the polynucleotide. In some embodiments a residue (e.g. an amino acid residue) in the peptide is chemically modified for attachment to the polynucleotide. In some embodiments a residue (e.g. an amino acid residue) in the peptide is enzymatically modified for attachment to the polynucleotide.
[0029] The conjugation chemistry between the polynucleotide and the peptide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react with amines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like.
[0030] Other suitable chemistry for conjugating the peptide to the polynucleotide includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0031] (a) copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);
[0032] (b) strain-promoted azide-alkyne cycloadditions; including alkene and azide [3+2] cycloadditions; alkene and tetrazine inverse-demand Diels-Alder reactions; and alkene and tetrazole photoclick reactions;
[0033] (c) copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);
[0034] (d) the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and
[0035] (e) the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond. Any reactive group may be used to form the conjugate. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis- maleimidotriethyleneglycol; 3,3’-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4’- diisothiocyanatostilbene-2, 2’ -disulfonic acid disodium salt; Bis[2-(4- azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acid N- hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N- hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO2010 / 086602, particularly in Table 3 of that application.
[0036] In some embodiments the reactive functional group is comprised in the polynucleotide and the target functional group is comprised in the peptide prior to the conjugation step. In other embodiments the reactive functional group is comprised in the peptide and the target functional group is comprised in the polynucleotide prior to the conjugation step. In some embodiments the reactive functional group is attached directly to the peptide. In some embodiments the reactive functional group is attached to the peptide via a spacer. Any suitable spacer can be used. Suitable spacers include for example alkyl diamines such as ethyl diamine, etc.
[0037] In some embodiments the polynucleotide is conjugated directly to the peptide. In some embodiments the polynucleotide is ligated to a polynucleotide linker which is conjugated to the peptide. The use of a linker may be beneficial to facilitate the facile conjugation of the polynucleotide to a target peptide. In the example shown in Fig. 1, the composite polymer 106 includes a second polymer 114 connected to the peptide 110. For example, the bridging polymer 112 may be a first polymer 112 connected at the C-terminusof the peptide 110, and the second polymer 114 may be connected to the peptide 110 at its N-terminus. The second polymer 114 may also be a polythymine or polythymidine associated with, for example, anywhere between 5 and 100 bases, some thymine bases. In other examples, the second polymer may be a polyadenosine or a poly cytosine. Due to the relative placement of the polymers with respect to the peptide 110 and the nanopore 102, the first and second polymers may be referred to as the left polymer (LPOLY) and the right polymer (RPOLY), respectively.
[0038] In the example shown in Fig. 1, the composite polymer 106 includes a second polynucleotide 116 connected to the second polymer 114. For example, the second polynucleotide 116 may be the same polynucleotide, such as DNA, as the first polynucleotide 108. The second polynucleotide 116 may be connected to the second polymer 114 at the 3’ end. In this manner, the composite polymer 106 may be associated with a degree of symmetry: in order, the components of the composite polymer 106 may be a first polynucleotide 108, a first polymer 112, a peptide 110, a second polymer 114, and a second polynucleotide 116. Measurements of the second polynucleotide 116 may not be included in the signal, for example, if the second polynucleotide 116 is lowered through the pore too quickly to measure when the motor enzyme skips over the peptide 110 during translocation. Thus, regardless of whether the composite polymer 106 includes a second polynucleotide 116, measurements of a part of the composite polymer 106 including the first polynucleotide 108, the peptide 110, and (optionally) the first polymer 112 and the second polymer 114 may be obtained.
[0039] The system 100 may include electrical circuitry 118 to perform measurements during translocation of the composite polymer 106 through the nanopore 102. The electrical circuitry may include electrodes for providing a potential difference across the nanopore 102 for translocating the composite polymer 106 through the nanopore 102. In such examples, the nanopore 102 may facilitate translocation under the control of enzymatic activity, such as under the influence of a motor enzyme, as described hereinbefore. A binding protein such as helicase may be provided in the nanopore 102 to provide for movement of the composite polymer 106 through the nanopore 102 in a stepwise fashion. In further examples, a combination of applied potential difference and enzymatic activity may be used to carry out controlled translocation of the composite polymer 106.
[0040] The electrical circuitry 118 may include a sensor 120 to measure a current flowing through the nanopore 102 during translocation of the composite polymer 106. For example, a current flowing in a transverse direction of the nanopore 102 or in a surrounding region of the nanopore 102 may be measured during translocation. The current may vary in response to the passage of each nucleotide within the nanopore 102, and may do so to an extent that corresponds with a particular nucleotide or multiple nucleotides being translocated through the nanopore 102. For example, variations of current may be on the order of picoamperes, nanoamperes, or microamperes. The ionic current may be measured by the sensor 120, and the resulting measurement signal may optionally be amplified using an amplifier 122, such as a low-noise amplifier system. In alternative examples, the measurable property may be a voltage difference or a change in resistivity or conductance or other ionic property of the nanopore 102 or its surrounding medium.
[0041] Suitable conditions for measuring ionic currents through transmembrane protein pores are known in the art. The method may be carried out with a voltage applied across the membrane and pore. The voltage used may vary from +2 V to -2 V, or from -400 mV to +400 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential. The methods may be carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salt, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3 -methyl imidazolium chloride, potassium chloride, sodium chloride, caesium chloride or rubidium chloride. The salt concentration may be 3 M or lower and is typically from 0.1 to 2.5 M. Hel308, XPD, RecD and Tral helicases surprisingly work under high salt concentrations which advantageously provide a high signal to noise ratio.
[0042] The methods may be carried out in the presence of a buffer. Any suitable buffer may be used in the method of the invention such as HEPES or Tris-HCl buffer. The pH may vary from 4.0 to 12.0 and is preferably about 7.5. The methods may be carried out at a temperature that supports enzyme function.
[0043] The membrane may be supported on a support structure separating cis and trans compartments comprising ionic solution. An array of membranes each comprising a nanopore may be provided on a support structure comprising a common cis chamber and multiple trans chambers such as disclosed by WO2014 / 064443. The potential difference may be applied across the nanopore between electrodes provide in the cis and trans compartments. Suitable electrode materials include Pt, Pd and Ag. The electrodes may be reference electrodes such as Ag / AgCl. The ionic solution may comprise a redox couple such as potassium ferri / ferrocyanide.
[0044] Possible electrical measurements also include: nanopore tunnelling measurements, FET measurements and optical measurements combined with electrical measurements such as disclosed by WO2005 / 124888, WO2020 / 183172, Soni GV et al., Rev Sci Instrum. 2010 Jan;81(l):014301, Ivanov AP et al., Nano Lett. 2011 Jan 12;1 l(l):279-85 and WO2016 / 009180.
[0045] Polynucleotide as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. In some embodiments the polynucleotide is a single-stranded DNA- RNA hybrid. DNA-RNA hybrids can be prepared by ligating single-stranded DNA to RNA or vice versa. The polynucleotide is most typically double-stranded deoxyribonucleic acid (DNA) or double-stranded ribonucleic nucleic acid (RNA). In some embodiments the polynucleotide may be single-stranded DNA. In some embodiments the polynucleotide is double-stranded RNA. In some embodiments the polynucleotide is a double-stranded DNA-RNA hybrid. Double-stranded DNA-RNA hybrids can be prepared from single-stranded RNA by reverse transcribing the cDNA complement.
[0046] The polynucleotide can be of any length. For example, the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs in length. The polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs in length or 100000 or more nucleotides or nucleotide pairs in length. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-translational modification, for example 5 ’-capping with 7-methylguanosine, 3 ’-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Sizes of nucleic acids, also referred to herein as “polynucleotides” are typically expressed as the number of base pairs (bp) for double-stranded polynucleotides, or in the case of single-stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called “oligonucleotides” and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).
[0047] Nucleotides can have any identity, and include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5 -methylcytidine monophosphate, 5- hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP. A nucleotide may be abasic (i.e. lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e. is a C3 spacer).
[0048] The polynucleotide may be a concatemer comprising multiple sets of complementary strands each linked by a bridging molecule or moiety. The concatemer may be provided by methods such as disclosed by US9910956.
[0049] Any transmembrane pore may be used in the methods provided herein. The pore may be biological or artificial. Suitable nanopores include, but are not limited to, protein pores, polynucleotide pores, and pores formed in solid state substrates such as silicon nitride. A solid-state pore may comprise a nanochannel. The pore may be a DNA origami pore such as disclosed in WO2013 / 083983.
[0050] The protein pore may be a monomer or an oligomer. The pore may be made up of several repeating subunits, such as at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 subunits. The pore may be a hexameric, heptameric, octameric, nonameric or decameric pore. The pore may be a homo-oligomer or a hetero-oligomer. The transmembrane protein pore may be modified from the wild type.
[0051] The transmembrane protein pore may be derived from P-barrel pores or a-helix bundle pores such as a -hemolysin, anthrax toxin and leukocidins, and outer membrane proteins / porins of bacteria, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC, MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A and Neisseria autotransporter lipoprotein (NalP) and other pores, such as lysenin, inner membrane proteins and a outer membrane proteins, such as WZA and ClyA toxin. The transmembrane protein pore may be derived from derived from Spl or haemolytic protein fragaceatoxin C (FraC). Examples of suitable pores derived from CsgG are disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318 and W02019 / 002893, each of which is hereby incorporated by reference in its entirety. The pore and / or enzyme may be modified to improve interaction between the two, such as disclosed in WO2015 / 166276. In further examples, an alternative type of pore may be used for translocating and obtaining measurements of the composite polymer 106.
[0052] The nanopore may be supported in a membrane having cis and trans openings on opposite sides of the membrane. The membrane may preferably be an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers such as disclosed in WO2014 / 064444. The block copolymer may be a diblock (consisting of two monomer sub-units), but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphiphiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. The amphiphilic molecules may be chemically-modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer may be planar. The amphiphilic layer may be curved. The membrane may be a lipid bilayer. The lipids may comprise a head group, an interfacial moiety and two hydrophobic tail groups which may be the same or different. Suitable lipid bilayers are disclosed in W02008 / 102121, W02009 / 077734 and W02006 / 100484.
[0053] Suitable polynucleotide binding proteins for controlling translocation of the polynucleotide are known in the art and include polymerases, exonucleases, helicases and topoisomerases. A preferred enzyme is a helicase which may be or be derived from a Hel308 helicase, a RecD helicase, such as Tral helicase or a TrwC helicase, a XPD helicase or a Dda helicase. The helicase may be any of the helicases, modified helicases or helicase constructs disclosed in WO2013 / 057495, WO2013 / 098562,
[0054] WO20 13 / 098561, WO2014 / 013259; WO2014 / 013262 and W02014 / 013260. The helicase may be added to the polynucleotide during sample preparation and stalled by one or more spacers as disclosed in WO2014 / 135838. Motor enzymes other than helicase may alternatively be used.
[0055] Translocation of the polynucleotide may be controlled by alternative methods disclosed for example in W02019 / 006214 and W02020 / 016573. The polynucleotide may comprise a polymer leader sequence which preferentially threads into the pore. The leader is preferably negatively charged and may be a polynucleotide, such as DNA or RNA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide. The leader sequence may form part of a Y adapter that my comprise (a) a double-stranded region and (b) a single-stranded region or a region that is not complementary at the other end. Leader sequences and Y adapters suitable for use are disclosed for example in WO2017 / 149316.
[0056] As discussed herein, a leader may be comprised in the composite polymer, i.e. “conjugate”. A leader may be comprised in the conjugate by being attached to the peptide. For example, a first end of the peptide may be conjugated to a polynucleotide and a second end of the peptide may be attached to the leader. As explained in more detail herein, the second end of the peptide may be attached to the leader by any suitable means.
[0057] In some embodiments the leader is directly attached to the second end of the peptide. In some embodiments the leader is attached to the second end of the peptide by a linker. In some embodiments the leader is charged. In some embodiments the leader is uncharged. In some embodiments the leader is negatively charged or positively charged, typically negatively charged.
[0058] In some embodiments the leader is polymeric. In some embodiments the leader is a charged polymer, e.g. a negatively charged polymer. In some embodiments the leader comprises a polymer such as PEG or a polysaccharide. In such embodiments the leader may be from 10 to 150 monomer units (e.g. ethylene glycol or saccharide units) in length, such as from 20 to 120, e.g. 30 to 100, for example 40 to 80 such as 50 to 70 monomer units (e.g. ethylene glycol or saccharide units) in length.
[0059] In some embodiments the leader is or comprises a polynucleotide. In embodiments wherein the leader is a polynucleotide the leader may be the same sort of polynucleotide as the polynucleotide used in the conjugate, or it may be a different type of polynucleotide. For example, the polynucleotide in the conjugate may be DNA and the leader may be RNA or vice versa. In some embodiments the polynucleotide in the conjugate comprises or consists of DNA (e.g. dsDNA) and the leader does not consist of DNA, although in some embodiments the leader may comprise one or more nucleotides in addition to other monomer units; for example the leader may comprise one or more spacers as described herein). In some embodiments the polynucleotide in the conjugate comprises or consists of ssDNA and the leader does not consist of ssDNA (e.g. the leader may comprise one or more spacers as described herein).
[0060] In some embodiments the leader comprises one or more spacers as described herein. In some embodiments the leader may comprise one or more abasic spacers i.e. one or more spacers in which the bases are removed from one or more nucleotides in the polynucleotide adapter. In some embodiments the leader may comprise peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or a synthetic polymer with nucleotide side chains. In some embodiments the leader may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuri dines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxy -thymidines (ddTs), one or more dideoxy-cytidines (ddCs), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2’-O-Methyl RNA bases, one or more Iso-deoxycytidines (Iso-dCs), one or more Iso-deoxyguanosines (Iso-dGs), one or more C3 (OC3H6OPO3) groups, one or more photo-cleavable (PC) [OC3H6-C(O)NHCH2-CeH3NO2-CH(CH3)OPO3] groups, one or more hexandiol groups, one or more spacer 9 (iSp9) [(OCFbCFbjsOPCh] groups, or one or more spacer 18 (iSpl8) [(OCFbCFbjeOPCh] groups; or one or more thiol connections. A leader may comprise any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9 and iSpl8 spacers are all available from IDT®. A leader may comprise any number of the above groups. For example, a leader may comprise about 5 to Ibout 100, such as from about 10 to about 50, e.g. from about 20 to about 40 spacers as described herein, e.g. C3, iSp9 and / or iSp 18.
[0061] In some embodiments the leader is directly attached to the peptide i.e. to the second end of the peptide in the conjugate. In some embodiments the leader is attached to the second end of the peptide by a chemical bond. In some embodiments the leader is attached to the second end of the peptide by a covalent bond. In some embodiments the leader is attached to the second end of the peptide by a linker. Any suitable linker may be used. In some embodiments the linker is a polynucleotide as described herein. In some embodiments the linker is a synthetic polymer such as a PEG. In some embodiments the linker is a polynucleotide of the same type of polynucleotide as the polynucleotide portion of the conjugate (e.g. in some embodiments the polynucleotide portion of the conjugate comprises DNA and the linker comprises DNA) and the leader comprises one or more nucleotides of a different sort to those comprised in the polynucleotide portion of the conjugate. In some embodiments the polynucleotide portion of the conjugate comprises DNA, the linker comprises DNA, and the leader comprises one or more non-DNA nucleotides as described herein, e.g. one or more spacers as described herein.
[0062] In some embodiments the leader is attached to the second end of the peptide at the N-terminus or the C-terminus of the peptide. In some embodiments the leader is attached to the second end of the peptide side chain group of a residue (e.g. an amino acid residue) in the peptide.
[0063] In some embodiments the target peptide has a naturally occurring reactive functional group which can be used to attach the leader. For example, a cysteine residue can be used to form a disulphide bond to the leader or to a modified group thereon.
[0064] In some embodiments the target peptide is modified in order to facilitate its attachment to the leader. For example, in some embodiments the peptide is modified by attaching a moiety comprising a reactive functional group for attaching to the leader. For example, in some embodiments the peptide can be extended at the N-terminus or the C-terminus by one or more residues (e.g. amino acid residues) comprising one or more reactive functional groups for reacting with a corresponding reactive functional group on the leader. For example, in some embodiments the peptide can be extended at the N-terminus and / or the C-terminus by one or more cysteine residues. Such residues can be used for attachment to the leader e.g. by maleimide chemistry (e.g. by reaction of cysteine with an azido-maleimide compound such as azido-[Pol]-maleimide wherein [Pol] is typically a short chain polymer such as PEG, e.g. PEG2, PEG3, or PEG4; followed by coupling to an appropriately functionalised leader, e.g. via a BCN group for reaction with the azide). For avoidance of doubt, when the peptide comprises an appropriate naturally occurring residue at the N- and / or C-terminus (e.g. a naturally occurring cysteine residue at the N- and / or C-terminus) then such residue(s) can be used for attachment to the leader. In some embodiments a residue in the target peptide is modified to facilitate attachment of the target peptide to the leader. In some embodiments a residue (e.g. an amino acid residue) in the peptide is chemically modified for attachment to the leader. In some embodiments a residue (e.g. an amino acid residue) in the peptide is enzymatically modified for attachment to the leader.
[0065] The attachment chemistry between the leader and the peptide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react with amines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like.
[0066] Other suitable chemistry for attaching the leader to the peptide includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0067] (f) copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);
[0068] (g) strain-promoted azide-alkyne cycloadditions; including alkene and azide [3+2] cycloadditions; alkene and tetrazine inverse-demand Diels-Alder reactions; and alkene and tetrazole photoclick reactions;
[0069] (h) copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);
[0070] (i) the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and (j) the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond. Any reactive group may be used to attach the leader to the peptide. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethyleneglycol; 3,3 ’-dithiodipropionic acid di(N- hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4’ -diisothiocyanatostilbene-2, 2’ -disulfonic acid disodium salt; Bis[2-(4- azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acid N- hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N- hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, particularly in Table 3 of that application.
[0071] In some embodiments the reactive functional group is comprised in the leader and the target functional group is comprised in the second end of the peptide prior to the attachment step. In other embodiments the reactive functional group is comprised in the second end of the peptide and the target functional group is comprised in the leader prior to the conjugation step. In some embodiments the reactive functional group is attached directly to the peptide. In some embodiments the reactive functional group is attached to the peptide via a spacer. Any suitable spacer can be used. Suitable spacers include for example alkyl diamines such as ethyl diamine, etc.
[0072] The rate of capture of polynucleotide during translocation through the pore can be enhanced by coupling it to the membrane. Suitable coupling moieties are disclosed for example in WO2017 / 149316 and WO2012 / 164270. Reference is also made to methods disclosed for example in WO 2021 / 111125.
[0073] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.
[0074] After amplification, the signal may be digitised and provided to a data processing system 124 for analysis. The data processing system 124 may contain hardware and software components capable of estimating the sequence of nucleotides based on the measurement signal to identify the polynucleotide 108, and / or the bridging polymer 112 and / or the second polymer 114 included in the composite polymer 106. The data processing system 124 may further include stored software components, such as machine learning models or algorithms, capable of characterising the peptide 110 included in the composite polymer 110. Methods of characterising the aforementioned components of the composite polymer 106 are as described below.
[0075] The data processing system 124 may be a desktop, a laptop, a tablet, or a server, or any combination of the above. In some examples, the data processing system may be a dedicated device for sequencing polymers, such as produced by Oxford Nanopore Technologies (RTM). The data processing system 116 may include a central processing unit (CPU) having one or more processors including integrated circuit microprocessors that may optionally be multithreaded, such as Intel (RTM) Xeon or Intel Core i3 / i5 / i7 series processors. The data processing system 116 may also optionally include one or more graphic processing units (GPUs) such as Nvidia (RTM) Al 00 or Hl 00 GPUs or AMD GPUs. Similarly, the data processing system 116 may contain software that may include source code, object code, firmware, etc. in any suitable language. For example, the source code may be written in Python, C, C++, Rust, Julia, etc and may use specific development frameworks, libraries, or packages, including PyTorch, TensorFlow, Keras, CUDA, etc. The examples provided herein for the hardware and software components do not constitute an exhaustive list; many more similar components may be alternatively or additionally included. The data processing system 124 may be configured to provide instructions to the electrical circuitry 118, for example to control the translocation of the composite polymer 106 with respect to the nanopore 102.
[0076] Fig. IB shows an example measurement signal 126 obtained by translocating the composite polymer 106 through the nanopore 102. In this example, the composite polymer 106 comprises a (DNA) polynucleotide 108 connected to a peptide 110 by a first (polythymine) polymer 112, and the peptide 110 is further connected to a second (polythymine) polymer 114. In the example measurement signal 126, current measurements have been normalised so that the 10% quantile is 0 and the 90% quantile is 1. The normalisation allows for correction of small variations in current levels between different pores on the instrument. The time on the x-axis is measured in samples taken at 1.5kHz. The region 128 of the measurement signal 126, i.e. before sample 6100, corresponds to the polynucleotide 108. The measurement signal 126 also includes a region 130 corresponding to the bridging polymer 112, a region 132 corresponding to the peptide 110. The signal 126 further comprises a region 134 corresponding to the second polymer 114. The uniformity of polythymine polymers lead to the regions 130 and 134 corresponding to being relatively flat in comparison to the rest of the signal 126. These relatively flat regions are at approximately 6100 to 6600 samples and 6700 to 7100 samples in the diagram. The relatively short region 132 between approximately 6600 to 6700 samples corresponds to the peptide 110.
[0077] The apparatus 100 may process the measurement signal 126 using the data processing system 124 to characterise the peptide 110. The processing may involve the following stages: (i) signal segmentation to identify a peptide region, and (ii) peptide classification based on processing the peptide region. During the segmentation stage, a region connecting the polynucleotide 108 to the peptide 110 may be identified. For example, the signal 126 may be sequenced or “basecalled” to identify an end of the polynucleotide 108 connected to the peptide 110. Alternatively, the signal 126 may be processed to identify a first amino acid in the sequence associated with the peptide 110 to identified the region connecting the peptide 110 to the polynucleotide 108. Similarly, when a bridging polymer 112 is included in the composite polymer 106, parts of the bridging polymer 112 may be identified as the region connecting the polynucleotide 108 and the peptide 110. Upon identification of the region, the signal 126 may be segmented to isolate a peptide region 134 including measurements of the peptide and any flanking regions, such as the polymer regions 130 and 132 corresponding to the bridging polymer and the second polymer. Segmentation methods described may be supplemented or replaced by either (a) a hidden Markov model segmentation, or (b) statistical segmentation.
[0078] The peptide region 134 determined using segmentation may be processed to classify and / or characterise the peptide 110 using a classifier. Classification of the peptide 110 may include identification of the amino acids to specifically identify the peptide, or categorisation of the peptide 110 under a broad class of peptides. The passages below describe methods used for segmentation of the composite polymer and classification of the peptide included in the composite polymer. Further, the methods described also cover protein labelling based on characterising peptides associated with multiple composite polymers. The segmentation step (i) and the classification step (ii) may also be employed during training of the classification model used by the classification algorithm.
[0079] Fig. 2 shows a method of characterising a peptide 110 in a composite polymer 106. A signal 226 illustrative of the measurement signal 126 received using the system 100 is shown. The signal 226 may comprise a polynucleotide region 230 including measurements of the polynucleotide 108, and a peptide region 234 including measurements of the peptide 110 present in the composite polymer 106. The signal 226 may further comprise a first region 230 including measurements of the first (bridging) polymer 112 present in the composite polymer 106, and may also comprise a second region 232 including measurements of the second polymer 116. In the passages below, the method of characterising an unlabelled peptide by segmenting the signal 226 followed by processing using a classification algorithm is described. Subsequently, a method of training the classification model used by the classification algorithm based on measurements of a peptide either associated with a predetermined label or a polynucleotide barcode is described.
[0080] The method shown in Fig. 2 illustrates an example method of segmenting the signal 226 and classifying the segmented signal. A number of alternative methods may be employed for segmentation as described further below. In the example shown in Fig. 2, a machine learning (ML) model 236 is used to basecall the signal 226 to thereby estimate a read sequence 238. For example, the read sequence 238 may include estimates of nucleotide units present in the composite polymer 106. The ML model 236 may, for example, comprise a neural network that is trained on signal data having labelled polymer units. For example, the ML model 236 may include a recurrent neural network (RNN) such as a long short-term memory (LSTM) network, or a transformer, or any other neural network capable of processing sequence data, and may have been trained to predict polymer units associated with measurement signals. In some examples, a commercially available basecaller, such as the Dorado basecaller developed by Oxford Nanopore Technologies, may be adopted to this end. In this manner, the polymer units included in the composite polymer 106 may be determined in the form of the read sequence 238.
[0081] Once the read sequence 238 is determined, the method may proceed to segment the signal 226 for characterising the peptide 110 associated with the signal 226. For example, segmenting the signal 226 may include identifying a region connecting the polynucleotide 108 with the peptide 110. In the example shown, the first polymer region 230 corresponding to measurements of the first (bridging) polymer 112 in the signal 226 is identified as a region connecting the polynucleotide 108 with the peptide 110. By identifying the region 230 connecting the polynucleotide 108 with the peptide 110, the method may determine a segmentation point for the signal 226.
[0082] Determining the first polymer region 230 can be challenging. For example, measurements of the polynucleotide 108 may include measurements of a nucleotide base that is also present in the first (bridging) polymer 112. As described above, the first polymer 112 may be a polythymine comprising a sequence of thymine bases. The polynucleotide 108 may be DNA and may also thereby include thymine bases. In such examples, measurements of individual thymine bases at the nanopore 102 may span long stretches of the signal 226 as the time taken for a base to translocate through the nanopore 102 may experience random variations. As a result, the same composite polymer 106 may exhibit completely different current traces in the signal 226 for separate translocations. This can make it challenging to accurately distinguish the polynucleotide region 228 from the first polymer region 230 within the signal 226. The example shown in Fig. 2 addresses the technical problem by identifying positions of the sequence of polymer units included in the first (bridging) polymer 112 within the read sequence 238. The basecalled version of the signal may be able to distinguish measurements corresponding to an actual sequence of repeating bases in the bridging polymer from measurements corresponding to the same sequence of repeating bases apparently associated with the polynucleotide. For example, a polythymine bridging polymer associated with a sequence of repeating thymine DNA bases may be distinguished from measurements of the polynucleotide associated with the same sequence of thymine DNA bases. In this manner, basecalling the signal may enable accurate identification of the region corresponding to the bridging polymer.
[0083] Further, the ML model 236 may process the signal 226 to produce a mapping 242 that associates positions within the read sequence 238 to corresponding points within the signal 226. In some examples, the ML model 236 may produce the mapping 242 as an output in the same step as producing the read sequence 238. In other examples, the mapping 242 may be produced by the ML model 236 in a different step separately from the step that involves generating the read sequence 238. The mapping may be used to obtain a position in the signal 226 corresponding to a join of the polynucleotide 108 and the peptide 110. Upon locating positions of the sequence of polymer units of the first (bridging) polymer 112 within the read sequence 238, a position corresponding to an end of the first (bridging) polymer 112 may be determined. For example, a position within the read sequence 238 corresponding to the end of the first polymer 112 connected to the polynucleotide 108 may be determined. In other examples, a position close to or near the boundary of the first polymer 112 and the polynucleotide 108, but within the polynucleotide 108, may be determined. In any case, the mapping 242 may subsequently be used to map the location of end of the first polymer 112 to a segmentation point 244 within the signal 226. In cases where the first polymer 112 comprises nucleotide units, the accuracy of base estimation within the read sequence 238 may be relatively higher close to the boundary between the first polymer 112 and the polynucleotide 108 in comparison to the boundary between the first polymer 112 and the peptide 110. As a result, selecting a location corresponding to a connection or boundary between the first polymer 112 and the polynucleotide 108 may result in a higher accuracy of segmentation.
[0084] Once the segmentation point 244 corresponding to the location within the read sequence 238 is determined, the signal 226 may be segmented at the segmentation point 244. In this manner, a segment 246 comprising measurements of the peptide 110 may be obtained. As shown in Fig. 2, the segment 246 may also include one or more measurements of the first (bridging) polymer 112. Furthermore, where a second polymer 114 is included in the composite polymer 106, the segment 246 may include measurements of the second polymer 114. In any case, by segmenting the signal 226 at the segmentation point 244, the peptide 110 may be isolated within a small and computationally less memory intensive section of the signal 226. As a result, the computational efficiency of subsequent processing of the signal segment 246 may be enhanced. Furthermore, a classification model that is used by the classification algorithm and trained specifically on peptide fragments may be applied to the segment 246; in this manner, the accuracy and generalisability of peptide classification may be enhanced in comparison to applying the classification algorithm to the entirety of the signal 226. Fig. 2 shows that the signal segment 246 comprises measurements of the peptide 110 as well as the first and second polymers. The signal segment 246 is processed by a classifier algorithm 248. For example, the classifier algorithm 248 may be a machine learning algorithm relying on a machine learning model with a neural network architecture. In some examples, a classification model used by the classifier algorithm 248 may be trained based on labelled examples of measurements of peptide amino acid sequences obtained using the nanopore 102 as described further below. A trained classification model of the classifier algorithm 248 may be used to process the segment 246 to determine a peptide classification 250 associated with measurements included in the segment 246.
[0085] In some examples, the machine learning model employed may comprise a recurrent neural network such as a Convolutional Neural Network with Long Shortterm Memory (CNN-LSTM) units. In such examples, LSTM units may process the signal segment 246 to produce outputs in an embedding space. The outputs thus produced by the LSTM units may be projected using a linear layer to an output space comprising multiple channels, wherein each channel encodes a score or a probability for the signal segment 246 being associated with a given peptide class. Outputs within the output space may subsequently be condensed to a single set of outputs indicating the peptide classification using a max pooling layer. In one example, the neural network used by the classifier algorithm may consist of four ID CNN layers, each with 64 features and with kernel size 5, followed by a 128-channel, two-layer bidirectional LSTM. In such examples, the machine learning model may be coded in Pytorch and trained using standard batch-gradient-descent methods. For example, the AdamW optimiser may be used to train the machine learning model over 400 epochs using a learning rate decreasing exponentially from 0.02 to 0.0004 over the training period. In other examples, a different number of CNN layers, features, kernel size, channel size may be utilised. Other non-standard gradient-descent methods may be used. Alternative optimisers may also be used, and training may proceed for a different number of epochs with learning rate varying in a different range over the training period. For example, model training may be performed using a loss function based on a cross-entropy loss. Several variations of the specific examples provided thus far may be adopted for peptide classification of the segment 246. For example, neural network architectures such as transformers, gated recurrent units, and multi-layer perceptrons may be used to processing the segment 246 and classify the peptide 110. Furthermore, machine learning methods such as decision trees and support vector machines may be utilised for the purpose of classification. Regardless of the specific classifier model used by the classification algorithm 248 to perform classification, the methods describe may lead to determining of the peptide classification 250 for the signal 226.
[0086] The segmentation and classification steps described thus far in the context of inference may also be adopted for training the classification model used by the classification algorithm 248. Training may comprise processing a signal 226 associated with a composite polymer 106 including a labelled peptide. For example, the peptide 110 may have a predetermined ground truth classification. A ground truth classification may, for example, refer to a desired output from the classification algorithm as provided within a dataset comprising peptides with known labels. For example, in machine learning, the term ground truth may refer to a reality you want to model with your supervised machine learning algorithm. In this manner, the ground truth classification may be a target for training or validating the classification model used by the classification algorithm 248 using the labelled dataset. The classification model used by the classifier algorithm 248 may then be trained based on the predetermined label for the peptide 110. For example, signal 226 obtained using the nanopore 102 may be segmented and classified as described previously to estimate a peptide classification for the peptide 110. Training of the classification model associated with the classification algorithm 248 may involve iterating on a number of signals associated with labelled peptides. For each such signal, the training algorithm may compare the estimated classification 250 for the peptide with the ground truth label. A loss function may quantify the difference between the estimated classification 250 and the ground truth label and propagate the error in each iteration to update parameter values associated with the classification model. For example, the weights and biases of a neural network classifier may be updated in this manner. By iterating through a number of such examples, the classification model used by the classifier algorithm 248 may be trained to characterise a peptide accurately. In this manner, a classification for the peptide 110 may be estimated with enhanced accuracy. Fig. 3 shows a schematic of a further method for characterising a peptide 110 in a composite polymer 106. In the example shown, a polynucleotide barcode 340 may be connected with the peptide 110. Polynucleotide barcoding is a method of sample identification based on species-specific or individual-specific polynucleotide barcodes. The polynucleotide barcode may be a segment of the polynucleotide that serves as a marker for the species or individual. For example, the polynucleotide may be a standardised DNA sequence. The polynucleotide barcode 340 may be serve as a marker for the connected peptide 110. For example, the methods described thus far may be used to process multiple signals, each signal associated with a peptide connected to a respective polynucleotide barcode. As an example, the signals may have been obtained by translocating different composite polymers through the nanopore 102 or through multiple nanopores. The polynucleotide barcode 340 for a given signal may be identified by using at least part of the signal 326. For example, the polynucleotide barcode 340 may be identified based on the read sequence 338 determined by the ML model 336. Further, the corresponding peptide 110 may be characterised by employing the segmentation and classification methods described above. As a result, the characterised peptide 350 may be associated with the identified polynucleotide barcode 340. In this manner, peptides associated with a given polynucleotide barcode may be grouped for further analysis or experimentation. The polynucleotide barcode 340 may also be used to prescribe a ground truth classification label for the peptide 110 for training the classification model used by the classification algorithm 248. For example, the polynucleotide barcode 340 may be identified by using at least part of the signal 326. As an example, the ML model 236 may be trained to identify the polynucleotide barcode 340 based on the read sequence 338 for the signal 326. In the example shown in Fig. 2, the polynucleotide 340 may be a DNA barcode; in other examples, an RNA barcode or another relevant barcode may be used. Training the classification model may include gathering a dataset of signals each associated with a peptide and a polynucleotide barcode. As described for training using labelled peptides, training of the classification model used by the classification algorithm 348 in this example may involve iterating on a number of signals associated with different peptides. Each different peptide may be connected to a distinct polynucleotide barcode. In this manner, the polynucleotide barcode 340 for a given signal 326 may serve as a marker for distinguishing between different peptides, or for assigning a corresponding label for the associated different peptides. During a training iteration, the classification algorithm 348 may provide an estimated classification label 350 for the peptide 110 associated with a given signal 326. For each such signal, the training algorithm may compare the estimated classification 350 for the peptide with the label assigned using the polynucleotide barcode 340. A loss function may quantify the difference between the estimated classification 350 and the ground truth label and propagate the error to update parameter values associated with the classification model. For example, the weights and biases of a neural network classifier may be updated in this manner. By iterating through a number of such signals, the classification model associated with the classifier algorithm 348 may be trained to characterise a given peptide accurately. In this manner, a classification for the peptide 110 may be estimated with enhanced accuracy. Although the polynucleotide barcode 340 was identified above based on the read sequence 248 corresponding to the signal 326, in some examples, polynucleotide identification may be performed subsequent to segmentation of the signal 326. For example, segmentation of the signal 326 based on the segmentation point 344 may result in a first segment 346 comprising measurements of the peptide 110 alongside a second segment (not shown in Fig. 2) comprising measurements of the polynucleotide barcode 340. The second signal segment may thereafter by processed using a machine learning model, such as the ML model 336 to determine the polynucleotide barcode 340 and to thereby determine a ground truth label for the peptide 110. In this manner, the polynucleotide barcode 340 may be identified based on basecalling after segmentation of the signal 326 rather than prior to segmentation.
[0087] The passages below describe further methods signal segmentation that can be used during inference and / or training of the classification model.
[0088] Fig. 4 shows a schematic of a segmentation method 300 based on a hidden Markov model (HMM). In this method, the signal segment 446 obtained by segmenting the signal 226 using the mapping 242 may be further segmented to further isolate a peptide region comprising measurements of the peptide 110. By segmenting the signal segment 446, the accuracy of classification may be enhanced further still. The method shown schematically in Fig. 4 involves partitioning the segment 446 of the signal 226 into a number of partitions 452 (or chunks). For example, the partitions may be of equal length and may therefore each comprise an equal number of measurements of the composite polymer 106. Moreover, the partitions may overlap with neighbouring partitions, i.e. each partition may include measurement data from neighbouring partitions. In some examples, each partition may include 30 samples including an overlap of 15 samples between neighbouring partitions. The measurements included in each partition may subsequently be processed to compute a respective statistic associated with the measurements included in the partition. In some examples, a mean of the measurements included in each partition may be computed. In other examples, a variance or another high-order statistic of the measurements may be computed. In further examples, a combination of multiple statistical quantities may be computed.
[0089] Once a statistic is computed for each of the partitions 452, one or more likelihood functions of HMMs 454 may be evaluated based on the computed statistic. Each such likelihood function may be associated with a respective specific component of the composite polymer 106. For example, each likelihood function may be a function that maps a statistic for a given partition to a score or probability of the partition being associated with a specific component of the composite polymer 106. As an example, a likelihood function associated with a component comprising the peptide 110 may be used. In such examples, an estimated statistic for measurements of the peptide 110 may be predetermined based on previously processed peptide signals. Subsequently, the likelihood function may be evaluated to assign a score to each partition based on whether the statistic for the partition corresponds with the predetermined statistic for the peptide 110. In this manner, partitions associated with the peptide 110 in the segment 446 may be estimated. Subsequently, contiguous partitions indicating a high score or probability of including measurements of the peptide 110 may be located to determine regions of the segment 446 associated with the peptide 110. The isolated region may thereafter be processed using the classifier algorithm 248 to identify the peptide 110 associated with the composite polymer 106. By evaluating the likelihood function in this manner, the classification algorithm 248 may be required to process a smaller portion of the segment 446 focusing on measurements of the peptide 110. The resulting classification may therefore be more computationally efficient and accurate. In alternative examples, a component comprising a combination of the first polymer 112 and the second 114 may be used. In such examples, the component may not include a contiguous set of partitions. Each partition may then be assigned a score associated with either the first polymer 112 or the second polymer 114 or a combination of both polymers. The partitions associated with the first polymer 112 and / or the second polymer 114 may then be used to identify the partitions associated with the peptide 110. In this manner, a segment of the signal including measurements of the peptide may be determined for classification. In further examples, the specific component may correspond to the polynucleotide 108.
[0090] Fig. 4 shows four HMM likelihood functions Co, Ci, C2 and C3 being evaluated as part of HMM segmentation. In this example, the composite polymer 106 comprises the polynucleotide 108 as well as a first and a second polymer connected to the peptide 110. The likelihood functions Co, Ci, C2 and C3 may therefore be associated with specific components corresponding to or comprising the polynucleotide 108, the first (bridging) polymer 112, the peptide 110, and the second polymer 114, respectively. Each likelihood function may be constructed such that it maps statistics of different partitions to scores or probabilities indicating the likelihood of the measurements included in the partition being associated with the specific component for the likelihood function. For example, a likelihood function Co for the polynucleotide 108 may map the computed statistics to scores or probabilities indicative of a likelihood that the partition contains measurements of the polynucleotide 108. Similarly, the likelihood function Ci may generate map the statistics to scores indicating likelihoods of the measurements being associated with the first (bridging) polymer 112. Once the scores for each of the four likelihood functions are obtained, Fig. 4 shows that the scores are combined to identify regions 456 in the segment 446 corresponding to measurements of one or more specific components of the composite polymer 106. For example, contiguous partitions that are assigned high scores by the likelihood function C2 may be combined to locate a peptide region within the segment 446. Similarly, partitions assigned high scores by the remaining likelihood functions may be respectively combined to locate regions corresponding to the respective specific components of the composite polymer 106. Evaluations of these remaining likelihood functions Co, Ci and C3 may be employed to improve upon the estimated peptide region identified based on the likelihood function C2 for the peptide 110. The segment 446 may subsequently be segmented to isolate the peptide region 458 containing measurements of the peptide 110. As a result, the isolated peptide region 458 may be processed by the classifier algorithm 448 to identify the peptide 450 in the composite polymer 106.
[0091] The likelihood functions evaluated as discussed above may be determined based on predetermined distribution functions. For example, the likelihood functions Comay be determined using a Gaussian distribution defined by a mean value corresponding to an average normalised current level expected for the polynucleotide 108, and a variance value corresponding to an expected variation in average normalised values current levels for the polynucleotide 108. The Gaussian distribution may then map an input statistic value, such as a mean value corresponding to a partition, to a probability that the partition is associated with measurements expected for a polynucleotide 108. Likelihood functions for other specific components of the composite polymer 106 may be determined in a similar manner. In some examples, combinations of Gaussian distribution functions may be adopted. In other examples, the Gaussian distribution used for evaluating the likelihood function for the polynucleotide 108 may be algebraically manipulated to obtain the likelihood function for the remaining specific components of the composite polymer 106. For example, a set of predetermined rules may be used to set the unique parametric dependence of each of the likelihood functions on the Gaussian distribution. In such examples, the likelihood function Ci for the first (bridging) polymer may be the same as the likelihood function C3 for the second polymer. Further, the likelihood functions for the peptide and the polynucleotide 108 may be determined using the same rules distinguishing the corresponding partitions from the partitions associated with the first and second polymers. As an illustrative example, the four likelihood functions Co, Ci, C2 and C3 may be defined using a Gaussian normal distribution that is a function of the mean m and / or the standard deviation 5 of measurements in each partition of the set of partitions 452. The Gaussian normal distribution may be defined as: where p indicates the mean of the Gaussian distribution and c indicates the standard deviation of the Gaussian distribution. The Gaussian distribution in this example is normalised to have maximum value 1. In this example, the mean of the Gaussian distribution may be set top = a = 0.66 to reflect the average normalised current level in a polythymine segment, and the standard deviation of the Gaussian distribution may be set to c = b = 0.2 to reflect the possible variation in average normalised currents across different polythymine segments. For different values of the mean m and / or standard deviation 5 provided as the function’s argument x, the Gaussian distribution or a function thereof may provide a corresponding score for the associated partition. For example, the likelihood functions may be defined as:
[0092] CQ= 1 — 0.7 G(m; a, b
[0093] Ci = G(m; a, b
[0094] C2= 1 — 0.7 G(m; a, b and
[0095] C3= G (m; p, <J) .
[0096] The table below indicates evaluations of each of the likelihood functions for two different mean signal values, m = 0.66 and m = 0.1, provided as argument:
[0097] Based on evaluating these functions, a partition with a mean signal value m close to the value of a = 0.66 may be assigned scores close to 1 to indicate a high likelihood of the measurements being associated with the first or the second polymer, and relatively low scores ~0.3 for the other specific components of the composite polymer 106. Similarly, a partition that has a mean signal value m very different from the value of a (say m = 0.1) may be assigned a low score (-0.02) indicative of a low likelihood of the measurements being associated with the first or second polymer. In this manner, partitions associated with the first or second polymer may be identified, and as a result partitions associated with the peptide 110 may be determined.
[0098] In further examples, other distributions or combinations thereof may be used to determine the likelihood functions Co, Ci, C2 and C3, involving a set of predetermined rules. For instance, a multiplicative combination of two or more Gaussian distributions may be used to determine the likelihood functions Ciand C3 for the polymers.
[0099] As an illustrative example, the table below shows a set of likelihood functions defined using the Gaussian normal distribution.
[0100] Such combinations of Gaussian distributions may address scenarios where measurement values of an end of the peptide signal connected to the first polymer are close to measurement values of the first polymer level, and include noise. In such scenarios, the relatively simpler likelihood functions discussed previously rules may misclassify part of the peptide signal as the first polymer. The additional complexity of the rules for defining likelihood functions above allow the segmentation procedure to prevent signal partitions with high standard deviation comprising measurements of the peptide from being classified as partitions associated with the first polymer.
[0101] In yet further examples, however, the rules determining the likelihood function for each specific component of the composite polymer 106 may be encoded by a statistical model. For example, a machine learning model may be trained to encode the rules that map statistics determined for the partitions to respective scores indicating a probability of a given specpfic component being included in the partition. Such statistical models may be trained using labelled samples of each of the specific components of the composite polymer 106, and may be used during segmentation to accurately locate the corresponding regions within the segment 446. Regardless of how the likelihood functions are determined, all four likelihood functions may be evaluated to obtain scores indicating the partitions associated with each of the specific components of the composite polymer 106.
[0102] Since the specific components of the composite polymer 106 may be expected to be arranged sequentially in a predetermined order, i.e. a polynucleotide 108 followed by a polymer 112, followed by the peptide 110 and the second polymer, this a priori information may be utilised to further improve upon the accuracy of locating the peptide region 458. For example, scores obtained using the likelihood functions Co, Ci, C2 and C3 may be processed by taking into account values of a transition matrix together with group scores determined by the likelihood functions. In such examples, the transition matrix may be configured to estimate an optimal assignment of a specific component to a given partition based on the sets of scores determined using the four likelihood functions. In particular, the optimal assignment estimated using the transition matrix may be constrained to be consistent with the sequence or order of the corresponding specific components within the composite polymer 106. In some examples, a Viterbi algorithm may be used to determine the optimal assignments from the transition matrix. In such examples, because the Viterbi algorithm is able to take account of the correct sequence of states (polynucleotide-first polymer-peptide-second polymer), as encoded by the transition matrix, and also the likelihoods of each state at each partitions as evaluated by the likelihood functions, it is able to correctly determine the location of the peptide region within the signal 226.
[0103] For example, having calculated likelihood scores for each of the four specific components of the composite polymer 106 and for each signal partition, the Viterbi algorithm may be used to calculate the best assignment of a specific component of the composite polymer 106 to each partition, consistent with the transition sequence given as: polynucleotide 108, first polymer 112, peptide 110 and second polymer 114. The Viterbi algorithm is described, for example, in Durbin et al, Biological sequence analysis, Cambridge University Press 1998. The result of the Viterbi algorithm is a segmentation of the signal 226, locating the peptide portion for further analysis.
[0104] By using the transition matrix to the scores, the computational complexity of locating regions in the segment 446 corresponding to the specific components may be reduced due to the consistency achieved between the located regions and the corresponding specific components in the composite polymer 106. In examples, the computational complexity of using the transition matrix to locate the peptide region may depend linearly on the length of signal segment 446. In contrast, locating the peptide region by individually evaluating and comparing the likelihood of each state for each partition may have a computational complexity depending exponentially on the length of the signal segment 446. In this manner, using the transition matrix may significantly reduce the computational resources required for carrying out segmentation and, accordingly, for peptide classification. Moreover, the technical problem of distinguishing a thymine base within the polynucleotide 108 from a repeating stretch of thymine bases in the first polymer is also addressed by using the transition matrix, which incorporates information relating to a transition sequence of composite polymer’s specific components.
[0105] Thus, a benefit of using the HMM together with the Viterbi algorithm is that it provides a principled method of finding an optimal, or close to optimal, assignment of states to locations in time, taking into account allowed transitions between specific components of the composite polymer 106 as well as assigned likelihoods for individual locations (partitions). The Viterbi algorithm can be executed in a time which depends linearly on the length of the signal, while evaluating separately the likelihood for each possible state sequence would require a time depending exponentially on the signal length.
[0106] Fig. 5 shows example outputs of the likelihood functions and the transition matrix for a signal segment. The signal segment 546 shown includes a few leading measurements of the polynucleotide 108 remaining after the initial segmentation, followed by a relatively flat region corresponding to the first (bridging) polymer 112. Next, a small region comprising significant variations in the measurements corresponds to the peptide 110, followed by another relatively flat region corresponding to measurements of the second polymer 114. Evaluations of the likelihood functions Co, Ci, C2 and C3 are respectively shown in panels 560, 562, 564 and 566. Evidently, the likelihood functions Ciand C2 in panels 562 and 566 assign relatively lower scores to partitions of the segment 546 associated with either the polynucleotide 108 or the peptide 110. While these evaluations may already provide a reasonable indication of the location of the peptide region, Fig. 4 shows an example wherein the transition matrix is used by a Viterbi algorithm to group the scores in a manner consistent with the sequence of the specific components in the composite polymer 106. Panel 568 shows an output of the Viterbi algorithm illustrating the regions (from left to right) corresponding the polynucleotide 108, the first polymer 112, the peptide 110, and the second polymer 114. In this manner, the peptide region 558 may be located. Returning to Fig. 4, having determined a peptide region, the segment 446 may be further segmented to isolate the peptide region 458, which may then be processed by the classifier algorithm 448 to identify the peptide 450. In this manner, since an isolated peptide region 458 is processed by the classifier algorithm 448, training of the classification model used by the classifier algorithm 448 may involve labelled samples of peptides rather than peptide molecules flanked by first and second polymers. As a result, the training process may be simplified and the classification accuracy may be increased. Although the example described involves applying the HMM segmentation to the signal segment 446, in alternative examples HMM segmentation may be performed directly on the signal 226.
[0107] The examples described above have relied on a machine learning model 236 and / or an HMM to segment the signal 226 for peptide identification. In alternative examples, segmentation may alternatively be performed using statistical techniques. In further examples, the signal 226 may be segmented solely based on a single statistical model. For example, statistical signal processing techniques may rely on the relative uniformity of the first (bridging) polymer 112 and / or the second polymer 114 to locate the corresponding regions within the signal 226. The signal 226 may locate the peptide region 458 as the region of the signal 226 flanked by the statistically uniform measurements corresponding to the polymer regions. For example, the polymer regions may be identified by applying a filtering threshold to rolling averages and rolling standard deviations computed for either the signal 226 or the signal segment 446. The located peptide region 458 may then be processed using the classification algorithm 448 to identify the peptide 450. In further examples, other statistical techniques may be relied upon to isolate the peptide region 458 within the signal 226. In this manner, peptide identification may be achieved without recourse to a machine learning model 236 and / or HMM likelihood functions 454 for signal segmentation. Regardless of the techniques used for signal segmentation, the methods described thus far enable the peptide 110 included in a composite polymer 106 to be characterised. Further, by relying on a polynucleotide barcode 240 included in the composite polymer 106, the segmentation methods described above can be used for training the classification model used by the classification algorithm 248 as described above. The methods described hereinbefore may be adapted for characterising polypeptides, including proteins. For example, multiple composite polymers may be translocated through the nanopore 102, each including a different peptide fragment of a polypeptide to be barcoded. As a result, measurement signals associated with each of the composite polymers may be received by the data processing system 124. The methods described above may thereafter be employed to characterise the peptide fragment included in each composite polymer. As a result, the polypeptide or protein may be characterised on the basis of the collection of peptide fragments identified within the cluster.
[0108] Fig. 6 shows a schematic for a method of polypeptide identification based on an unsupervised machine learning model. The method involves receiving a dataset comprising multiple signals 670. Each of the plurality of signals 670 may have been obtained from the sensor unit 120 during translocation of a respective composite polymer 106. In this manner, a plurality of signals 670 each including measurements of a different composite polymer 106 may be received for processing. The composite polymers corresponding to each measurement may include different peptides 110 and polynucleotides 108. For example, some of the composite polymers may include a target polynucleotide corresponding to a particular sample or a specific individual. The dataset 670 may therefore include a set of target signals 672 corresponding to composite polymers comprising the target polynucleotide. Additionally, the dataset 670 may include signals 674 corresponding to composite polymers that do not comprise the target polynucleotide, instead comprising a different polynucleotide. Still other signals may be included in the dataset.
[0109] Having received the dataset including a plurality of signals 670, the data processing system 124 may perform the methods described above to characterise the peptide included in each signal. However, in order to characterise the peptides associated with the target signals, the data processing system 124 may be tasked with identifying the target signals 672 from amongst all the signals included in the dataset 670. In one approach, the polynucleotide 108 associated with each signal may be identified first, for example, using a machine learning model 236 as described hereinbefore. The target signals 672 may then be identified as the signals for which the target polynucleotide is detected using the methods above. In an alternative example shown in Fig. 6, an unsupervised machine learning model 676 is used to identify the set of target signals 672. In some examples, the model 676 may be trained on unlabelled signals, i.e. signals for which the polynucleotide 108 and / or the peptide 110 has not been determined, to recognise signals including the same polynucleotide 108. For example, the model 676 may be trained to recognise a statistical characteristic or signature of the target polynucleotide based on a priori information.
[0110] By recognising target signals based on such statistical characteristics, the set of target signals 674 may be readily grouped, following which the segmentation and classification methods described thus far in relation to a single composite polymer 106 may be performed separately on each of the signals in the set of target signals 674. This approach of identifying the set of target signals 672 may not rely on processing each of the target signals, thereby saving valuable computational resources.
[0111] In further examples, the model 676 may be trained in dependence on a loss function that includes a contrastive loss term, which penalises classification of signals associated with different polynucleotides into the same grouping. In this manner, the unsupervised machine learning model 676 may learn to recognise any given signal Atand its time-warped versions as belonging to the same cluster. For example, as shown in Fig. 6, target signals At, At+ p and At+ 2P may be placed in the same cluster by the trained model 676. Moreover, all other signals in the dataset 670 may also be readily clustered in this manner. Fig. 6 shows an example output 678 of the unsupervised model 676 presented within a two-dimensional embedding space. The dimensions of the embedding space in this example relate with timings of individual measurements included in the set of signals 670 translocated through the nanopore. As a result, clusters of signals segregated by the unsupervised model 676 can be clearly distinguished in the embedding space. In particular, the output 678 shows clusters in which each point corresponds to signal associated with the same polynucleotide as others within the cluster, i.e. time-warped versions of the same signal are clustered together. For example, the cluster 680 may include the set of target signals 672 and may therefore be associated with the target polynucleotide. As an example, the target polynucleotide may be a polynucleotide barcode. Other clusters may be associated with other polynucleotides. The shading of the points in each cluster of the output 670 is indicative of the peptide identity determined for the corresponding signal. In this manner, the different peptides corresponding to signals associated with the same (target) polynucleotide may be identified. Subsequently, the protein 680 associated with the peptide fragments may be determined. For example, some proteins may be uniquely determined by the given combination of peptide fragments identified within a cluster. In other examples, the protein 680 may be determined up to an equivalence class comprising combinations of the same peptide fragments. In either case, the concentration of a given protein 680 may be determined based on the number of peptide fragments associated with the protein 680 present in the cluster.
[0112] The method described thus far with reference to Fig. 5 allows large datasets including many signals 670 to be readily clustered, with each cluster being subsequently processed for characterising the peptides. This allows large sample sets including signals associated with different target polynucleotides to be processed in parallel. In this manner, the processing efficiency of the methods described is greatly enhanced. For example, this may aid in the preparation of assays for multiple patients simultaneously.
[0113] Fig. 7 shows a flow diagram of the method 700 for characterising a peptide connected to a polynucleotide in a composite polymer. The polynucleotide may be connected to the peptide by a bridging polymer having a predetermined sequence of polymer units. For example, the bridging polymer may be a polythymine molecule include a predetermined sequence of thymine DNA bases. The method 700 includes at step 784 providing the composite polymer to an apparatus and thereafter at step 786 translocating the composite polymer through the nanopore. The nanopore may include a sensor for measuring a current or voltage in the nanopore during translocation of the composite polymer. The method 700 includes at step 788 receiving a signal indicative of measurements of the composite polymer by the sensor unit in the nanopore.
[0114] At step 792 the method includes determining a segment comprising measurements of the peptide by identifying a region of the signal corresponding to the predetermined sequence of polymer units associated with the bridging polymer. The segment may, for example, exclude a majority of measurements of the polynucleotide and thereby isolate a minority of measurements from the signal including measurements of the peptide. In some examples, the method 700 may rely on a statistical model to determine the signal segment comprising measurements of the peptide. Alternatively, the method 700 may employ a machine learning model to determine the signal segment. For example, the machine learning model may determine a read sequence and locate an end of the predetermined sequence of polymer units associated with the bridging polymer within the read sequence. In this manner, the machine learning model may locate a segmentation point for segmenting the signal and determining the segment comprising measurements of the peptide.
[0115] In further examples, the segment may be further segmented using a hidden Markov model (HMM). For example, HMM likelihood functions associated with specific components such as the polynucleotide, bridging polymer and the peptide may be evaluated to determine scores (or probabilities) for a plurality of signal partitions. The scores may, for example, be indicative of the likelihood of each signal partition including measurements of the specific component associated with the likelihood function. In this manner, partitions including measurements of the peptide may be further isolated from within the signal segment to determine a peptide region for further processing. In some examples, the HMM segmentation may be employed without segmenting based on the basecalled read sequence.
[0116] A signal segment obtained using any of the segmentation methods discussed may, in step 792 of the method 700, be processed using a classification algorithm to characterise the peptide. The method 700 may further be used to label proteins associated with composite polymers including the same polynucleotide connected to peptide fragments of the protein to be labelled. For example, the composite polymers associated with a same target polynucleotide may be included in a sample containing many other composite polymers associated with different polynucleotides aimed at alternative assays. Signals obtained by translocating each of the composite polymers may be processed using the methods described above to determine an association between the polynucleotide and different peptides included in each of the composite polymers. In such examples, an unsupervised machine learning model may be used to cluster the composite polymers associated with the same polynucleotide in a same class, thereby grouping the target composite polymers prior to applying the segmentation and classification methods discussed. In this manner, the method 700 enables rapid identification of proteins associated with peptides associated with a given target polynucleotide such as a polynucleotide barcode. The above embodiments are to be understood as illustrative examples of the invention. Further embodiments of the invention are envisaged. For example, deep learning techniques may be employed to perform the signal segmentation, either entirely replacing or supplementing the basecalling segmentation and HMM segmentation methods described. Deep learning methods may be trained using supervised methods relying on human-annotated datasets and / or using semi -supervised methods. It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.
Claims
CLAIMS1. A computer-implemented method comprising: receiving a signal indicative of measurements of a composite polymer by a sensor unit comprising a nanopore, the composite polymer comprising a polynucleotide connected to a peptide; determining, by identifying a region of the composite polymer connecting the polynucleotide to the peptide, a segment of the signal comprising measurements of the peptide; and processing the determined segment of the signal using a classification algorithm to characterise the peptide.
2. The method of claim 1, wherein: the classification algorithm uses a classification model; and characterising the peptide comprises estimating a characteristic of the peptide, the method further comprising updating parameter values of the classification model based on a comparison between the estimated characteristic of the peptide and a ground truth characteristic of the peptide.
3. The method of claim 2, comprising: identifying, using at least part of the signal, the polynucleotide; and determining the ground truth characteristic of the peptide based on the identified polynucleotide.
4. The method of claim 1, comprising identifying, using at least part of the signal, the polynucleotide, whereby to determine an association between the polynucleotide and the peptide.
5. The method of any preceding claim, wherein the polynucleotide is connected to the peptide by a bridging polymer, the bridging polymer being associated with a predetermined sequence of polymer units, and wherein identifying the regionof the composite polymer comprises identifying the predetermined sequence of polymer units.
6. The method of claim 5, wherein the determined segment of the signal is a first segment of the signal, and determining the first segment of the signal comprises: processing the signal using a machine learning model to determine: a read sequence indicative of an estimated sequence of polymer units of the composite polymer; and a mapping that associates positions within the read sequence with corresponding points within the signal; estimating a position within the read sequence corresponding to an end of the predetermined sequence of polymer units; and determining, using the mapping and the estimated position within the read sequence, a segmentation point within the signal, the segmentation point separating the first signal segment and a second signal segment comprising measurements of the polynucleotide.
7. The method of claim 5, wherein identifying the polynucleotide is based on the determined read sequence.
8. The method of claim 5, wherein identifying the polynucleotide comprises basecalling the second signal segment.
9. The method of any of claims 5 to 8, wherein processing the determined segment of the signal comprises: determining a plurality of partitions of the determined segment of the signal; processing the plurality of partitions to compute, for each partition, a respective statistic associated with measurements in the partition; evaluating a likelihood function of a hidden Markov model (HMM), the likelihood function mapping the respective statistic for each partition to arespective score indicative of the partition being associated with a specific component of the composite polymer; locating, based on the respective scores for the plurality of partitions, a peptide region within the determined segment of the signal; and processing the peptide region using the classification algorithm to characterise the peptide.
10. The method of claim 9, wherein: the bridging polymer is a first polymer; the peptide is connected to the first polymer at a first terminus and to a second polymer at a second terminus; the determined segment of the signal comprises measurements of each of the first polymer, the peptide, and the second polymer; the specific component of the composite polymer comprises the first polymer and / or the second polymer; and the peptide region is located based on identifying, within the signal, measurements of the first polymer and / or measurements of the second polymer.
11. The method of either claim 9 or claim 10, wherein: the likelihood function is a first likelihood function of a plurality of likelihood functions of the HMM, each likelihood function of the plurality of likelihood functions being associated with a respective specific component of the composite polymer; and locating the peptide region is based on scores determined by evaluating the plurality of likelihood functions.
12. The method of claim 11, wherein locating the peptide region comprises: processing the scores determined by evaluating the plurality of likelihood functions using a transition matrix to group the scores in dependence on a transition sequence of the associated specific components of the composite polymer; andlocating the peptide region by locating a grouping of scores indicative of the peptide.
13. The method of either claim 11 or claim 12, comprising determining a statistical model to encode rules for evaluating the plurality of likelihood functions, wherein evaluating the plurality of likelihood functions is in dependence on the determined statistical model.
14. The method of any of claims 9 to 13, wherein neighbouring partitions of the plurality of partitions overlap with one another.
15. The method of any of claims 9 to 14, wherein the respective statistic for a given partition is a mean or a variance of measurements included in the given partition.
16. The method of claim 5, wherein: the bridging polymer is a first polymer; the peptide is connected to the first polymer at a first terminus and to a second polymer at a second terminus; the determined segment of the signal comprises measurements of each of the first polymer, the peptide and the second polymer; and processing the determined segment of the signal comprises: locating, using a statistical model, measurements of the first polymer and measurements of the second polymer within the determined segment of the signal; locating, based on the located measurements of the first polymer and the second polymer, a peptide region within the determined segment of the signal; and processing the peptide region using the classification algorithm to characterise the peptide.
17. The method of any of claims 5 to 16, wherein the bridging polymer is a polythymine.
18. The method of any preceding claim, wherein the classification algorithm uses a neural network classifier.
19. The method of any preceding claim, wherein the polynucleotide is DNA.
20. A computer-implement method comprising: receiving a plurality of signals each indicative of measurements of a respective composite polymer by a respective sensor unit comprising a respective nanopore; processing the plurality of signals to identify a set of target signals, the respective composite polymer for each target signal comprising a target polynucleotide; performing, for each target signal, the method of any of claims 1 to 19 to characterise a respective peptide; and characterising a target protein based on the respective peptides.
21. The method of claim 20, wherein processing the plurality of signals to identify the set of target signals uses an unsupervised machine learning model.
22. The method of claim 21, wherein the unsupervised machine learning model is trained to identify the set of target signals based on a contrastive loss function.
23. A data processing system comprising means for carrying out the method of any of claims 1 to 22.
24. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 22.
25. A method comprising:providing a composite polymer, the composite polymer comprising a polynucleotide connected to a peptide; translocating the composite polymer through a nanopore to obtain a signal indicative of measurements of the composite polymer using the nanopore; and performing the method of any of claims 1 to 19 to characterise the peptide.
26. A method comprising: providing a plurality of composite polymers each comprising a respective polynucleotide connected to a respective peptide; translocating the plurality of composite polymers through at least one nanopore to obtain a plurality of measurement signals; and processing the plurality of measurement signals using the computer- implemented method of any of claims 20 to 22 to characterise the target protein.
27. The method of either claim 25 or claim 26, wherein translocation is controlled by a polynucleotide binding protein.
28. The method of claim 27, wherein the polynucleotide binding protein is a helicase.
29. Apparatus comprising: a sensor unit comprising a nanopore; means for translocating a composite polymer through the nanopore, the composite polymer comprising a polynucleotide connected to a peptide; and a data processing system arranged to process measurements of the composite polymer by the sensor unit in accordance with any of claims 1 to 19.
30. The apparatus of claim 29, wherein said means comprise electrodes for providing a potential difference across the nanopore for translocating the polynucleotide through the nanopore.
31. The apparatus of either claim 29 or claim 30, wherein said means comprise a motor enzyme for controlling translocation of the composite polymer through the nanopore.
32. A kit for use in the characterisation of a peptide, the kit comprising the apparatus of any one of claims 29 to 31.
Citation Information
Patent Citations
Sequencing using concatemers of copies of sense and antisense strands
US9910956B2
Suspended carbon nanotube field effect transistor
WO2005124888A1
Deliver of molecules to a li id bila
WO2006100484A2
Formation of lipid bilayers
WO2008102121A1
Formation of layers of amphiphilic molecules
WO2009077734A2