method

By integrating a variable region flanking sequence (VRFS) between barcode and UMI sequences in polynucleotide arrays, the challenges of accurate UMI and barcode identification in long-read sequencing are addressed, leading to improved sequencing accuracy and reduced data loss.

WO2025133135A1PCT designated stage expired Publication Date: 2025-06-26OXFORD UNIVERSITY INNOVATION LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/087930
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The identification of unique molecular identifiers (UMIs) and barcodes in long-read sequencing is challenging due to errors in bead synthesis, PCR, and sequencing, leading to inaccurate assignment and significant data loss.

Method used

Incorporating a first variable region flanking sequence (VRFS) between the barcode and UMI sequences in polynucleotide arrays, which allows for more accurate delimitation of these sequences by aligning reads to a reference sequence containing the VRFS.

Benefits of technology

The use of VRFS enhances the accuracy of UMI and barcode assignment, reducing errors and data loss, and improving the overall resolution and efficiency of single-cell sequencing analyses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000071_0001
    Figure IMGF000071_0001
  • Figure IMGF000072_0001
    Figure IMGF000072_0001
  • Figure 00000101_0000
    Figure 00000101_0000
Patent Text Reader

Abstract

The invention relates to polynucleotides comprising sequences that flank variable regions, such as barcode sequences (BCs) and unique molecular identifier sequences (UMIs). The invention also relates to methods of delimiting the BC and / or the UMI in sequencing reads of said polynucleotides.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD

[0002] Field of the Invention

[0003] The present invention relates to arrays of polynucleotides comprising a first variable region flanking sequence between a barcode and a unique molecular identifier sequence, methods of sequencing the polynucleotides, methods for analysing the sequence data produced by comprising a computer readable media and methods of making the arrays.

[0004] Background of the Invention

[0005] Droplet-based sequencing has rapidly advanced to be the gold standard method in singlecell transcriptomics, leveraging its capability to offer higher throughput. Droplet-based sequencing methods such as Drop-seq (Macosko et al. (2015))1, InDrops (Klein et al. (2015)) and 10X Chromium (Zheng et al. (2017))3heralded a significant advancement in the throughput of single-cell sequencing. While Drop-seq gained traction as a favourable academic approach, 10X Genomics Chromium platform capitalised on this momentum and commercialised droplet-based single-cell sequencing, ultimately securing its position as the market front-runner.

[0006] Central to this technology are polyA mRNA capture beads that are synthesised and contain a PCR primer, a cell barcode and a unique molecular identifier (UMI). Historically, the techniques and strategies underpinning droplet-based sequencing were developed with Illumina sequencing platforms in mind. Typically, the Illumina sequencing protocol employs a fragmentation step so that sequencing of Read 1 contains both the barcode and the UMI sequence, while, Read 2 contains the sequencing information for either the 3’ or 5’ end of the captured transcript.

[0007] Recently, the emergence of long-read sequencing, which allows for complete end-to-end sequencing, has made it more challenging to identify barcode and UMI sequences. This challenge mainly stems from the reliance on computational pattern matching to locate the primer site before these sequences, ensuring accurate barcode and UMI assignment. PCR and sequencing errors further complicate this, making pattern matching less accurate and leading to a significant proportion of reads being discarded.

[0008] It is an aim of the present invention to improve the accuracy of barcode and UMI assignment.

[0009] Summary of the Invention

[0010] The present inventors surprisingly found that errors in bead synthesis, PCR and sequencing have made the identification of UMIs and barcodes significantly more difficult. Indeed, the examples demonstrate that sequence information produced using a variety of different methods (for example 10X Chromium or single cell nanopore sequencing) contained a number of errors in the UMI and barcode region. In particular, the inventors determined that for UMIs near the 3 ’ end of the polynucleotide attached to the bead, the UMI predicted based on the sequencing information tended to include a larger than normal proportion of T nucleotides as the 3’ base. Based on this information, it was hypothesised that errors, in particular bead synthesis errors, had shortened the UMI and that the computational methods for detecting the UMI had therefore assumed that one of the two bases corresponding to the reverse complement of the polyA tail was part of the UMI. The examples also demonstrate that the errors caused by bead synthesis, PCR and sequencing could be reduced by introducing at least one variable region flanking sequence (VRFS). Where VRFS(s) are used, a user may delimit the BC and / or UMI in a sequence read by aligning it to a reference sequence that comprises the VRFS.

[0011] Accordingly, the invention provides an array of polynucleotides, wherein each polynucleotide of the array comprises (a) a PCR handle sequence; (b) a barcode sequence (BC); (c) a unique molecular identifier sequence (UMI); and (d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI.

[0012] The invention also provides a micro-particle comprising a micro-bead and an array of polynucleotides of the invention, wherein each polynucleotide is bound to the micro- bead. The invention further provides a plurality of micro-particles of the invention, wherein the BC sequence of each polynucleotide of each micro-particle is the same as the BC sequence of essentially each other polynucleotide of the same micro-particle, and different from the BC sequence of the polynucleotides of each other micro-particle.

[0013] The invention additionally provides a method for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises (a) a PCR handle sequence, (b) a barcode sequence (BC), (c) a unique molecular identifier sequence (UMI), and (d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI, and wherein the method comprises (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS, and (ii) predicting the positions of the BC and / or the UMI of the read and / or its reverse complement based on the alignment.

[0014] The invention also provides a method for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises (a) a PCR handle sequence, (b) a barcode sequence (BC), (c) a unique molecular identifier sequence (UMI), and (d) a second variable region flanking sequence (VRFS), wherein the second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI, and wherein the method comprises (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and (ii) predicting the positions of the BC and / or the UMI based on the alignment. The invention further provides a method for delimiting barcode sequences and / or unique molecular identifier sequences in a library of reads, each read corresponding to the sequence of at least a portion of a polynucleotide, wherein each polynucleotide or its reverse complement comprises (a) a PCR handle sequence, (b) a barcode sequence (BC); (c) a unique molecular identifier sequence (UMI), (d) a first variable region flanking sequence (VRFS) between the BC and the UMI and / or a second VRFS located 3’ of the portion of the polynucleotide comprising the BC and the UMI and / or between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI; and (e) optionally, an analyte capture sequence, wherein the method comprises, for each read, performing the method of the invention for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide.

[0015] The invention also provides a computer readable medium comprising instructions for performing any one of the methods of the invention. The invention also provides a computer comprising the computer readable medium of the invention.

[0016] The invention also provides a method of making the array of polynucleotides of the invention, wherein the method comprises using split-and-pool polynucleotide synthesis to make the BC and / or using degenerate polynucleotide synthesis to make the UMI.

[0017] The invention further provides a method of producing a library of polynucleotides, wherein the polynucleotides are amplified from the polynucleotides of a sample and / or tag non-polynucleotide analytes of a sample, the method comprising (a) capturing analytes in the sample on an array of polynucleotides, a micro-particle or a plurality of micro-particles according to the invention, (b) generating copies of the array of polynucleotides, including any (I) polynucleotides of the sample or (II) polynucleotides that tag non-polynucleotide analytes of the sample, captured by the array polynucleotides; and (c) amplifying the number of copies of each polynucleotide to produce a library of polynucleotides amplified from or tagging analytes in the sample. The invention also provides a library of polynucleotides obtained or obtainable by the method of the invention of producing a library of polynucleotides.

[0018] The invention also provides a method of generating a library of reads, comprising sequencing the library of polynucleotide produced by method of the invention of producing a library of polynucleotides, or the library of polynucleotides of the invention.

[0019] Brief Description of the Figures

[0020] Figure 1 - Significant data loss in single-cell sequencing analysis. a, A comparative schematic illustrates the bead designs of both Drop-seq and 10X Chromium platform v3.1. While both designs incorporate a barcode and UMI, they vary in length, the positioning of the V base (A, C, or G), and PCR primer sequences. N signifies a random base, J denotes a semi-random base, and T, C, G and A represent Thymidine, Cytosine, Guanine, Adenosine, respectively, b, Data from multiple public datasets using the 10X Chromium v3.1 was analyzed. The percentage of Thymidine at each position in read 1 of the Illumina sequencing data was plotted, c, The percent of A, C, G and T base using data for the 5k Human PBMC 3’ v3.1 10X Chromium was plotted for each base within the fastq upstream of the end of the PCR primer position, d, Drop-seq data downloaded from GSE63473 was analysed. The percentage of each base in read 1 was plotted, e, A mixed species Dropseq single-cell experiment sequenced on a PromethlON™ device was analysed. The percent of A, C, G and T base was plotted for each base upstream of the end of the PCR primer position.

[0021] Figure 2 - Base proportions in 10X chromium and Drop-seq sequences across shortread and ONT long-read platforms. a, Data from multiple public datasets using the 10X Chromium v3.1 was analyzed. Length of each oligonucleotide was measured using ONT sequencing, b, Data from multiple public datasets using Dropseq was analyzed. Length of each oligonucleotide was measured using ONT sequencing, c, d, The frequency of UMIs captured within the 10X (left panel) and the Dropseq (right panel) datasets.

[0022] Figure 3 - Inclusion of a VRFS enhances the ability to identify the start of the CMI sequence. a. Schematic representation comparing two CMI selection methods: positional and ‘spacer’. The positional method identifies the PCR primer's end through local alignment and selects a site 16 base pairs downstream. In contrast, the ‘spacer’ method utilizes a VRFS to pinpoint the CMI's start position. As used in this context, the term ‘spacer’ may be used interchangeably with ‘VRFS’ to refer to a structural feature flanking the BC and UMI. b. Percentage of CMIs perfectly matching the anticipated CMI length. Data analysed using Unpaired t-test assuming both populations have the same standard deviation. ** P<0.01

[0023] Figure 4 - The inclusion of a VRFS in the mRNA capture beads improves the UMI recovery. a, A schematic showing the scCOLOR-seq v2 bead design that now incorporates a VRFS between the barcode and UMI, in addition to a homodimer UMI followed by a limited variable V base between the poly(dT) capture region, b, The percent of A, C, G and T base plotted for each base within the fastq file upstream of the PCR primer end position. c, We use the homodimer repeats as a proxy for measuring improved UMI selection. The percent of perfect homodimer repeats across the full length of the UMI is plotted using the positional and the ‘spacer’ approach.

[0024] Figure 5 - Including a VRFS into the beads increases the number of UMIs and transcripts detected. a, The log 10 density of UMIs per cell comparing the positional and ‘spacer’ UMI selection methods, b, The loglO density of transcripts per cell for both positional and ‘spacer’ UMI selection approaches, c, Correlation of transcripts in human-origin cells between the positional and ‘spacer’ approaches. Transcripts with enhanced detection using the spacer method are circled, d, Correlation of mouse-origin transcripts between the positional and ‘spacer’ approach, with the circle identifying transcripts more readily detected using the ‘spacer’ UMI identification approach, e, Expression profile of the transcript ENSMUST00000034270 for the positional and ‘spacer’ approach, f, UMAP plots showing the increased expression of ENSMUST00000034270 using the ‘spacer’ approach. For plots c-f, each dot represents a single cell.

[0025] Detailed Description

[0026] General Definitions

[0027] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person skilled in the art to which this invention belongs.

[0028] In general, the term “comprising” is intended to mean including but not limited to. For example, the phrase “each polynucleotide of the array comprises a PCR handle sequence ” should be interpreted to mean that each polynucleotide of the array comprises the PCR handle sequence, but that each polynucleotide may comprise further components (for example a barcode sequence, a unique molecular identifier sequence, etc.,).

[0029] In some embodiments of the invention, the word “comprising” is replaced with the phrase “consisting of” . The term “consisting of” is intended to be limiting. For example, the phrase “each polynucleotide of the array consists of a PCR handle sequence" should be understood to mean that each polynucleotide of the array contains a PCR handle sequence and no further components.

[0030] In some embodiments of the invention, the word “comprising” is replaced with the phrase “consisting essentially of” . The term “consisting essentially of” means that specific further components can be present, namely those not materially affecting the essential characteristics of the subject matter. For example, the phrase “each polynucleotide of the array consists essentially of a PCR handle sequence ” indicates that each polynucleotide of the array may further comprise one or more nucleotides that have no particular function.

[0031] The term “about” or “around” when referring to a value refers to that value but within a reasonable degree of scientific error. Optionally, a value is “about X” or “around X” if it is within 10%, within 5% or within 1% of X.

[0032] The singular forms “a”, “an” and “the ” include plural references unless the context clearly dictates otherwise.

[0033] All publications, patents and patent applications cited herein, whether Supra or Infra, are hereby incorporated by reference in their entirety.

[0034] Polynucleotides

[0035] The present invention relates to polynucleotides and arrays of polynucleotides.

[0036] The terms “polynucleotide", “oligonucleotide” or “oligo” may in some cases be used herein interchangeably, and refer to a string of nucleotide monomers in a chain typically linked by phosphodiester bonds. As used herein, a polynucleotide may be a chain of nucleotides of any length, whilst an oligonucleotide typically comprises up to 50 nucleotides.

[0037] Polynucleotides have a chemical orientation defined by the position of the linking carbon in the five-carbon sugar of each consecutive nucleotide in the chain. Polynucleotides may be manufactured by the addition of nucleotides at either the 5’ end (manufacture in a 5’ direction) or the 3’ end (manufacture in a 3’ direction) to elongate the chain.

[0038] Likewise, sequence elements along the length of a polynucleotide have a sequential order defined by the directionality of the chain of nucleotides that is either 5’ to 3’ or 3’ to 5’. Polynucleotides may comprise any combination of natural or canonical nucleotides (z.e., “naturally occurring" or “natural” nucleotides), which include adenosine, guanosine, cytidine, thymidine and uridine. The polynucleotides may also comprise nucleotide analogues. For example, the polynucleotide may include one or more peptide nucleotides, in which the phosphate linkage found in DNA and RNA is replaced by a peptide-like A-(2-aminoethyl)glycine. Peptide nucleotides undergo normal Watson- Crick base pairing and hybridize to complementary DNA / RNA with higher affinity and specificity and lower salt-dependency than normal DNA / RNA oligonucleotides and may have increased stability. The polynucleotide may include one or more locked nucleotides (LNA), which comprise a 2'-(9-4'-C-methylene bridge and are conformationally restricted. LNA form stable hybrid duplexes with DNA and RNA with increased stability and higher hybrid duplex melting temperatures. The polynucleotide may include one or more Propynyl dU (also known as pdU-CE Phosphoramidite, or 5'- Dimethoxytrityl-5-(l-Propynyl)-2'-deoxyUridine,3'-[(2-cyanoethyl)-(N,N-diisopropyl)]- phosphoramidite). The polynucleotide may include one or more unlocked nucleotides (UNA), which are analogues of ribonucleotides in which the C2'-C3' bond has been cleaved. UNA form hybrid duplexes with DNA and RNA, but with decreased stability and lower hybrid duplex melting temperatures. LNA and UNA may therefore be used to finely adjust the thermodynamic properties the polynucleotides in which they are incorporated. The polynucleotide may include one or more triazole-linked DNA oligonucleotides, in which one or more of the natural phosphate backbone linkages are replaced with triazole linkages, particularly when click chemistry is used for synthesising the polynucleotide. The polynucleotide may include one or more 2’-O-methoxy-ethyl bases (2’ -MOE), such as 2-Methoxyethoxy A, 2-Methoxyethoxy MeC, 2-Methoxyethoxy G and / or 2-Methoxyethoxy T. The polynucleotide may include one or more 2'-O- Methyl RNA bases. The polynucleotide may include one or more 2’-fluoro bases, such as fluoro C, fluoro U, fluoro A, and / or fluoro G. Other specific examples of nucleotide analogues include 2-Aminopurine, 5-Bromo dU, deoxyUridine, 2,6-Diaminopurine (2- Amino-dA), Dideoxy-C, deoxyinosine, Hydroxymethyl dC, Inverted dT, Iso-dG, Iso-dC, 5-Methyl dC, 5 -Nitroindole, 5-hydroxybutynl-2’-deoxyuridine (Super T) and 8-aza-7- deazaguanosine (Super G). In some cases the polynucleotide may include super T 2,6- Diaminopurine (2-Amino-dA) and / or 5-Methyl dC. An Array of Polynucleotides

[0039] The present invention relates to arrays of polynucleotides. An array of polynucleotides is a collection of polynucleotides that share common features (for example comprise a PCR handle sequence, a barcode sequence, a unique molecular identifier sequence and a first variable region flanking sequence). In this case, each polynucleotide of the array comprises:

[0040] (a) a PCR handle sequence;

[0041] (b) a barcode sequence (BC);

[0042] (c) a unique molecular identifier sequence (UMI); and

[0043] (d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI. Optionally, each polynucleotide of the array further comprises (e) an analyte capture sequence. Optionally, each polynucleotide of the array further comprises a second VRFS.

[0044] The PCR handle is typically 5’ of the BC, UMI, first VRFS, the optional analyte capture sequence and the optional second VRFS. This is to ensure that all of these features, plus any analyte that is captured / hybridised / ligated to the polynucleotide is amplified when PCR is performed. Typically, the analyte capture sequence is 3’ of the second VRFS where present, or otherwise 3’ of the portion of the polynucleotide comprising the PCR handle sequence, the BC, the UMI and the first VRFS. For example, each polynucleotide of the array may comprise in a 5’ to 3’ direction:

[0045] (a) the PCR handle sequence;

[0046] (b) the BC or the UMI;

[0047] (c) the first VRFS; and

[0048] (d) the BC or the UMI.

[0049] Each polynucleotide of the array may comprise in a 5’ to 3’ direction:

[0050] (a) the PCR handle sequence;

[0051] (b) the BC or the UMI;

[0052] (c) the first VRFS; (d) the BC or the UMI; and

[0053] (e) the second VRFS.

[0054] Each polynucleotide of the array may comprise in a 5’ to 3’ direction:

[0055] (a) the PCR handle sequence;

[0056] (b) the BC or the UMI;

[0057] (c) the first VRFS;

[0058] (d) the BC or the UMI;

[0059] (e) the second VRFS; and

[0060] (f) the analyte capture region.

[0061] Each polynucleotide of the array may comprise in a 5’ to 3’ direction:

[0062] (a) the PCR handle sequence;

[0063] (b) the BC or the UMI;

[0064] (c) the first VRFS;

[0065] (d) the BC or the UMI; and

[0066] (e) the analyte capture region.

[0067] The BC or the UMI are generally included only once in the examples provided above and may be in either order. For example, the BC may be 5’ of the UMI, or the UMI may be 5’ of the BC. Typically, the BC is 5’ of the UMI, for example, each polynucleotide of the array may comprise in a 5’ to 3’ direction:

[0068] (a) the PCR handle sequence;

[0069] (b) the BC;

[0070] (c) the first VRFS;

[0071] (d) the UMI;

[0072] (e) the optional second VRFS; and

[0073] (f) the optional analyte capture sequence.

[0074] The polynucleotides of the array are typically DNA (single-stranded DNA) but could also be RNA. The array may comprise any number of polynucleotides. The array may comprise at least 1,000 polynucleotides, at least 10,000 polynucleotides, at least 100,000 polynucleotides, at least IxlO6, at least IxlO7, at least IxlO8, at least IxlO9, or at least IxlO10polynucleotides. It will be appreciated that there is no real upper limit to the number of polynucleotides that could be included in the array. In some cases, the array may comprise IxlO3to IxlO12polynucleotides, such as IxlO5to IxlO11polynucleotides.

[0075] Typically, the polynucleotides in the array of the invention are at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90 nucleotides in length. For example, the polynucleotides in the array of the invention may be 30 to 300 polynucleotides in length, such as 40 to 250, 50 to 200, 70 to 150, or 80 to 120 nucleotides in length.

[0076] The polynucleotides in the array are typically DNA.

[0077] In addition to the features described herein, the polynucleotides of the arrays of the present invention may in some cases comprise any other suitable feature described in PCT Publication Nos: WO 2021 / 229230, WO 2022 / 118027, or WO 2023 / 194714, or in PCT Application No: PCT / GB2023 / 052028, each of which is specifically incorporated by reference herein.

[0078] In some cases, the polynucleotides of the array do not dimerise with each other, i.e. do not form dimers with other polynucleotides of the array, for example under conditions of lx phosphate buffered saline, pH 7.4, 25 °C. In some cases, each polynucleotide of the array does not form a hairpin or dimerise, for example under conditions of lx phosphate buffered saline, pH 7.4, 25 °C. It is advantageous for the polynucleotides of the array to not dimerise nor form hairpins, as such structures could negatively influence the capture of analyte molecules, the amplification of the polynucleotides by PCR and the sequencing of the amplified polynucleotides. In some cases, a proportion of the polynucleotides of the array may comprise a 3’ hairpin sequence, such as those discussed in PCT application no. PCT / GB2023 / 052028 and in the section entitled “hairpin” below. Typically, none of the polynucleotides of the array comprise a hairpin sequence, such as a 3’ hairpin sequence, i.e. each polynucleotide of the array does not comprise a hairpin sequence, such as a 3’ hairpin sequence.

[0079] PCR Handle Sequences

[0080] Each polynucleotide of the array comprises a PCR handle sequence. The methods of the invention are performed on polynucleotides comprising a PCR handle sequence.

[0081] A PCR handle sequence hybridizes to PCR oligonucleotide primers during a PCR reaction. Typically, a PCR handle sequence may be at least 15, 16, 17, 18, 19, or 20 nucleotides in length and / or up to 21, 22, 23, 24, 25, 30, or 35 nucleotides in length, for example about 15 to 30, or 18 to 25 nucleotides in length. In some cases, the PCR handle sequence(s) may comprise one or more nucleotide analogues, such as analogues described herein, that form double-stranded hybrids with higher stability than natural nucleotides. In this case, the PCR handle sequence(s) could be shorter, such as at least 3, 4, 5, 6, 8, 9, 10, 11, 12, 13 or 14 nucleotides in length, provided that the PCR handle sequence(s) was capable of hybridizing to PCR oligo as described herein. Hence, a PCR handle sequence(s) may include nucleotide analogues as described herein.

[0082] Suitable PCR handle sequences are discussed in Example 7. In some cases, one or more of the PCR handle sequences may comprise or consist of one of the following sequences: 5’- GTGGTATCAACGCAGAGTAC-3’; 5’-GTCCGAGCGTAGGTTATCCG-3’. In some cases, one or more of the PCR handle sequences may comprise or consist of one of the following sequences: 5’- GTGGTATCAACGCAGAGTAC-3’; 5’- GTCCGAGCGT AGGTT ATCCG-3 ’ , 5 ’ - AAGC AGTGGTATC AACGC AGAGTAC-3 ’ or 5’-TACACGACGCTCTTCCGATCT-3’. The PCR handle sequences may be compatible with standard Illumina sequencing to eliminate the need for custom read 1 sequencing primers to read BC / UMI.

[0083] In some embodiments, the array comprises fewer than 10, fewer than 8, or fewer than 5 different PCR handle sequences. In some embodiments, essentially each polynucleotide of the array comprises the same PCR handle sequence. The phrase “ essentially each polynucleotide of the array comprises the same PCR handle sequence" may be interpreted to mean that some polynucleotides within the array may comprise a different PCR handle sequence due to e.g. errors in synthesis. However, these different sequences are considered for all intents and purposes to be the same PCR handle as they were designed to comprise identical sequences. Other similar phrases such as “essentially each polynucleotide of the array comprises the same BC’ should be interpreted in the same way.

[0084] Optionally, the polynucleotides of the array do not comprise a PCR handle. Thus, further embodiments are contemplated in which the PCR handle sequence is not present. For example, the present invention provides an array of polynucleotides, wherein each polynucleotide of the array comprises:

[0085] (a) a barcode sequence (BC);

[0086] (b) a unique molecular identifier sequence (UMI); and

[0087] (c) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI.

[0088] Barcode and Unique Molecular Identifier Sequences

[0089] Each polynucleotide of the array comprises a barcode sequence (BC) and a unique molecular identifier sequence (UMI). BCs and UMIs are variable regions of the polynucleotide.

[0090] The polynucleotides of the array comprise a barcode sequence (BC). The methods of the invention are performed on polynucleotides which comprise a BC. Barcode sequences are well understood in the art and are typically used to tag / index multiple analytes from a shared source. The present invention initially, or partly, uses BCs in the conventional way during sample analyte library generation. An array of polynucleotides (e.g. on a micro-bead as described herein) may be used to capture sample analyte to generate a library. Each of the polynucleotides may comprise the same or essentially the same BC, or one of the same limited group of BCs. The array of polynucleotides may be conveniently compartmentalised together with one or more discrete samples (for example single cells). Analytes from the sample may be subsequently captured by the polynucleotides and tagged with the barcode. Hence, in cases where single samples (e.g. single cells) are co-compartmentalised with single micro-particles, or the capture polynucleotides thereof, analytes from the same sample can be identified by a shared BC of the polynucleotide that captured the analyte.

[0091] The polynucleotides of the array comprise a unique molecular identifier sequence (UMI). UMIs are typically added to oligonucleotides in micro-particles such as those described herein using degenerative synthesis, as known in the art. In some cases, a different UMI is added to each polynucleotide in the sample to which the UMI is added. Generally, if the polynucleotides are used to capture analytes for sequencing, the user will capture the analytes on the polynucleotides and then amplify the polynucleotides comprising the captured analytes. The amplification step generates multiple copies of each polynucleotide comprising a captured analyte. In which case it is useful if each analyte is captured by a polynucleotide comprising a different UMI as this allows the user to differentiate sequence reads corresponding to multiple copies of the same analyte from sequence reads corresponding to analytes that are similar but not identical in sequence. However, the same UMI may be added to more than one polynucleotide / analyte. This is because once an analyte, such as an mRNA molecule from a cell, has been captured by the polynucleotide of the array and amplified, the UMI can be used in combination with the identity of the analyte to identify sequence reads with the same UMI as being derived from different sources or the same source. As such, it is possible that the diversity of the UMI in the array is similar to, or lower than the number of analyte molecules typically captured by the array, as the diversity of the analyte molecules in combination with the diversity of the UMI minimises the chances that analyte molecules with identical sequence are captured by polynucleotides with identical UMIs. Even if this occurs, for example due to the abundance of analyte molecules of identical sequence, this has minimal effect on the counts of unique analyte molecules comprising the sequence (typically, less than 10%, more typically less than 5%, 4%, 3%, 2%, or 1%). In some cases, where the number of analytes in a sample captured by the array is lower than the number of polynucleotides in the array, the diversity of the UMIs in the array may be designed such that essentially every analyte captured by the array receives a different UMI, but the same UMI sequence may be added to two or more polynucleotides in the array. In this scenario, the user may design the UMIs such that there is a sufficiently high number of different UMIs to ensure that it is unlikely that two analytes which are similar in sequence will be captured by polynucleotides comprising the same UMI. Where polynucleotides having the same UMI sequence each capture an analyte, these may be further distinguished based on their BC and the identity of the captured analyte. Accordingly, in some cases the diversity of the UMIs in the array may be desired such that there is 0.01 to 100 unique UMIs for every analyte molecule captured by the array, such as 0.05 to 20, 0.1 to 10, 1 to 100, 1 to 10, or 0.1 to 1 unique UMIs for every analyte molecule captured by the array. Hence, the diversity of UMIs needed depends on the experiment. In some cases, a different UMI sequence is added to each of at least 1000, 4000, 1.6xl04, 6.4xl04, 2.6xl05, IxlO6, 4xl06, or 1.6xl07different polynucleotides in the array.

[0092] BCs and UMIs have conventionally consisted of a string of single nucleotides, generated randomly, for example using split-and-pool synthesis or degenerate polynucleotide synthesis, as described further elsewhere herein. In other cases, in accordance with the invention, the BC and UMIs may have the features described in WO 2022 / 118027, for example in the claims of that application. Instead of consisting of a string of single nucleotides, the BC and / or UMI, may comprise a series of discrete nucleotide blocks. Typically, the blocks are two or three nucleotides in length, i.e. are dimers or trimers. Longer blocks can also be used. A BC and / or UMI could also be built up from multiple blocks of different sizes. Typically, however, each block in the same or an equivalent position in each BC and / or UMI is the same length. Each discrete nucleotide block has one of a pool of known sequences. A UMI and / or BC, e.g. of an individual polypeptide, may comprise or consist of any combination of the same or different nucleotide blocks within the limitations described herein. However, the ability to differentiate between BCs or between UMIs is improved by using a pre-determined and limited pool of different nucleotide block sequences, either across the full sequence of the BC or UMI, or at each same or equivalent position of the sequence across different polynucleotides. The sequence of each block can be compared across different polynucleotides, i.e. because they are at the same known or otherwise identifiable position in each polynucleotide. These pre-defined nucleotide block sequences may be referred to herein as a nucleotide block pool or nucleotide block sequence pool. All of the nucleotide block sequences within the pre-defined nucleotide block pool typically differ from every other nucleotide block sequence within the pool by at least two, or at least three nucleotide substitutions. Thus, each nucleotide block in the pool differs from each other nucleotide block in the pool by a Hamming distance of at least two or at least three. For example, where dimer nucleotide blocks are used, wherein each nucleotide is selected from a, c, g and t, the pre-defined nucleotide block pool may only comprise four different nucleotide blocks whilst ensuring that each nucleotide block in the pool differs from each other block in the pool by at least two nucleotide substitutions.

[0093] The sequence of each of the nucleotide blocks is otherwise not particularly limited unless otherwise provided herein and provided that the different nucleotide block sequences can be distinguished when sequenced.

[0094] A typical BC may comprise or consist of at least 4, 5, 6, 7, 8, 9, 10, 11 or 12 nucleotides or nucleotide blocks. For example, a BC may comprise or consist of about 4 to 20, 6 to 18, 8 to 16, 10 to 14, or 11 to 13 nucleotides or nucleotide blocks. This may be, for example, where each nucleotide block pool comprises four different sequences (e.g., where each nucleotide block consists of one or two natural nucleotides). A BC may consist of about 4 to 9, 5 to 8, or 6 to 7 nucleotide blocks where each nucleotide block pool comprises twelve different sequences (for example, where each nucleotide block consists of three natural nucleotides). The BC is typically synthesised using a number of rounds of split-and-pool synthesis corresponding to the number of nucleotides or nucleotide blocks in the BC.

[0095] A typical UMI may comprise or consist of at least 4, 5, 6, 7, 8, 9, 10, 11 or 12 nucleotides or nucleotide blocks. For example, a UMI may comprise or consist of about 4 to 20, 4 to 16, 5 to 14, 6 to 10, 8 to 12, or 7 to 9 nucleotides or nucleotide blocks. This may be, for example, where each nucleotide block pool comprises four different sequences (e.g., where each nucleotide block consists of one or two natural nucleotides). A UMI may consist of about 4 to 9, 5 to 8, or 6 to 7 nucleotide blocks where each nucleotide block pool comprises twelve different sequences (for example, where each nucleotide block consists of three natural nucleotides). The UMI is typically synthesised using a number of rounds of degenerative nucleotide synthesis corresponding to the number of nucleotides or nucleotide blocks in the UMI.

[0096] A UMI or BC may comprise at least 4, more typically at least 5, 6, 7, or 8 sequence units (where a sequence unit is either a single nucleotide or a nucleotide block as described herein) and up to 12, 13, 14, 15, or 16 sequence units or more, more typically 6 to 18, 7 to 17, 8 to 16, or 8 or 12 sequence units. 8 to 12 sequence units are most typical for UMIs and 12 to 16 sequence units are most typical for BCs. The sequence units are added to, or present in, the relative polynucleotide at successive, consecutive or non-consecutive unit or block positions. The BCs and UMIs are defined by the sequence of the unit or block at each position. The total diversity of possible BCs or UMIs is determined by (i) the number of sequence units included in the BC or UMI; and (ii) the number of different units or nucleotide block sequences included in the pool of sequences that can be used at each unit / block position. If the same pool is used at every position, then the total possible diversity of sequences is equal to [the size of the pool]A[the number of sequence units or nucleotide blocks in the BC or UMI], In some cases, the total diversity of possible BCs or UMIs is at least 10, or at least 20, 50, 100, 200, or 500 times in excess of the number of analytes in a sample. In other cases, the total diversity of possible BC or UMIs is exceeded by at least 10 times by the number of analytes in sample. For example, 10 to 12 sequence units, with four possible options for each unit, provides about le+6 to 1.6e+7 possible unique BCs or UMIs. Hence, the chances of two identical sample RNA transcripts getting an identical pairing of BCs and / or UMIs becomes vanishingly small. In other cases, such as in single-cell mRNA sequence experiments, where the experiment is designed to capture the transcripts from many cells, but with each cell labelled with a different barcode, the combinatorial diversity generated by the UMI and the BC may be sufficient to distinguish the source of the sequence read (i.e. to associate sequencing reads derived from the same cell and / or the same mRNA molecule by PCR amplification). In this case, a UMI may comprise for example, 8 nucleotides or nucleotide blocks, resulting in a maximum total diversity of 65,536 combinations. A single mammalian cell typically comprises IxlO5to IxlO6mRNAs, however, not all of these will be captured. A typical single cell mRNA sequencing experiment may capture around 30,000 mRNA molecules. Arrays used in single-cell sequence applications may have up to around 2xlO10polynucleotides forming an array on a micro-bead and associated with a single cell and thus have high redundancy in the UMIs present when, for example, a total diversity of 65,536 UMIs are used. In practice, even where mRNA molecules of identical sequence are captured, it is uncommon for these to be captured by polynucleotides with the same UMI. Whilst this may occur more often for abundant mRNA transcripts, this is unlikely to have any significant impact on the overall experiment as it will typically alter the count of mRNA molecules by up to 1-2 %. Accordingly, the diversity in the mRNA transcripts in combination with the diversity in the UMIs is sufficient to distinguish the source of the sequencing read to an acceptable degree (i.e. to associate sequencing reads derived from the same cell and / or the same mRNA molecule by PCR amplification).

[0097] Hence, the number of sequence units chosen will be influenced by the total diversity or number of different possible identifier sequences that are needed for a particular purpose.

[0098] In some cases, nucleotides or nucleotide blocks may be used consecutively to form a single longer nucleotide block corresponding to the full identifier sequence, i.e. a series of consecutive nucleotides or nucleotide blocks. The term “consecutive ” is used to refer to sequential nucleotides or nucleotides blocks in a polynucleotide which immediately follow the previous sequence unit or nucleotide block without intervening nucleotides. For example, an identifier sequence comprising 6 nucleotide blocks, wherein the nucleotide blocks are selected from di -adenosine, di-guanosine, di-cytidine, di-thymidine and di -uridine may have the sequence “AAGGCCTTAAGG”.

[0099] In other cases, one or more spacers or other sequence elements may be included between the nucleotides or nucleotide blocks that make up the identifier sequence. The identifier sequence or the region of the polynucleotide containing all of the nucleotides or nucleotide blocks of the identifier sequence, optionally with other intervening sequence elements, may in some cases be up to 20, 22, 24, 26, 28, 30, 25, 40, 45, 50, 70, 100, or 200 nucleotides in length.

[0100] In some cases, one or more or each of the nucleotide block in a pool consists of two or more of the same nucleotide. For example, the nucleotide block pool may in some cases comprise or consist of the blocks AA (di-adenosine), TT (di-thymidine), GG (diguanosine) and CC (di-cytidine) (or UU (di-uridine)) or / or the blocks AAA (triadenosine, TTT (tri-thymidine), GGG (tri-guanosine) and CCC (tri-cytidine) (or UUU (tri-uridine)); or the blocks AAA, TTT, CCX and GGY, wherein X is A, T or G, and Y is C, A or T; or any combination thereof. In other cases one or more or each nucleotide block sequence may comprise a duplicate or triplicate or other multiple of a nucleotide analogue. Thus, BC and / or UMI may comprise or consist of homo-dimer nucleotide blocks. The BC and / or UMI may comprise or consist of homo-trimer nucleotide blocks. In other cases, such as where sequencing methods are used that have difficulty distinguishing between runs of the same nucleotide, the BC and / or UMI may comprise or consist of nucleotide blocks that do not have two consecutive nucleotides that are the same, e.g. the nucleotide blocks may be hetero-dimers or hetero-trimers.

[0101] The term “nucleotide ” as used herein, particularly in relation to the nucleotide blocks of an identifier sequence, may refer to natural or canonical nucleotides (z.e., “naturally occurring” or “natural” nucleotides), which include adenosine, guanosine, cytidine, thymidine and uridine, or to a non-canonical nucleotide or a nucleotide analogue. The polynucleotides or nucleotide blocks described herein may comprise any combination of natural nucleotides. Alternatively, any non-canonical nucleotides or nucleotide analogues may appear in one or more identifier sequence.

[0102] Variable Region Flanking Sequences

[0103] The polynucleotides of the array comprise at least one variable flanking sequence (VRFS). The polynucleotides of the array may further comprise a second VRFS.

[0104] A VRFS is a nucleotide sequence that flanks (is immediately 5’ or 3’ of) a variable region (such as a BC or a UMI). A VRFS is the same (comprises the same nucleotides) in a substantial proportion (for example in at least 10%, at least 25%, at least 50%, at least 90%, at least 99% or in substantially all) of the polynucleotides of the array. The term "variable region” refers to the sequence elements that are designed to vary to enable the source of a sequencing read to be distinguished from sequencing reads derived from other sources. For example, the BC and the UMI described above are examples of “variable regions” .

[0105] The purpose of the VRFS is to act as an identifiable marker to delimit the variable regions. As described elsewhere herein, it has been identified that BC and UMI from sequencing reads can be incorrectly determined due to errors in the original synthesis of the BC and / or UMI, errors in PCR amplification and errors in sequencing. This can lead to misassignment of reads to cells, UMI inflation and / or the discarding of sequencing reads. This results in lower resolution data obtained from experiments, for example due to the loss of data from a sample derived from a single cell. The use of VRFSs reduces these errors, and thus enables greater resolution experiments with lower wastage of resources such as reverse transcription, PCR and sequencing reagents.

[0106] Each polynucleotide of the array comprises a first VRFS between the UMI and the BC. For example, each polynucleotide of the array may comprise a first VRFS 3’ of the BC and 5’ of the UMI. Alternatively, each polynucleotide of the array may comprise a first VRFS 5’ of the BC and 3’ of the UMI. For example, each polynucleotide of the array may comprise in the direction from 5’ to 3’ : a PCR handle, a BC, a first VRFS and a UMI. Each polynucleotide of the array may comprise in the direction from 5’ to 3’ a PCR handle, a UMI, a first VRFS and a BC.

[0107] Each polynucleotide of the array may comprise a second VRFS. The second VRFS is typically located 3’ of the portion of each polynucleotide of the array that comprises the BC and the UMI, i.e. both the BC and the UMI are upstream of the second VRFS. For example, each polynucleotide of the array may comprise in the direction from 5’ to 3’ : a PCR handle, a BC, a first VRFS, a UMI and a second VRFS. Each polynucleotide of the array may comprise in the direction from 5’ to 3’ a PCR handle, a UMI, a first VRFS, a BC and a second VRFS. Where each polynucleotide of the array comprises an analyte capture sequence, the second VRFS may be between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI, i.e. upstream of the analyte capture sequence and downstream of the BC and the UMI. For example, the second VRFS may be between the UMI and the analyte capture sequence. The second VRFS may be between the BC and the analyte capture sequence. Each polynucleotide of the array may comprise in the direction from 5’ to 3’: a PCR handle, a BC, a first VRFS, a UMI, a second VRFS and an analyte capture sequence. Each polynucleotide of the array may comprise in the direction from 5’ to 3’ a PCR handle, a UMI, a first VRFS, a BC, a second VRFS and an analyte capture sequence.

[0108] A VRFS (e.g. the first and / or the second VRFS) may be any number of nucleotides in length, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. A VRFS may be 2 to 20 nucleotides in length, such as 2 to 15 nucleotides in length, 2 to 10 nucleotides in length, 2 to 8 nucleotides in length, 4 to 8 nucleotides in length, 4 to 6 nucleotides in length, or about 4 nucleotides in length. As discussed in the summary of the invention, a user may delimit a BC and / or a UMI in a sequence read corresponding to at least a portion of a polynucleotide by aligning it to a reference sequence that comprises the VRFS. The length of a VRFS is a balance between increasing the power of the alignment versus increasing efficiency of the synthesis of each polynucleotide in the array. The longer the VRFS, the more nucleotides there are to align, and longer VRFSs can be identified more easily if they contain errors such as sequencing or synthesis errors. However, longer polynucleotides are less efficient to synthesise and increasing the length of the VRFS(s) also increase the chances of errors being introduced during the synthesis.

[0109] A VRFS (e.g. the first and / or the second VRFS) may comprise 1 or more, 2 or more, 3 or more, or 4 or more constant nucleotides. In some cases, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90% of the nucleotides of the VRFS are constant nucleotides. In some cases, all of the nucleotides of the VRFS are constant nucleotides, i.e. the VRFS consists of constant nucleotides. As used herein, the term “constant nucleotide" refers to one or more positions in the sequence of the VRFS which is the same in all instances of a VRFS of the same sequence. For example, a VRFS comprising the sequence ACNCT in each polynucleotide of the array, where N is variably selected from A, C, G or T, comprises four constant nucleotides. If the array comprises 5 different VRFSs of the sequences ACNCT, AGNGC, TCNTC, TTNGG, and CANCA, each different VRFS sequence has 4 constant nucleotides (ACCT, AGGC, TCTC, TTGG and CACA respectively). In some cases, the term “ constant nucleotide" refers to a position in the sequence of the VRFS which is the same in all polynucleotides of the array (i.e. all polynucleotides of the array have the same VRFS), or is the same in all polynucleotides of sub-arrays.

[0110] A VRFS (e.g. a first and / or second VRFS) may comprise 1 or more variable nucleotides. As used herein, the term “variable nucleotide" refers to one or more positions in the sequence of the VRFS which is variable across polynucleotides comprising said VRFS. A variable nucleotide in a VRFS may be added to each polynucleotide in an array by degenerative nucleotide synthesis, e.g. from a pool of nucleotides comprising A, C, G and T. The variable nucleotide may be a limited variable nucleotide. The term “limited variable nucleotide" refers to one or more positions in the sequence of the VRFS which has limited variability across polynucleotides comprising said VRFS. For example, a limited variable nucleotide in a VRFS may be added to each polynucleotide in an array by degenerative nucleotide synthesis from a pool of nucleotides in which one or more of A, C, G or T are absent. The limited variable nucleotide may be M (A or C), R (A or G), W (A or T / U), S (C or G), Y (C or T / U), K (G or T / U), V (A or C or G; not T / U), H (A or C or T / U; not G), D (A or G or T / U; not C) or B (C or G or T / U; not A). Limited variable nucleotides may be particularly useful on the 3’ and / or 5’ ends of the VRFS where the identity of the adjacent nucleotide of the UMI, BC or PCR handle is known or is known to be restricted. For example, the limited variable base ‘V’ may be used in a VRFS adjacent to a polyT analyte capture region. Accordingly, the second VRFS may consist of a ‘V’ or may comprise a ‘V’ at the 3’ end of the second VRFS.

[0111] In some cases where a VRFS comprises a variable base, the variable base is not adjacent to a variable region, e.g. a BC or UMI. In some cases where a VRFS comprises a variable base, the variable base is flanked on 5’ and 3’ sides by a constant nucleotide of the VRFS.

[0112] In some cases, a VRFS (e.g. a first and / or second VRFS) does not comprise the same nucleotide in three consecutive positions (e.g. does not comprise a CCC, TTT, AAA or GGG sequence). In some cases, a VRFS does not comprise the same nucleotide in two consecutive positions. To put it another way, a VRFS may consist of a sequence in which adjacent nucleotides are different (e.g AT, AG, AC, TA, TG, TC, GA, GT, GC, CA, CT, or CG sequences). Repeats of the same nucleotide consecutively can lead to errors, for example when synthesising the polynucleotide and when sequencing the polynucleotide. For example, the first nucleotide (5’ nucleotide) of the first VRFS may be different to the second nucleotide of the first VRFS, the last nucleotide (3’ nucleotide) of the first VRFS may be different from the penultimate nucleotide of the first VRFS, the first nucleotide (5’ nucleotide) of the second VRFS may be different to the second nucleotide of the second VRFS, and / or the last nucleotide (3’ nucleotide) of the second VRFS may be different from the penultimate nucleotide of the second VRFS.

[0113] Where a polynucleotide comprises a first VRFS and a second VRFS, the first VRFS typically comprises a different sequence to the second VRFS. This is advantageous because it reduces the risk that a sequencing read is misaligned to a reference read. In some cases, the first VRFS has a Levenshtein edit distance of at least one when compared to the second VRFS, such as a Levenshtein edit distance of at least two, at least three, or at least four. The Levenshtein edit distance between the first and second VRFS may be two less than the total number of nucleotides in the first VRFS, such as one less than the total number of nucleotides in the first VRFS, or the same as the total number of nucleotides in the first VRFS. In some cases, the first VRFS has a Hamming distance of at least one when compared to the second VRFS, such as a Hamming distance of at least two, at least three, or at least four. The Hamming distance between the first and second VRFS may be two less than the total number of nucleotides in the first VRFS, such as one less than the total number of nucleotides in the first VRFS, or the same as the total number of nucleotides in the first VRFS. The measure of difference between the first and second VRFS is preferentially determined by Levenshtein edit distance. In some cases, the nucleotide at each position of the first VRFS is different to the nucleotide at each corresponding position of the second VRFS. As used in this context, the term “ corresponding position" means that the first nucleotide of the first VRFS read in a 5’ to 3’ direction corresponds to the first nucleotide of the second VRFS read in a 5’ to 3’ direction, the second nucleotide of the first VRFS read in a 5’ to 3’ direction corresponds to the second nucleotide of the second VRFS read in a 5’ to 3’ direction, and so on. If the first VRFS and the second VRFS are not the same length (e.g. AGTC and CTGAA) then the 5’ nucleotide of each VRFS will be considered to correspond, i.e. the corresponding nucleotides in the above example will be A and C, G and T, T and G, and C and A; the final A of the second VRFS will be ignored.

[0114] Where a polynucleotide comprises a first VRFS and a second VRFS, the first VRFS is typically not the reverse complement of the second VRFS. This is advantageous because it reduces the risk that a sequencing read or its reverse complement is misaligned to a reference read, and also reduces the chances that the polynucleotide forms a hairpin. In some cases, the first VRFS does not hybridise to the second VRFS, for example, under conditions of lx phosphate buffered saline, pH 7.4, 25°C.

[0115] The array of polynucleotides may comprise fewer than 10, fewer than 8, or fewer than 5 different first VRFSs. To put it another way, each polynucleotide of the array comprises a first VRFS selected from a pool of fewer than 10 different first VRFSs, fewer than 8 different first VRFSs or fewer than 5 different first VRFSs. Typically, essentially every polynucleotide, or every polynucleotide in the array comprises the same first VRFS. The array of polynucleotides may comprise fewer than 10, fewer than 8 or fewer than 5 different second VRFSs. To put it another way, each polynucleotide of the array comprises a second VRFS selected from a pool of fewer than 10 different second VRFSs, fewer than 8 different second VRFSs or fewer than 5 different second VRFSs. Typically, essentially every polynucleotide, or every polynucleotide in the array comprises the same second VRFS. Where the VRFS is designed to comprise a variable or limited variable base, for the purposes of the present disclosure each VRFS falling within the definition of the VRFS including the variable or limited variable base is considered to be the same VRFS. For example, if the VRFS is VACG, then each of AACG, CACG and GACG are considered the same VRFS. Furthermore, it is acknowledged that due to errors in the synthesis of polynucleotides, it is possible that some polynucleotides of the array may have different VRFSs than the intended pool of VRFSs: for the purposes of the disclosure, errors in polynucleotide synthesis are not counted towards whether a VRFS is considered as a different VRFS sequence.

[0116] The 5’ and / or 3’ ends of the variable regions (i.e. the BC and the UMI) typically comprise a different nucleotide to the adjacent nucleotide of the flanking region (e.g. the adjacent nucleotide of the PCR handle, the first VRFS, the second VRFS and / or the analyte capture region). In some cases, the 5’ nucleotide of the BC and / or the UMI is different to the 3’ nucleotide of its 5’ flanking sequence, wherein its 5’ flanking sequence is the PCR handle or the first VRFS. In some cases, the 3’ nucleotide of the BC and / or the UMI is different to the 5’ nucleotide of its 3’ flanking sequence, wherein its 3’ flanking sequence is the first VRFS, the second VRFS, or the analyte capture region. For example, the first VRFS may be immediately 3’ of the BC, and the 5’ nucleotide of the first VRFS is different to the 3’ nucleotide of the BC. The second VRFS may be immediately 3’ of the UMI, and the 5’ nucleotide of the second VRFS is different to the 3’ nucleotide of the UMI. In some cases, the first VRFS is immediately 3’ of the UMI and the 5’ nucleotide of the first VRFS is different to the 3’ nucleotide of the UMI. In some cases, the second VRFS is immediately 3’ of the BC, and the 5’ nucleotide of the second VRFS is different to the 3’ nucleotide of the BC.

[0117] In some cases, the PCR handle may be immediately 5’ of the BC and the 3’ nucleotide of the PCR handle is different to the 5’ nucleotide of the BC. Alternatively, the PCR handle may be immediately 5’ of the UMI and the 3’ nucleotide of the PCR handle is different to the 5’ nucleotide of the UMI.

[0118] In some cases, the analyte capture region may be immediately 3’ of the UMI and the 5’ nucleotide of the analyte capture region is different to the 3’ nucleotide of UMI. Alternatively, the analyte capture region may be immediately 3’ of the BC and the 5’ nucleotide of the analyte capture region is different to the 3’ nucleotide of BC. In some cases, the second VRFS is immediately 5’ of the analyte capture sequence, and the 3’ nucleotide of the second VRFS is different to the 5’ nucleotide of the analyte capture sequence.

[0119] In some cases, the nucleotide at the 5’ end and / or the 3’ end of the BC and / or the UMI is a limited variable base. For example, the 3’ nucleotide of the second VRFS may be a nucleotide other than a T (the 3’ is not a T), e.g. the 3’ nucleotide may be selected from the group of A, C and G, or may be a limited variable base selected from the group of V, M, R and S. This is particularly useful where the analyte capture region comprises a 5’ T, such as wherein the analyte capture region comprises or consists of a polyT sequence.

[0120] Analyte capture sequence / Analytes

[0121] Each polynucleotide, or a portion of the polynucleotides, of the array described herein may comprise an analyte capture sequence.

[0122] An analyte capture sequence may be any nucleotide sequence suitable for capturing analyte in a sample. In some cases, the analyte(s) may be biological analytes or may be selected from polynucleotides DNA and / or RNA, or from oligonucleotides, DNA, cDNA, RNA, mRNA, rRNA, tRNA, snRNA, siRNA and / or ribozymes, proteins, polypeptides and / or peptides, cell surface receptors or cells. Most typically the analytes are mRNA. Other examples of analytes that may be captured by suitable analyte capture sequences include amino acids, metal ions, inorganic salts, polymers, nucleotides, oligonucleotides, polynucleotides, dyes, bleaches, pharmaceuticals, diagnostic agents, recreational drugs, explosives and / or environmental pollutants. Such analytes may be captured, for example, by an aptamer or other types of analyte capture sequences as are known in the art.

[0123] Typically, the analyte capture sequence is at least 10, or at least 15, 20, 25 or 30 nucleotides in length, such as from about 15 to about 50, from about 20 to about 40 or from about 25 to about 35 nucleotides. Most typically the analyte capture sequence is 20 to 40 nucleotides in length. In some cases, the analyte capture sequence may comprise one or more nucleotide analogues, such as analogues described herein, that form double- stranded hybrids with higher stability than natural nucleotides. In this case, the analyte capture sequence could be shorter, such as at least 3, 4, 5, 6, 8 or 9, for example between 3 and 50, or 40 or 30 or 20 nucleotides in length, provided that the analyte capture sequence is capable of binding or hybridizing to target analyte, e.g. such that analyte sequence can be amplified as described herein. An analyte capture sequence may include nucleotide analogues as described herein.

[0124] In some cases, an analyte capture sequence may be a DNA capture sequence, an RNA or mRNA capture sequence, or a polypeptide capture sequence.

[0125] An analyte capture sequence, particularly a 3’ analyte capture sequence, may be a polythymidine sequence (a polyT sequence). Polythymidine may hybridise to and capture any polynucleotide in the sample that comprises a suitable polyadenosine, such as polyadenylated mRNA. Typically, the polythymidine sequence is at least 10, or at least 15, 20, 25 or 30 thymidines in length, such as from about 15 to about 50, from about 25 to about 35 or most typically from about 20 to about 40 thymidines. The analyte capture sequence may be (about) 30 thymidines in length.

[0126] An analyte capture sequence, particularly a 3’ analyte capture sequence, may be a primer sequence. In this case, each polynucleotide of the array may comprise two PCR handle sequences, one 5’ of the BC and the UMI (to ensure that these are amplified with the analyte) and a primer sequence 3’ of the BC and the UMI (to capture / amplify an analyte sequence). Where the analyte capture sequence is a primer, the analyte is captured by hybridising to the primer sequence and is incorporated into the polynucleotide by polymerase extension. In this case, the analyte is typically a polynucleotide, such as a DNA sequence, an RNA sequence or an mRNA sequence. The primer may be designed to target a particular genomic region or mRNA transcript. Where the analyte capture sequence is a primer, a number of different analyte capture sequence-primers may be included in the polynucleotides of the array, such as at least 5, at least 10, at least 20, at least 50 or at least 100 different analyte capture sequence-primers. This may allow for the subsequent targeted amplification of a number of different analyte sequences, e.g. from within genomic DNA. In some cases, the analyte capture sequence-primer is a random or semi-random primer, e.g. designed to non-specifically capture polynucleotides.

[0127] In other cases, the analyte capture sequence(s) may comprise or consist of an aptamer. Aptamers can be produced using SELEX (Stoltenburg, R. et al.. (2007), Biomolecular Engineering 24, p381-403; Tuerk, C. etal., Science 249, p505-510; Bock, L. C. etal., (1992), Nature 355, p564-566) or NON-SELEX (Berezovski, M. et al. (2006), Journal of the American Chemical Society 128, p 1410-1411). Typically, an aptamer may be at least 15 nucleotides in length, such as from about 15 to about 50, from about 20 to about 40 or from about 25 to about 30 or nucleotides in length. An aptamer may bind to analyte such as small molecules, proteins, nucleic acids or cells. Aptamers may be designed or selected to bind to pre-determined target analyte(s). In one example, the aptamer may bind to a Coronaviridae protein or SARS-CoV-2 protein, as described in described in PCT Publication No. WO 2021 / 229230.

[0128] In some cases, an analyte capture sequence may comprise or consist of a biotinylated nucleotide sequence. Nucleotides or polynucleotides may be biotinylated using methods known in the art. Typically, the biotinylated sequence may be at least 10, or at least 15, 20, 25 or 30 nucleotides in length, such as from about 15 to about 50, from about 20 to about 40 or from about 25 to about 35 nucleotides. A biotinylated capture sequence may be used to capture any suitable target analyte comprising streptavidin or avidin.

[0129] In some cases, an analyte capture sequence may comprise or consist of a nucleotide sequence designed to hybridise to a complementary sequence in a target polynucleotide / analyte. In some cases, the capture sequence is for capturing / hybridising to transposed DNA. In this case an analyte capture sequence may comprise or consist of a sequence that is complementary to transposed DNA in a sample, for example to a transposed MEDS DNA sequence. In other cases, the sequence may be gene or transcript-specific, such as a polynucleotide sequence that is complementary to, or at least 80%, 85%, 90%, 95%, 98% or 99% complementary to, a viral sequence, a bacterial sequence or a sequence associated with a disease or disorder, such as a sequence from a cancer-associated antigen or a neoantigen. In some cases, the analyte capture sequence(s) may hybridise to a nucleotide sequence that encodes a part of a Coronaviridae protein or SARS-CoV-2 protein, as described in described in PCT Publication No. WO 2021 / 229230.

[0130] In other cases, the sequence may be designed to capture a polynucleotide tag added to analyte of interest.

[0131] In some cases, all of the polynucleotides, or all of the polynucleotides in an array may have the same analyte capture sequence. In other cases, a combination of different analyte capture sequences may be used.

[0132] Typically, the analyte capture sequence is 3’ of the PCR handle sequence. Typically, the analyte capture sequence is 3’ of the portion of the polynucleotide comprising the PCR handle sequence, the BC, the UMI, the first VRFS and the optional second VRFS. More typically, the analyte capture sequence is at the 3’ end of the polynucleotide. In some cases, the analyte capture sequence may comprise a short sequence at its 3’ end, such as a single limited variable base (e.g. a ‘v’ nucleotide).

[0133] In some cases, the polynucleotides of the array described herein do not comprise an analyte capture sequence. Analyte may be ‘captured’ by the polynucleotides of the array by other means, for example by ligation. Where a second VRFS is used, splint ligation may be used, for example, by using a splint oligonucleotide complementary to the second VRFS and to an analyte sequence.

[0134] In some cases, some of the polynucleotides of the array are hybridised to, or further comprise a sequence corresponding to, an analyte. In this case, this array is defined after contact with the analyte, and may be before or after amplification of the polynucleotide and the analyte by reverse transcription and / or PCR. In some cases, at least 0.000001% of the polynucleotides of the array are hybridised to, or further comprise a sequence corresponding to, an analyte. For example, at least 0.00001%, at least 0.0001%, at least 0.001%, at least 0.01%, at least 0.1%, at least 10%, at least 50%, at least 90%, at least 95% or at least 99% of the polynucleotides of the array are hybridised to, or further comprise a sequence corresponding to, an analyte. In a typical single-cell sequence experiment, a single micro-particle comprising a micro-bead and an array of polynucleotides may comprise 2xlO10polynucleotides. A mammalian cell typically comprises IxlO5to IxlO6mRNA molecules. In some cases, the invention relates to a library of polynucleotides, wherein each polynucleotide of the library may be as defined as according to the array of polynucleotides described herein, wherein each polynucleotide is hybridised to, or further comprises a sequence corresponding to, an analyte. The library may be generated by contacting an array of polynucleotides as described herein with a sample comprising analyte molecules, such as mRNA molecules. The generation of the library may comprise performing reverse transcription and / or PCR amplification. The generation of the library may comprise performing splint ligation of the polynucleotides of the array with the analyte molecules and / or further performing PCR amplification.

[0135] Hairpins

[0136] Optionally, a proportion of or each polynucleotide of the array comprises a hairpin sequence. However, in some embodiments, a proportion of or each polynucleotide of the array does not comprise a hairpin sequence.

[0137] A hairpin sequence includes palindromic elements that will anneal together to form a hairpin structure, as is well understood in the art. The hairpin structure may optionally include additional nucleotides in between the pair of palindromic (reverse complementary) elements, e.g. a loop sequence. If the hairpin structure is melted / denatured, then two different polynucleotides having the same hairpin sequence can anneal to each other to form a dimer. Dimers form preferentially and at a higher temperature than hairpins having the same sequence because the dimers have twice as many base pairs. The hairpin length and sequence can be selected by those skilled in the art to have a desired melting temperature for the hairpin structure and / or corresponding dimers. For example, the hairpin sequence could be designed to have an unstable hairpin structure, but be stable as a dimer, at room temperature (about 20°C, or about 18 to 23°C). A typical hairpin sequence may contain palindromic sequence elements, each about 5, or about 6, 7, 8, 9, 10, 12, or 15 to about 30, or about 27, 25, 20, or 18 nucleotides in length, with an optional loop sequence element or spacer in between the two palindromic regions. Most typically, each of the pair of palindromic sequences is about 6 to 20 nucleotides in length. In some cases, the palindromic regions may comprise one or more nucleotide analogues, such as analogues described herein, that form double-stranded hybrids with higher stability than natural nucleotides. In this case, the palindromic regions could be shorter, such as at least 3, 4, or 5, nucleotides in length, provided that the hairpin sequence is able to form both a hairpin structure and dimers.

[0138] In some cases, the hairpin sequence may contain sequence elements that are substantially palindromic. The term substantially palindromic as used herein is intended to refer to sequence elements that comprise a degree of complementarity sufficient to form a stable hairpin structure. The stability of the hairpin may be assessed, for example, at 20°C, 25°C, 30°C or 37°C. The hairpin may be considered stable if for example, less than 10 %, less than 5 %, less than 1% or less than 0.1 % of hairpin sequences are linear in solution under the conditions tested. A hairpin sequence may therefore contain first and second sequence elements, wherein the second sequence element comprises (a) a complementary sequence to the sequence of the first sequence element, or (b) comprises a sequence having one, two, three, four or five nucleotide substitutions to the sequence of (a). Typically, the hairpin sequence / palindromic region spontaneously forms a hairpin structure.

[0139] In some cases, no loop sequence per se is included, although a small number of bases at the centre of a palindromic sequence may spontaneously form a non-base pairing loop. Not including a non-palindromic loop sequence increases the number of paired bases when two “hairpin” polynucleotides form a dimer. When a non-palindromic loop sequence is present, it is typically less than about 50, or less than about 40, 35, 30, 25, 20, 15 or 10 nucleotides in length and is typically about 4 to 12, or 6 to 10 nucleotides in length. The loop may typically have a simple structure with no base-pairing between nucleotides of the loop. However, sequences in between the palindromic elements that provide additional secondary structure are also contemplated, for example a clover structure. Such structures fall within the term “hairpin” as used herein, provided the sequences are also able to dimerise, as described herein.

[0140] Sub-Arrays, Micro-Particles and other Products

[0141] In some cases, the array may be sub-divided into a plurality of sub-arrays, wherein the BC of each polynucleotide is the same as the BC of essentially each other polynucleotide of the same sub-array, but different from the BC of the polynucleotides of essentially every different sub-array. The BC of each sub-array may be used to distinguish the source of each polynucleotide. Sub-arrays of polynucleotides as described herein are typically used to probe different parts or sub-elements of a sample, such as individual cells of a cellular sample or different spatial positions on a surface of a tissue sample. Different capture-polynucleotides of the same sub-array may capture different analytes in the same part or sub-element of the sample. Typically the capture-polynucleotides of each sub-array are separated from other sub-arrays before being contacted with the analytes of the part / sub-element of the sample. For example, each sub-array may be isolated and / or contacted with the analytes in a separate fluidic compartment as described further herein.

[0142] All of the polynucleotides of the same initially isolated sub-array may have the same BC. Hence, the BC allows a mixed pool of polynucleotides from different sub-arrays to be identified. For example, different polynucleotides may be shared between different micro-beads / micro-particles, or different wells or discrete pre-defined positions on a surface or the like. Each of the polynucleotides associated with each micro-bead, microparticle, well, or position may be a sub-array of the polynucleotides. All of the polynucleotides of a sub-array may have the same BC as each other, and a different BC to the polynucleotides associated with each other sub-array. Typically, each sub-array is contacted with a different sample or part of a sample. For example, each sub-array or micro-particle might be contacted with the analytes of different single cells, or each subarray of each well or position of a surface may be contacted with a different part of a tissue sample laid over the surface. The BC of the polynucleotides may be used to tag the analytes from the sample. After analyte capture, the polynucleotides from the different sub-arrays can be combined for amplification and / or sequencing en masse. The BC associated with each sequenced polynucleotide can subsequently identify the sample or source of the corresponding captured analytes.

[0143] The phrase “the BC of each polynucleotide is the same as the BC of essentially each other polynucleotide of the same sub-array", encompasses scenarios where some polynucleotides within the same sub-array may comprise a different BC due to e.g. errors in synthesis. However, these different BCs are considered for all intents and purposes to be the same BC as they were designed to comprise identical sequences. The phrase “different from the BC of the polynucleotides of essentially every different sub-array', encompasses the possibility that the number of sub-arrays exceeds the number of different barcodes, and / or that two or more sub-arrays may have the same barcode, but that this overall would not significantly affect the data generated by the experiment. For example, fewer than 1 % of sub-arrays may comprise the same BC, such as fewer than 0.1%, fewer than 0.01%, fewer than 0.001%, or fewer than 0.0001% of sub -arrays may comprise the same BC. In some cases, a different BC is added to each of at least 1000, 4000, 1.6xl04, 6.4xl04, 2.6xl05, IxlO6, 4xl06or 1.6xl07different sub-arrays.

[0144] The invention also provides a micro-particle comprising a micro-bead and an array (or sub-array) of polynucleotides as described herein. Each polynucleotide is bound to the micro-bead. In some cases, the array (or sub-array) of polynucleotides bound to the micro-bead comprises fewer than 10, fewer than 8, fewer than 5 or the same BC sequence of essentially each other polynucleotide in the array (or sub-array). This allows polynucleotides derived from the same micro-particle to be associated as coming from the same source.

[0145] Microbeads are typically less than 500 pm, or less than 400 pm, 300 pm or 200 pm in diameter, for example, between 10 and 500 pm, 20 and 400 pm, 40 and 300 pm , or 50 and 200 pm. Typically a micro-bead is approximately spherical or sphere-like. Microbeads with surface-attached polynucleotides are well-known in the art and may be made from, for example, a biocompatible polymer such as polystyrene, polyacrylamide or hydroxylated methacrylic polymer, or from controlled pore glass. The micro-particles described herein may comprise a micro-bead as described herein with an array of polynucleotides as described herein, wherein each polynucleotide in the array is attached to the bead at the 3’ or the 5’ end of the polynucleotide. The opposite end (5’ or 3’) is typically free in solution.

[0146] Dissolvable beads and hydrogel beads have also been described and are encompassed in the present disclosure. Dissolvable beads may, for example, be made from crosslinked acrylamide with disulfide bridges that are cleaved with dithiothreitol. An array of polynucleotides may be embedded in the bead matrix and released when the bead is dissolved. A micro-particle comprising a micro-bead and an array of polynucleotides bound to the micro-bead, as described herein, may be a typical micro-bead with surface bound polynucleotide or a dissolvable bead with embedded polynucleotide.

[0147] Polynucleotides may be synthesised on the bead using methods known in the art or described herein. For example, the phosphoramidite method may be used, in which one nucleotide or pre-synthesised block is added per synthesis cycle. The identifier sequence(s) may be generated as described elsewhere herein. Other and / or longer sequence elements may in some cases be added using enzymatic ligation methods, such as using DNA ligase, chemical ligation methods, such as phosphoramidate ligation, and / or click chemistry ligation methods, such as the azide-alkyne cycloaddition reaction. Suitable methods are known in the art.

[0148] In some cases, the invention relates to a plurality of micro-particles as described herein. The number of micro-particles may be selected for a specific purpose or experiment, such as to capture sample analytes for the generation one or more libraries. Typically the plurality of micro-particles comprises at least 1000, or at least 10,000, at least 100,000, at least 1,000,000 or at least 10,000,000 micro-particles. For example, the plurality of micro-particles may comprise between 1000, 10,000, or 100,000, 1,000,000 and IxlO9, such as between about 10,000 and about IxlO9,or between 100,000 and IxlO8microparticles. Typically the set of polynucleotides associated with each micro-particle have a different shared BC from the polynucleotides associated with essentially each of the other micro-particle. Hence, the analytes that are captured by the set of polynucleotides associated with each micro-particle can be distinguished. The potential diversity of BC is dependent on the number of nucleotides or nucleotide blocks that make up the BC and the size of the nucleotide block pools, as described above. The BC will typically be long enough that the potential sequence diversity is on a similar scale to the number of microparticles in the plurality (such as with lOx less or more), or well in excess of the plurality of micro-particles in the plurality, such as at least 50x or lOOx in excess, or least 10, 20, 50, 100, 200 or 500 times in excess.

[0149] The invention may also relate to a surface comprising a plurality of wells or discrete predetermined positions, wherein each well or discrete pre-determined position is associated with a (sub-)array of polynucleotides as described herein. Typical examples include solid supports such as well plates (for example, one or more 6, 12, 24, 48, 96, 384, 1536 or 3456 well plates) microplates, or slides, as are known in the art.

[0150] The polynucleotides may be attached to the surface using chemistry known in the art. The polynucleotide may be attached to the solid support through a linker of non-specific bases. The surface may be made from any suitable materials, for example, a biocompatible polymer such as polystyrene, polyacrylamide or hydroxylated methacrylic polymer, or from controlled pore glass.

[0151] In some cases, the polynucleotides may be synthesised on the solid support in a similar way to synthesis on a micro-bead, as described above. In other cases, the polynucleotides may be pre-synthesised and then divided by aliquoting different sub-sets (sub-arrays) of the polynucleotides and / or micro-particles, as described herein, into the wells or onto the pre-defined discrete spatial positions.

[0152] Solid-supports or surfaces comprising an array or plurality of sub-arrays of polypeptides as described herein are particularly useful for capturing and / or analysing sample analytes in a spatial manner, for example the analytes of a tissue or other two or three dimensional sample laid over or otherwise contacted with the surface / support in a manner that captures the spatial arrangement of the analytes as captured. The invention also relates to a kit for generating one or more libraries from one or more groups of analytes, the kit comprising an array of polynucleotides, a micro-particle, a plurality of micro-particles, or a surface / solid support of the invention as described herein. The invention also relates to the use of an array of polynucleotides, a microparticle, a plurality of micro-particles, or a surface / solid support of the invention, or a kit of the invention as described herein for generating one or more libraries from one or more groups of analytes. The kit may also comprise buffers, enzymes and other components used for generating the library as described herein and / or instructions for use in a method of generating the library. For example, the kit may comprise reagents for performing reverse transcription (such as free nucleotides, a reverse transcriptase enzyme and / or a template switch oligonucleotide), reagents for performing PCR (such as free nucleotides, PCR primers e.g. that are complementary to the PCR handle sequences on the polynucleotide and / or the template switch oligonucleotide, and a DNA polymerase), and / or reagents for attaching a sequencing adaptor.

[0153] Methods for delimiting a barcode sequence and / or a unique molecular identifier sequence

[0154] The polynucleotides described herein comprising at least one VRFS are particularly suitable for use in applications in which the BC and / or the UMI need to be extracted from a sequencing read. For example, polynucleotides of the invention comprising a BC and / or a UMI may be used to capture and sequence analytes. Sequencing the analytes generally involves amplifying polynucleotides comprising a BC, a UMI and the analyte (in order to provide multiple copies for sequencing). The resulting copies are sequenced resulting in sequence reads corresponding to the polynucleotide comprising the BC, UMI and the analyte. As discussed above, the BC and the UMI are indicative of the source of the read (for example, reads comprising a sequence corresponding to the same BC may come from a common source such as a common micro-bead, and all reads comprising the same UMI may correspond to the same analyte molecule). Thus, properly delimiting (identifying) the BC and / or the UMI allows the source of the read to be determined. However, errors in the BC and / or UMI may be introduced and the use of a VRFS may help to delimit (identify) the BC and / or the UMI in such cases. Furthermore, the use of a VRFS to delimit the BC and / or the UMI enables the use of data from reads comprises sequences corresponding to BC and / or UMI sequences of variable-length (e.g. UMIs whose length is different to that predicted due to addition or removal of a nucleotide during synthesis or sequencing), as described herein, to be utilised.

[0155] Whilst the improved delimitation (identification) of the BC and the UMI may not be improved in a single sequencing read where no errors have been introduced, it will be improved in a single sequencing read which comprises errors and will also be improved when defining or delimiting a BC and / or UMI in large groups of sequencing reads such as reads corresponding to arrays of polynucleotides, for example reads obtained by sequencing an array of polynucleotides as disclosed herein. The errors may be errors in synthesis of the polynucleotides of the array as described herein, such as the addition of one or more extra nucleotides, the incorporation of one or more incorrect nucleotides and / or the failed incorporation of one or more nucleotides, errors in the PCR amplification of the polynucleotides of the array as described herein, such as the addition of one or more extra nucleotides, the incorporation of one or more incorrect nucleotides and / or the failed incorporation of one or more nucleotides, or errors in the sequencing of the polynucleotides of the array as described herein, such as the skipping of one or more nucleotides, the mis-assignment of one or more nucleotides, or the incorrect incorporation of one or more nucleotides that are not present in the polynucleotide.

[0156] Accordingly, the invention provides a method for delimiting a BC and / or a UMI in a read, or a plurality of reads. The read or plurality of reads may be derived from an array of polynucleotides as described herein. The read or plurality of reads corresponds to the sequence of at least a portion of a polynucleotide. The portion of the polynucleotide, or its reverse complement, comprises (a) a PCR handle sequence, (b) a BC, (c) a UMI, and (d) a first VRFS. The first VRFS is located between the BC and the UMI of the polynucleotide or its reverse complement. The method comprises (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS. The method further comprises predicting the positions of the BC and / or the UMI of the read and / or its reverse complement based on the alignment. By performing an alignment of the read to a reference sequence comprising or consisting of the first VRFS, it may be possible to determine the ends, i.e. the 5’ or 3’ nucleotide, of the BC and / or the UMI of the read based on the intermediate position of the VRFS. The polynucleotide or its reverse complement may further comprise an analyte capture sequence as described herein. The polynucleotide or its reverse complement may further comprise a second VRFS as described herein. The second VRFS is typically located 3’ of the portion of the polynucleotide (or its reverse complement) comprising the BC and the UMI. Where the polynucleotide or its reverse complement comprises an analyte capture sequence, the polynucleotide (or its reverse complement) typically comprises the second VRFS between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI. Where the polynucleotide or its reverse complement further comprises a second VRFS, the reference sequence may comprise at least the sequence of the first VRFS and the sequence of the second VRFS. Using both VRFS elements improves the alignment and enables the positions of the BC or the UMI to be predicted with greater accuracy than a single VRFS, as either the BC or the UMI will be flanked with both a first VRFS and a second VRFS. For example, the polynucleotide or its reverse complement will comprise a sequence having, in a 5’ to 3’ direction, (A) a BC, a first VRFS, a UMI and a second VRFS, or (B) a UMI, a first VRFS, a BC and a second VRFS.

[0157] The invention also provides a method for delimiting a BC and / or a UMI in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises (a) a PCR handle sequence (b) a BC, (c) a UMI and (d) a second VRFS. The second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI. The method comprises (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and (ii) predicting the positions of the BC and / or the UMI based on the alignment. In this embodiment, it is not necessary for a first VRFS to be present. Instead, the second VRFS may be used to align with the second VRFS of a reference sequence, and it may be possible to delimit the end of the BC or the UMI, and / or further predict the positions of the BC and / or UMI based on the alignment with the second VRFS. The polynucleotide or its reverse complement typically further comprises an analyte capture sequence as described herein. Where the polynucleotide or its reverse complement comprises an analyte capture sequence, the polynucleotide (or its reverse complement) typically comprises the second VRFS between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI.

[0158] The term “delimiting a BC and / or a UMI” refers to determining the nucleotides in a read which correspond to the BC and / or the UMI. Once the nucleotides in a read which correspond to the BC and / or the UMI are known, it is possible to identify the BC and / or the UMI (for example to provide a sequence for the BC and / or the UMI or to use the BC and / or the UMI to determine the source of the read and group together reads of the same origin). Thus, in some cases the methods further comprise a step of identifying the BC and / or the UMI and / or a further step of using the delimited BC and / or UMI to determine the source of the read.

[0159] In some cases, the term “delimiting” may be replaced by the term “identifying”. For example, the invention further provides a method for identifying a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises:

[0160] (a) a PCR handle sequence;

[0161] (b) a barcode sequence (BC);

[0162] (c) a unique molecular identifier sequence (UMI); and

[0163] (d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI, and wherein the method comprises:

[0164] (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and

[0165] (ii) predicting the positions of the BC and / or the UMI of the read and / or its reverse complement based on the alignment.

[0166] Similarly, the invention further provides a method for identifying a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises:

[0167] (a) a PCR handle sequence;

[0168] (b) a barcode sequence (BC);

[0169] (c) a unique molecular identifier sequence (UMI);

[0170] (d) a second variable region flanking sequence (VRFS), wherein the second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI, and wherein the method comprises:

[0171] (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and

[0172] (ii) predicting the positions of the BC and / or the UMI based on the alignment.

[0173] The phrase “predicting the positions of the BC and / or the UMF may be interpreted to mean determining the nucleotides of the read that are considered to make up the BC and / or the UMI.

[0174] The read will typically correspond to at least a portion of a polynucleotide that has been generated by PCR amplification, and so it is generally equally likely for the read to be of a polynucleotide having the original sequence of the polynucleotide prior to PCR amplification, or the reverse complement of the original sequence. Although the term “reverse complement is explicitly used when explaining the described methods, this is also implicitly encompassed in these methods even when not mentioned, unless it is clear that the reverse complement should not be encompassed (for example if the original sense of the polynucleotide is selected over its reverse complement in one of the following steps).

[0175] In some cases, the methods described herein are performed on a plurality of reads. For example, the methods described herein may be performed on at least IxlO4, at least IxlO5, at least IxlO6, at least IxlO7, at least IxlO8or at least IxlO9reads. It should be understood that there is no real upper limit to the number of reads to which the method can be applied. For simplicity the methods are further described in relation to a single read, but it should be understood that said further embodiments can equally be applied with the method is performed on a plurality of reads.

[0176] The polynucleotide, or its reverse polynucleotide of the described methods may be further described according to the array of polynucleotides as described herein. For example, where the method relates to a read corresponding to a polynucleotide comprising a second VRFS but not a first VRFS, all of the features relating to the array of polynucleotides described above equally apply to this embodiment despite the lack of a first VRFS.

[0177] In some cases, the methods described herein are performed on reads produced by sequencing a polynucleotide or an array of polynucleotides, such as a polynucleotide or an array of polynucleotides as described herein. The sequencing may be performed by any means known to the skilled person. For example, the sequencing may be performed by long-read sequencing technologies, such as Nanopore or PacBio. The methods described herein are particularly useful for long-read sequencing technologies. As used herein, the term long-read sequencing technologies is intended to encompass technologies capable of sequencing polynucleotides of at least 1000 nucleotides in length, such as at least 10000 nucleotides in length. The sequencing may be performed by massively parallel sequencing technologies (otherwise known as next-generation sequencing (NGS)) may be used to sequence the library of cellular analytes, e.g. wherein the library comprises amplified DNA generated by PCR. The NGS technology may be pyrosequencing-based, reversible terminator chemistry-based, or sequencing-by-ligation- based.

[0178] The polynucleotide or its reverse complement that corresponds to a read may further comprise an analyte, i.e. a captured analyte, as described herein. The method may comprise capturing analyte molecules on a polynucleotide, or an array of polynucleotides, such as on an array of polynucleotides as described herein, and generating copies of the polynucleotide or the array of polynucleotides, including the captured analyte. The analyte is typically RNA, preferably mRNA. The analyte is typically captured using a polyT sequence as the analyte capture sequence. For example, the method may comprise capturing RNA molecules on a polynucleotide, or an array of polynucleotides, such as an array of polynucleotides as described herein, performing reverse transcription of captured RNA molecules using a reverse transcriptase and primed using the captured RNA molecules to generate cDNA polynucleotides having 3’ end non-templated nucleotides; annealing a set of template switch oligonucleotides (TSOs) to the 3’ end non-templated nucleotides; and amplifying the cDNA.

[0179] The capture of analyte can be performed under any conditions known to the skilled person. For example, the method may be performed as part of a single cell sequencing experiment. Reverse transcription may be performed using any means known to the skilled person. Generating copies of the polynucleotide may be performed by any means known to the skilled person, such as by reverse transcription, rolling circle amplification and / or PCR amplification. Following the generation of copies of the polynucleotide / cDNA, sequencing may be performed as described herein.

[0180] As noted above, the methods comprise aligning the read and / or its reverse complement to a reference sequence and / or its reverse complement and predicting the positions of the BC and / or the UMI based on the alignment. In some cases, the BC and / or the UMI is not part of the reference sequence, and the position of the BC and / or the UMI is based on the relative position from a part of the read that has been aligned to the reference sequence, for example the relative position from the first VRFS and / or the second VRFS. In other cases, the BC and / or the UMI is included in the reference sequence, e.g. as an undefined base, and the positions of the BC and / or UMI are predicted based on which residues best align with the BC and / or UMI in the reference sequence. In some cases, a combination of these two approaches is used.

[0181] In some cases, the reference sequence does not comprise the BC and / or the UMI, and the alignment may be used to determine the portion of the read which corresponds to at least the first VRFS and / or the second VRFS. If the user then knows the position of the UMI and / or the BC relative to the first VRFS and / or the second VRFS, they can determine which portion of the read corresponds to the UMI and / or the BC. For example, if a read corresponds to a polynucleotide known to have (e.g. was synthesised in such a way that it is likely to have) a first VRFS of sequence ATGC, and a BC which is 5 nucleotides in length and immediately 5’ of the first VRFS, and the read has the sequence AATTGTG7G7CGATGC, the user would understand that the BC has the sequence TGTCG as this corresponds to the 5 nucleotides 5’ to the sequence which aligns to the VRFS, i.e. the ATGC sequence. Thus, the method of delimiting a BC and / or a UMI may comprise a further step of determining the position of (nucleotides corresponding to) the BC and / or the UMI based on the position of at least the first VRFS and / or the second VRFS. The method may further comprise determining the position of (nucleotides corresponding to) to BC and / or the UMI based on the position of the PCR handle and / or the analyte capture sequence. This may be referenced herein as the ‘positional’ method. In the Examples, the ‘spacer’ approach is a version of the ‘positional’ method that utilises a VRFS.

[0182] In addition the user may wish to carry out a multi-step positional approach. For example, the user may firstly align the read to a reference sequence comprising the PCR handle, and then align the read to a reference sequence comprising a VRFS (such as the first VRFS or the second VRFS), and predict the positions of the UMI and / or the BC based on their predicted positions relative to the PCR handle and the VRFS. Thus, the method may further comprise additional aligning steps in which the read is firstly aligned to a reference sequence comprising the first VRFS or second VRFS, and then the read is aligned to a second reference sequence, such as a reference sequence comprising or consisting of the PCR handle, and the positions of the BC and / or the UMI are predicted based on all the alignments.

[0183] In other cases, the reference sequence may comprise the BC, the UMI and at least one VRFS. For example, the reference sequence may comprise the BC, the UMI and the first VRFS. The reference sequence may comprise the BC, the UMI and the second VRFS. The reference sequence may comprise the BC, the UMI, the first VRFS and the second VRFS. The sequences of the BC and the UMI are not typically known. Instead, the reference sequence may comprise a placeholder N nucleotide at each position of the BC and / or the UMI, where N can be A, C, G or T. Where the BC and / or the UMI comprises a limited variable base or a constant nucleotide, the sequence of the BC and / or the UMI in the reference sequence may be amended accordingly. For example, if the BC were to comprise A, C or G in position 1, then a V nucleotide could be added to position 1 of the BC in the reference sequence.

[0184] In such cases, the reference sequence may comprise AT GCNNNNNN where ATGC is a VRFS and NNNNNN corresponds to the positions of the predicted BC or UMI. If the sequence that aligns to the reference sequence is ATGCG' / 'G' / 'CGTTT. the user understands that GTGTCG is the BC or UMI as it aligns best with the NNNNNN portion of the reference sequence. In such cases, the methods may further comprise a step of delimiting the BC and / or the UMI as the positions corresponding to the positions that best align with the BC and / or the UMI in the reference sequence. This may be referenced herein as the ‘full alignment’ method.

[0185] In some cases, the method involves a combination of the ‘positional’ and ‘full alignment’ methods. For example, the method may involve a ‘full alignment’ of the UMI and the flanking elements (e.g. a first VRFS and a second VRFS), in combination with a prediction of the first and last nucleotides of the UMI based on the ‘positional’ method.

[0186] The reference sequence comprises at least one VRFS. Where a first VRFS and a second VRFS is present in the polynucleotide or its reverse complement, the reference sequence may comprise the first VRFS and the second VRFS. The reference sequence may comprise two or more of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence. For example, the reference sequence may comprise (A) and (B); (A) and (C); (B) and (D); (C) and (D); (A), (B) and (C); (A), (B) and (D); (B), (C) and (D); or (A), (B), (C) and (D). The positions of the BC and / or the UMI in the read may be predicted based at least in part on the positions of two or more of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence in the reference sequence. For example, the positions of the BC and / or the UMI in the read may be predicted based at least in part on the positions of all of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence in the reference sequence.

[0187] Adding additional features, such as the PCR handle, and / or the analyte capture sequence to the reference sequence can improve the delimitation (identification) of the BC and / or the UMI as they can (for example) improve the alignment to the reference sequence. For example, in cases where a UMI has a deletion error, this may be more readily identified by alignment to a reference sequence that comprises an upstream feature (like a VRFS) and a downstream feature (like an analyte capture sequence). For example, a reference sequence could be AT GC QV QV QVTTT . If the read has the sequence ATGCA4CCGTTT, the alignment to the reference sequence would show that the first 4 nucleotides were likely to correspond to the VRFS and the final 3 nucleotides to the analyte capture region, thus the UMI comprises AACCG, but with a single nucleotide deletion.

[0188] The reference sequence may further comprise the PCR handle and / or the analyte capture sequence (if present). For example, the reference sequence may comprise:

[0189] - the PCR handle, the BC, the UMI and the first VRFS;

[0190] - the PCR handle, the BC, the UMI and the second VRFS;

[0191] - the PCR handle, the BC, the UMI, the first VRFS and the second VRFS;

[0192] - the analyte capture sequence, the BC, the UMI and the first VRFS;

[0193] - the analyte capture sequence, the BC, the UMI and the second VRFS;

[0194] - the analyte capture sequence, the BC, the UMI, the first VRFS and the second VRFS;

[0195] - the PCR handle, the BC, the UMI, the first VRFS and the analyte capture sequence; - the PCR handle, the BC, the UMI, the second VRFS and the analyte capture sequence; or

[0196] - the PCR handle, the BC, the UMI, the first VRFS, the second VRFS and the analyte capture sequence.

[0197] The aligning step may be performed using any means known to the skilled person. For example, the aligning step may be performed using the Smith -Waterman algorithm. The aligning step may be performed using the Needleman-Wunsch algorithm. The alignment may be a local or a global sequence alignment.

[0198] The methods may further comprise delimiting (or identifying) the positions of the PCR handle sequence, the BC, the UMI, the first VRFS, the second VRFS and / or the analyte capture sequence in the read and / or its reverse complement. The positions are delimited based on the alignment. That is to say, the position and sequence of the PCR handle sequence, the BC, the UMI, the first VRFS, the second VRFS and / or the analyte capture sequence in the read and / or its reverse complement may be determined.

[0199] Typically, where more than one sequence element is included in the reference sequence, the alignment is weighted. Methods of specifying different weights to different sequence elements during an alignment are well-known to the skilled person. By specifying different weights to different sequence elements, the strength of the alignment can be improved by essentially adding less credibility to sequence elements where the alignment is expected to be poor. For example, where the analyte capture region is a polyT sequence and the reference sequence comprises the analyte capture region, the analyte capture region can be assigned a relatively low weighting when compared to other sequence elements in the reference sequence. This is because polyT sequences may be sequenced with low fidelity, introducing variability into the length of the polyT sequence in the read when compared to the reference sequence. If the reference sequence comprises the BC and / or the UMI, the alignment may utilise a penalty in cases where the BC and / or the UMI is not the correct (predicted) length, e.g. if the UMI was synthesised to be 9 nucleotides long and the alignment suggests a 10 nucleotide UMI, then the alignment may be considered less likely to be accurate and a penalty applied. Sequence elements with a fixed length and known sequence may be assigned a relatively high weighting when compared to other sequence elements in the reference sequence. For example, the VRFS(s) and / or the PCR handle sequence may be assigned a relatively high weighting when compared to the analyte capture sequence (e.g. polyT) and / or a UMI. The VRFS(s), the BC and / or the PCR handle sequence may be assigned a relatively high weighting when compared to the analyte capture sequence (e.g. polyT) and / or a variable length UMI. The VRFS(s), the BC, the UMI (where it is a fixed-length UMI) and / or the PCR handle sequence may be assigned a relatively high weighting when compared to the analyte capture sequence. The VRFS(s) and / or the PCR handle sequence may be assigned a relatively high weighting when compared to the BC, the UMI and / or analyte capture sequence.

[0200] In the methods described herein, it is typically important to include an alignment of the reverse complement of the read and / or the reverse complement of the reference sequence. As approximately 50% of reads produced after PCR amplification will be of the other sense, said reads will align poorly to the reference sequence, but the reverse complement of said reads will align better to the reference sequence (or the reads will align better to the reverse complement of the reference sequence). Accordingly, the method further may comprise a step of comparing the quality of two or more of (A) the read to the reference read, (B) the read to the reverse complement of the reference read, (C) the reverse complement of the read to the reference read, and (D) the reverse complement of the read to the reverse complement of the reference read. The method may also comprise a step of selecting the alignment of higher quality. This allows the methods to retain close to 100% of the reads that comprise the polynucleotide or its reverse complement.

[0201] In some cases, the quality is determined by assigning a score for the alignment of the PCR handle sequence, a score for the alignment of the BC, a score for the alignment of the UMI, a score for the alignment of the first VRFS, a score for the alignment of the second VRFS and / or a score for the alignment of the analyte capture sequence. The scope for the alignment may be determined by (A) the percentage of matches in each aligned section; or (B) an exponential of the percentage of matches in each aligned section. The quality may be determined by assigning a score for two or more of the alignment of the PCR handle sequence, the alignment of the BC the alignment of the UMI, the alignment of the first VRFS, the alignment of the second VRFS and the alignment of the analyte capture sequence, and taking a weighted average of the two or more scores. An exemplary weighted average is discussed in more detail in the Examples. The weights must add up to 1. For example, where the reference sequence comprises a PCR handle, a BC, a first VRFS, a UMI, a second VRFS and a polyT analyte capture region, exemplary weights may be 0.25, 0.2, 0.2, 0.075, 0.2 and 0.075, respectively. It will be understood that the exact weightings are not particularly limited and may be selected by the skilled person based on the identity of the sequence elements within the reference sequence.

[0202] The method typically comprises a step of determining a sequence for the BC and / or the UMI. The method typically comprises a step of determining a sequence of the analyte (i.e. after capture, reverse transcription and / or amplification). The method typically further comprises annotating the sequence of the analyte with the BC and / or the UMI. This enables the source of the analyte to be determined, e.g. as deriving from a single cell as another analyte and / or from the same analyte (e.g. mRNA) molecule as another analyte.

[0203] The method typically comprises a step of determining a sequence of the analyte (i.e. after capture, reverse transcription and / or amplification). The method may further comprise associating the sequence of the analyte with the sequence of other analytes comprising (a) the same BC, (b), the same UMI, and / or (c) the same BC and the same UMI. In some cases, it is not necessary to determine the sequence of the BC and / or the UMI. The method may comprise associating reads with (a) the same BC, (b) the same UMI, and / or (c) the same BC and the same UMI.

[0204] Where the sequence of the analyte is annotated with, or associated with, the BC and / or the UMI, the method may further comprise a step of aggregating reads with the same BC, the same UMI, and the same analyte sequence. Whether a BC or UMI is the same may be determined using the methods described herein, for example by the whitelisting and / or error correction methods described herein. An analyte sequence of a read may be considered the same if it comprises at least 50% sequence identity to the analyte sequence of another read, such as at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 90% sequence identity, at least 95% sequence identity or at least 95% sequence identity. This is because sequencing technologies, in particular long-read sequencing technologies, can introduce errors in sequencing reads. Furthermore, reads derived from the same analyte molecule may derive errors from PCR amplification. In some cases, the method may further comprise a step of determining a consensus analyte sequence for reads with the same BC, the same UMI and the same analyte sequence. A consensus sequence may be determined by aligning reads with the same BC, the same UMI and the same analyte sequence, and by assigning for each position in the analyte sequence, the nucleotide that is most abundant.

[0205] The sequence of the analyte may be determined as the sequence 3’ of the portion of the polynucleotide comprising the BC and the UMI, or may be determined as the sequence 3’ of the analyte capture sequence. The sequence of the analyte may be determined as the sequence 3’ of the reference sequence (or its reverse complement). The sequence of the analyte may be further determined based on sequences that are 3’ to the analyte, e.g. 5’ of the sequence of a template-switch oligonucleotide (TSO) sequence. Accordingly, in some embodiments, the method further comprises a step of alignment the read (or its reverse complement) to the sequence of a TSO (or its reverse complement).

[0206] The method may further comprise one or more steps to detect one or more errors in the read. The presence of errors in the read may be acceptable (e.g. are tolerated) and means that the read may be retained for further analysis. In other cases, the presence of errors in the read is not acceptable and the read may be discarded (i.e. removed from the dataset). In some cases, the presence of errors in reads leads to the reads being binned (otherwise referred to herein as being aggregated), e.g. for further processing, before deciding whether the read can be retained or should be discarded. Where errors occur outside of the BC and / or UMI and / or the analyte sequence, these can often be tolerated as they do not necessarily affect the incorrect assignment of a read to a particular source. Where errors occur in the BC and / or the UMI, it may be possible to correct the errors. These error correction steps allow for reads that may be incorrectly assigned to a particular source, e.g. using the BC and / or UMI sequences that have been determined to be removed from a dataset. These steps may also allow for reads with errors to be corrected and assigned to a particular source, e.g. using the BC and / or the UMI, to thereby prevent unnecessarily loss of reads from the data set.

[0207] It is important to identify errors in the BC and / or UMI, as these errors could lead to incorrect assignment of the source of the read. It may also be possible to correct errors in the BC and / or the UMI to correctly assign the source of the read and retain it in the dataset. Accordingly, the method may further comprise comparing the predicted length of the BC and / or the UMI in the read to the length of the BC and / or the UMI in the reference sequence, and determining that there is an error in the read if the predicted length of the BC and / or the UMI differs from the length of the BC and / or the UMI in the reference sequence. Where the BC and / or the UMI comprises nucleotide blocks, the method may further comprise comparing the predicted number of nucleotide blocks in the BC and / or the UMI in the read to the number of nucleotide blocks of the BC and / or the UMI in the reference sequence, and determining that there is an error in the read if the number of nucleotide blocks in the predicted BC and / or the UMI differs from the number of nucleotide blocks in the BC and / or the UMI of the reference sequence. The method may further comprise binning the read (i.e. aggregating reads containing errors in the BC and / or the UMI optionally for further processing) if the number of nucleotides or nucleotide blocks in the predicted BC and / or the UMI differs by at least one nucleotide or nucleotide block from the number of nucleotides or nucleotide blocks in the BC and / or the UMI in the reference sequence. The binned reads may be processed as described further herein to correct the error(s) or to discard the read. In some cases, the method comprises discarding the read if the number of nucleotides or nucleotide blocks in the predicted BC and / or the UMI differs by at least a specified number of nucleotides or nucleotide blocks in the reference sequence. This may be because the discrepancy in length indicates that it is unlikely that the BC and / or the UMI can be corrected. The specified number of nucleotides or nucleotide blocks that will result in the discarding of the read may vary depending on the length of the BC and / or of the UMI. The specified number of nucleotides or nucleotide blocks may vary depending on the diversity of the BC and / or of the UMI. For example, the specified number may be at least 3, such as at least 4 or at least 5 nucleotides or nucleotide blocks.

[0208] Where the BC and / or the UMI comprise a series of discrete nucleotide blocks, wherein the discrete nucleotide blocks are from a mixed pool of nucleotide blocks of known sequence, and wherein each nucleotide block sequence in the pool differs from each other nucleotide block in the pool by at least two nucleotide substitutions, the method may further comprise comparing the sequence of each nucleotide block to the sequence of each nucleotide block in the pool, and determining that there is an error in the read if the sequence of a nucleotide block in the read is not present in the pool. The method may further comprise binning the read (i.e. aggregating reads containing errors in the BC and / or the UMI) if at least one nucleotide block in the read is not present in the pool. The binned reads may be processed as described further herein to correct the error(s) or to discard the read. In some cases, the method comprises discarding the read if the number of nucleotide blocks in the read that are not present in the pool exceeds a specified number. This may be because the number of errors is such that it is unlikely that the BC and / or the UMI can be corrected. The specified number that will result in the discarding of the read may vary depending on the length of the BC and / or of the UMI. The specified number may vary depending on the diversity of the BC and / or of the UMI. For example, the specified number may be at least 3, such as at least 4 or at least 5 nucleotide blocks.

[0209] The method may further comprise comparing the predicted length of the analyte capture region in the read or its reverse complement to the length of the analyte capture region in the reference sequence and determining that there is an error in the read or its reverse complement if the predicted length of the analyte capture region differs from the length of the analyte capture region in the reference sequence. The analyte capture sequence is typically a polyT sequence. However, errors in the analyte capture sequence can often be well-tolerated or entirely ignored, e.g. if a high number of errors are introduced during sequencing of the analyte capture sequence (e.g. a polyT). Furthermore, errors in the analyte capture region sequence itself are relatively immaterial to the data set due to the lack of impact on the BC / UMI (and thus the identified source of the read) and the identity of the analyte. However, large numbers of errors in the analyte capture sequence could be indicative of low sequencing fidelity and thus of a low quality read that should be discarded. Accordingly, the method may comprise discarding the read if the predicted length of the analyte capture region differs by at least a specified number of nucleotides from the length of the analyte capture region in the reference sequence. The analyte capture region may be a polyT sequence, such as T30. The specified number may a percentage of the total length of the analyte capture region. For example, the specified number may be at least 90% of the total length of the analyte capture region in the reference sequence, such as at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80% of the total length of the analyte capture region in the reference sequence. The specified number may be at least 10, at least 15, at least 20, at least 25 or at least 30 nucleotides. The specified number may vary depending on the length of the analyte capture region and / or the identity (e.g. the sequence) of the analyte capture region.

[0210] The method may further comprise comparing the sequence of at least two of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence in the read or its reverse complement to the reference sequence, and discarding the read if greater than a specified number of the nucleotides in the sequences that were compared are different to the corresponding portions of the reference sequence. Preferably, the method may further comprise comparing the sequence of at least two of (A) the PCR handle, (B) the first VRFS, and (C) the second VRFS, in the read or its reverse complement to the reference sequence, such as all three of (A) to (C), and discarding the read if greater than a specified number of the nucleotides in the sequences that were compared are different to the corresponding portions of the reference sequence. The method may comprise comparing the sequence of at least three, or all four, of (A) to (D) in the read and / or its reverse complement to the reference sequence, and discarding the read if greater than a specified number of the nucleotides in the sequences that were compared are different to the corresponding portions of the reference sequence. For example, the specified number may be at least 50% of the total length of the sequences that were compared in the reference sequence, such as at least 40%, at least 30%, at least 20%, at least 10% or at least 5% of the total length of the sequences that were compared in the reference sequence. The specified number may be at least 5, at least 10, at least 15, or at least 20. The specified number may vary depending on the length of the sequences that were compared and / or their identity (e.g. their sequence).

[0211] The method may further comprise comparing the sequence of the reference sequence to the sequence of the read or its reverse complement, and discarding the sequence if greater than a specified number of the nucleotides in the sequence that was compared is different to the corresponding portion(s) of the reference sequence. The specified number may vary depending on the length and identity of the reference sequence. The specified number may be at least 50% of the total length of the sequences that were compared in the reference sequence, such as at least 5%, at least 10%, at least 20%, at least 30%, or at least 40%. The specified number may be at least 5, at least 10, at least 15 or at least 20. When determining whether a nucleotide is different, a nucleotide assigned to the BC and / or UMI region is typically always considered a match (z.e. the same) unless the position in the BC and / or UMI comprises a limited variable base or a constant nucleotide and the nucleotide does not match. Where an extra nucleotide is added to, or a nucleotide is missing from the sequence of the read when compared to the reference sequence, this is typically considered a mismatch (i.e. a difference).

[0212] In some cases, the polynucleotide comprises a limited variable base between the analyte capture region and the portion of the polynucleotide comprising the UMI, BC, PCR handle and / or first VRFS. In this case, the method may comprise discarding the read if the limited variable base is not identified in the sequence of the read or its reverse complement. In some cases, the polynucleotide comprises a limited variable base between the analyte capture region and the portion of the polynucleotide comprising the PCR handle and / or first VRFS.

[0213] The methods described herein may further comprise correcting errors in the BC and / or the UMI. For example, the methods may further comprise binning reads with errors in the BC and / or UMI as described above, and correcting errors in the BC and / or UMI. This may be done by whitelisting of barcodes or by other means (for example based on the deduplication of nucleotide blocks as described in WO 2022 / 118027). Means of correcting errors in the BC and / or UMI are described below.

[0214] The method may further comprise comparing the barcode sequence to a whitelist of known barcode sequences. Whitelisting of barcode sequences may be performed by any means known to the skilled person or as described herein. For example, the method may comprise generating a whitelist of barcode sequences using a knee-plot. In some cases, the method may comprise generating a whitelist of barcode sequences in real time, comprising: associating barcodes of similar sequence; ranking barcodes based on the number of reads of each barcode sequence, and assigning the highest-ranking barcode of each group of associated barcodes to the whitelist. A pair of barcodes may be considered as comprising a similar sequence if they have a Levenshtein edit distance of 5 or less, such as 4 or less, 3 or less, 2 or less or 1. A pair of barcodes may be considered as comprising a similar sequence if they have a Levenshtein edit distance / Hamming Distance of 1 or less, typically 0.5 or less, 0.4 or less, 0.3 or less, 0.2 or less or 0.1 or less. If the barcode sequence is not in the whitelist, the method may comprise correcting the barcode sequence by assigning it a new sequence corresponding to the barcode sequence in the whitelist which is the nearest match. In some cases, the barcode sequence is assigned the new sequence if the barcode sequence in the whitelist matches with a Levenshtein edit distance of 4 or less, such as 3 or less, 2 or less or 1. In some cases, the barcode sequence is assigned the new sequence if the barcode sequence in the whitelist matches with a Hamming distance of 6 or less, such as 5 or less, 4 or less, 3 or less, 2 or less or 1. In some cases, if the barcode sequence is not corrected as described herein, the read is discarded. This is because the read will not be able to be correctly assigned a source. In some cases where the barcode sequence is not in the whitelist, it is associated with the barcode sequence in the whitelist which is the nearest match, as described above, to assign reads from the same source.

[0215] In some cases, it is not necessary to calculate the Hamming distance and / or Levenshtein edit distance between a barcode sequence from the read and a barcode sequence from the whitelist. Any suitable method may be used to identify the differences between a barcode sequence from the read and a barcode sequence from the whitelist, and to either (i) assign the barcode sequence of the read to the barcode sequence of the nearest match on the whitelist, or associate thee barcode sequence of the read to the barcode sequence of the nearest match on the whitelist, or (ii) discard the read, based on the differences (e.g. the extent or type of the differences) between the barcode sequence of the read and that of the nearest match on the whitelist. The nearest match may be determined by any suitable means.

[0216] The method may further comprise identifying two or more reads with the same BC, and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with an edit distance (for example a Levenshtein edit distance or Hamming distance) of 5 or less, such as 4 or less, 3 or less, 2 or less or 1. In some cases, the method may further comprise identifying two or more reads with the same BC, and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with a Levenshtein edit distance / Hamming Distance (i.e. Levenshtein edit distance divided by Hamming distance) of 1 or less, typically 0.5 or less, 0.4 or less, 0.3 or less, 0.2 or less or 0.1 or less. Preferably, the method further comprises identifying two or more reads with the same BC and analyte sequence, and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with an edit distance (for example a Levenshtein edit distance or Hamming distance) of 5 or less, such as 4 or less, 3 or less, 2 or less or 1. In some cases, the method may further comprise identifying two or more reads with the same BC and analyte sequence, and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with a Levenshtein edit distance / Hamming Distance of 1 or less, typically 0.5 or less, 0.4 or less, 0.3 or less, 0.2 or less or 0.1 or less.

[0217] In some cases, the reads are derived from an array of polynucleotides comprising UMIs whose predicted lengths are not all identical (e.g. if the UMI length is higher or lower than intended due to synthesis or sequencing errors). In this case, Hamming distance is a less useful metric as different UMIs may comprise an identical series of nucleotides or nucleotide blocks except for the ‘extra’ nucleotides or nucleotide blocks in the longer sequence. In some cases, the method may further comprise identifying two or more reads with the same BC (and preferably the same analyte sequence), and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with a Levenshtein edit distance of 5 or less, such as 4 or less, 3 or less, 2 or less or 1, or with a Levenshtein edit distance / Hamming Distance of 1 or less, typically 0.5 or less, 0.4 or less, 0.3 or less, 0.2 or less or 0.1 or less.

[0218] Also disclosed is a method for delimiting barcode sequences and / or unique molecular identifier sequences in a library of reads, each read corresponding to the sequence of at least a portion of a polynucleotide, wherein each polynucleotide or its reverse complement comprises: (a) a PCR handle sequence; (b) a barcode sequence (BC); (c) a unique molecular identifier sequence (UMI); (d) a first variable region flanking sequence (VRFS) and / or a second VRFS, wherein the first VRFS is between the BC and the UMI, wherein the second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI and / or between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI; and (e) optionally, an analyte capture sequence, and wherein the method comprises, for each read, performing the method as described herein.

[0219] In some cases, the method described herein is a computer-implemented method. Also disclosed herein is a computer readable medium comprising instructions for performing the method as described herein. Also disclosed is a computer comprising the computer readable medium.

[0220] Methods for generating the array of, or library of polynucleotides

[0221] The invention also provides a method of making the array of polynucleotides described herein. Typically, the method comprises using split-and-pool polynucleotide synthesis to make the BC of each polynucleotide of the array. The method typically comprises using degenerate polynucleotide synthesis to make each UMI of the array. The polynucleotides of the array may be synthesised in the 5’ to 3’ direction, or in the 3’ to 5’ direction. The polynucleotides of the array may be synthesised directly on microbeads, in the 5’ to 3’ direction, or in the 3’ to 5’ direction. The polynucleotides of the array may be made in sections and joined, for example, using splint ligation. This is particularly useful where the polynucleotides of the array are long, for example 100 nucleotides in length or greater, such as 150 nucleotides in length or greater. This is also useful where different features, such as different VRFSs, PCR handles, analyte capture regions and / or hairpins are added to the polynucleotides of the array.

[0222] The invention also provides a method of producing a library of polynucleotides, wherein the polynucleotides are amplified from the polynucleotides of a sample and / or tag nonpolynucleotide analytes of a sample, the method comprising (a) capturing analytes in the sample on an array of polynucleotides, a micro-particle or a plurality of micro-particles as described herein; (b) generating copies of the array of polynucleotides, including any (I) polynucleotides of the sample and / or (II) polynucleotides that tag non-polynucleotide analytes of the sample, captured by the array polynucleotides; and (c) amplifying the number of copies of each polynucleotide to produce a library of polynucleotides amplified from or tagging analytes in the sample.

[0223] The method may comprises contacting the sample with a surface / solid support, microparticle or an array of polynucleotides as described herein and allowing analytes in the sample to bind to the 3’ and / or 5’ analyte capture regions of the array of polynucleotides.

[0224] In some cases, the analytes are from a single cell, such as a bacterial cell, or single cell nucleus, single cell vesicle, such as an exosome or mini-vesicle, or other compartment enclosed by a lipid membrane. In some cases, the method comprises isolating the single cell, nucleus, vesicle or other compartment, or a lysate thereof, and contacting the isolated cell, nucleus, vesicle, compartment or lysate with a single micro-particle or array of polynucleotides as described herein.

[0225] In some cases, the analytes are from a sample comprising a plurality of cells (such as bacterial, prokaryotic or eukaryotic cells), cell nuclei, vesicles or other compartment enclosed by a lipid membrane, or a two or three dimensional sample such as a tissue sample. Analytes from different cells, cell nuclei, vesicles or other compartments or from different spatial positions in the sample may be captured by separate sub-arrays of polynucleotides as described herein. The method may in some cases include preparing a single cell or single cell nuclei or vesicle suspension from the sample and / or isolating single cells, nuclei, vesicles or lysates thereof in separate compartments. A single cell, nuclei, vesicle or cell / nuclei / vesicle lysate of the sample and a single micro-particle may be isolated in each of a plurality of separate compartments. In some cases, there may be at least 500, or at least 5,000, or 25,000, or 50,000, or 100,000, for example between 500 and 500,000, or between 2,000 and 200,000, or between 20,000 and 100,000 separate compartments, each comprising a single cell, nuclei or cell / nuclei lysate and a single micro-particle or (sub-)array of polynucleotides as described herein. The method may in in other cases include preparing a two or three dimensional sample for contact with a surface / solid support as described herein.

[0226] Micro-particles or (sub-)arrays of polynucleotides may be contacted with sample or analyte in a compartment, typically a fluidic compartment. For example, the compartment may be a well, such as a well in a multi-well plate or microplate or slide, a discrete site / position on a microfluidic chip, or a (micro-)droplet, which may be formed in an oil emulsion. Such compartments may in some cases be made using a microfluidics device as known in the art. In some cases, the sample, a micro-particle, an array of polynucleotides, a cell, cell nuclei or cell / nuclei lysate may be encapsulated or coencapsulated in a fluidic compartment.

[0227] The cell and / or cell nucleus membrane(s) may be lysed, before or after contact with the micro-particle or array of polynucleotides. In some cases, the cell and / or nucleus may be lysed and target polynucleotides may be amplified, for example by PCR, before being brought into contact with the microparticle or array of polynucleotides. The analytes that bind to the 3’ and / or 5’ analyte capture regions may then include the PCR products. In some cases, a linker of the polypeptides may be cleaved or a micro-particle hydrogel may be dissolved to separate the polypeptides from (the bead of) a micro-particle or from a surface / solid support before or after contact with the cell, nucleus or lysate, for example by exposing the micro-particle to UV light or heat, or contacting the micro-particle with appropriate chemicals or enzyme(s), such as the enzyme(s) described herein. In some cases, it may be convenient to include the micro-particle or array or polynucleotides in cell / membrane lysis buffer, and / or the cell, nucleus or lysate in a buffer comprising agents needed for cleavage of the linker, such that mixing the two buffers exposes cell / nucleus to lysis buffer and / or micro-particle to the chemical milieu for linker cleavage, as well as bringing the array of polynucleotides into contact with analyte from the cell or nuclei. For example, a microfluidics device may be used to join two aqueous flows into discrete microfluidic droplets. One flow may comprise a single cell or single cell nuclei suspension in cell buffer and optionally chemicals or enzymes needed to cleave the polynucleotide linker. The other flow may comprise a suspension of microparticles as described herein, optionally in cell / membrane lysis buffer. Some of the droplets that are formed comprise both a cell or nuclei and a micro-particle, resulting in contact between the analytes and the array of polynucleotides.

[0228] Other components needed for downstream reactions and processes may be included in the cell / nucleus buffer and micro-particle / lysis buffer or be added to or used to wash the plate or other surface / solid support. In some cases, the cell / nucleus buffer may comprise template switch oligonucleotides. In some cases, the micro-particle / lysis buffer may comprise reverse transcriptase, in particular when the 3’ analyte capture region of the polynucleotides is an RNA capture region, such as a polythymidine.

[0229] RNA in the sample may bind to the 3’ RNA capture regions of the polynucleotides and reverse transcription using the bound RNA as template provides an RNA / cDNA hybrid at the 3’ end of the polynucleotide. Template switch oligonucleotides may be added at the end of the RNA / cDNA hybrid. The template switch oligonucleotides and the PCR handle sequence 5’ to the identifier sequence(s) (e.g. BC and / or UMI) provide a pair of PCR handle sequences for PCR amplification of the cDNA / RNA hybrid. PCR amplification may use a pair of oligonucleotide primers that hybridise to the template switch oligonucleotide sequence and the PCR handle sequence 5’ to the identifier sequence(s).

[0230] DNA in the sample may bind to a 5’ and / or 3’ DNA capture region of the polynucleotides. DNA captured at the 5’ end may be amplified using oligonucleotide primers that hybridise to a PCR handle sequence on the captured DNA and the complement of the PCR handle sequence 3’ to the identifier sequence(s) (e.g. BC and / or UMI). DNA captured at the 3’ end may be amplified using oligonucleotide primers that hybridise to a PCR handle sequence on the captured DNA and the PCR handle sequence 5’ to the identifier sequence(s).

[0231] When the array polynucleotides comprise both 3’ and 5’ analyte capture regions, a first round of PCR amplification using all of the relevant PCR oligonucleotide primers will generate PCR products comprising the identifier (e.g. the barcode and UMI) sequences and polynucleotide sequence from polynucleotide analyte bound to the 3’ or 5’ analyte capture regions. A third PCR product comprising the identifier (e.g. barcode and UMI) sequences, but not any sequence from bound analytes, will also be generated from the pair of primers that hybridise one to the PCR handle sequence 5’ to the barcode and UMI sequences and the other to the complement to the PCR handle sequence 3’ to the barcode and UMI sequences. One or more further rounds of PCR amplification using only one or the other pair of PCR oligonucleotide primers described above eliminates this additional PCR product from further amplification.

[0232] After amplification, the PCR products may be sequenced using methods known in the art. Barcoding means that different compartments containing different polynucleotide arrays bound to analyte, or reaction products thereof, may be merged after capture. For example, the different compartments, such as droplets, may be merged after reverse transcription and / or before PCR. Downstream processes may then be carried out in bulk. PCR product sequences having different barcodes can be digitally assigned to different samples or sample parts or sub-elements (for example cells or cell nuclei), as described herein. Analytes can be digitally counted by using the UMI sequences to identify duplicates derived from the same capture event.

[0233] The method may comprise further performing the method of delimiting the BC and / or the UMI of a read, as described herein. The invention also provides a library of polynucleotides obtained or obtainable by the method of producing the library of polynucleotides. The invention also relates to a method of generating a library of reads, comprising sequencing the library of polynucleotide produced by the method described herein, or the library of polynucleotides described herein.

[0234] Examples

[0235] Example 1 - Evidence of bead truncation in 10X chromium and drop-seq beads.

[0236] Droplet based single-cell methods process mRNA from individual cells encased in oil droplets in a highly parallel fashion. Here, cells are trapped within the droplet, which facilitate cell lysis and mRNA capture using oligonucleotide beads. However, bead designs vary between Dropseq and 10X Chromium (Fig. la). In Dropseq, the bead contains a PCR primer region, followed by a 12 base pair (bp) cell barcode created by a split and pool synthesis. Next is an 8bp Unique Molecular Identifier (UMI) sequence. A V base is also present preceding the start of a poly(dT) capture region that acts to bind the polyA tail of the mRNA. In contrast, 10X chromium bead design is distinct. It contains a 16bp barcode formed though combinatorial enzymatic ligation split and pool, followed by a 12bp UMI sequence. Unlike Dropseq, the 10X Chromium beads lack a V base between the UMI and the poly(dT) sequence. They uniquely incorporate a VN base at the poly(dT) end, ensuring mRNA capture near the polyA terminus.

[0237] To explore the possibility of synthesis errors leading to discarded reads, we analysed public datasets from 10X Genomics Illumina sequencing. Our initial assessment highlighted an elevated percentage of T bases in readl, corresponding to the barcode and UMI sequences. Particularly striking was the significant rise in T bases at the UMI's concluding base, potentially indicating truncated beads due to reading into the poly(dT) region (Fig. lb.). This trend was even more pronounced in long-read sequencing data (Fig. 1c). The heightened T base presence might be attributed to the different computational techniques used for analysis. In this method, after identifying the PCR primer, the barcode and UMI's location are ascertained based on their expected downstream position. Considering the commonality of indel errors from PCR, sequencing, or synthesis, this positional method might account for the observed T base increase in long-read data. When examining Dropseq, a distinct pattern emerged. We similarly observed an increase in the proportion of T throughout the UMI sequence in Illumina (Fig. Id) and ONT sequencing (Fig. le). Notably, neither Illumina nor ONT sequencing showed a heightened presence of T base at the final UMI position. The presence of the V base between the UMI and the poly(dT) capture sequence likely aids in a more accurate identification of the UMI sequence, explaining this observation.

[0238] We further validated these observations by producing an ONT sequencing library directly from the beads (Fig. 2a, b). This highlighted an anticipated dominant peak size of 28 bp for the 10X genomics data and 20 bp for Dropseq. Notably, a subset of beads had one or more deletions, while a fewer number had one or two insertions, confirming synthesis errors within the beads. In summary, it's evident that bead truncation is a significant concern. This conclusion is reinforced when analysing the most frequent UMIs in both 10X short read and long-read datasets (Fig. 2c, d). A common pattern emerges: an excess of Ts, particularly towards the UMI's tail end.

[0239] Example 2 - Enhancing barcode and UMI identification with an intervening VRFS

[0240] 10X chromium beads are synthesised in a semi random process using enzymatic ligation, while Dropseq barcodes are generated in a split and pool and have a higher complexity. It was found that while version 2 beads have a minimum edit distance of 2, edit distance of less than 2 is observed within the version 3 beads. This makes whitelisting for correcting barcodes more problematic, increasing the chance that reads are computationally whitelisted to the wrong cell. However, this is likely only problematic if the data suffers from high levels of errors. To investigate the effect of errors on barcode assignment we subsequently conducted a species mixing experiment in single-cell sequencing, employing both the 10X Chromium and Dropseq techniques and sequencing using the ONT platform. While truncation of the barcode might seem problematic, it was shown that it does not compromise the number of cells identified. This is due to the whitelisting correction process that can overcome any errors or truncations. To validate this, we computationally shortened the barcode sequence for both 10X Chromium and Dropseq methods. Our analysis showed no significant impact on discerning human and mouse cells, even when the barcode was shortened by one or two bases. Therefore, even with the observed barcode truncation, the whitelisting capability effectively mitigates potential complications stemming from synthesis, PCR, or sequencing errors. Given that UMI’s are inherently random, whitelisting cannot be employed for their correction. Consequently, any UMI error can artificially inflate the read count. We previously noted a pronounced increase in T bases at the UMI's tail in 10X data, indicative of truncation. Further probing the effect of shortening the UMI by a single base, which revealed 415 differentially expressed transcripts between UMIs of lengths 11 and 12. Notably, all these transcripts were upregulated in the dataset with a 12-length UMI compared to the 11 -length counterpart. This suggests that synthesis errors may result in bead sequence truncation, potentially skewing downstream gene expression analysis.

[0241] To summarise, our data underscores that while synthesis issues may not markedly hamper barcode delineation, they do have considerable implications for UMI identification. Our findings underscore the rationale for devising methods that can precisely detect truncated UMIs, a step pivotal for the reliability and efficiency of single-cell sequencing.

[0242] Example 3 - Inclusion of a VRFS improves the ability to detect the start and end of the UMI.

[0243] Having identified synthesis challenges associated with the bead oligonucleotides, we theorised that incorporating a VRFS between the barcode and UMI, and a V base between the UMI and the poly(dT) capture handle, could provide clearer demarcation of the UMI region. In our previous work (Sun et al. (2023)), we introduced the concept of a Common Molecular Identifier (CMI) sequence, designed to precisely quantify errors in a sequenced read. In the present study, we employed the CMI to evaluate merits of incorporating a VRFS (otherwise referred to as a spacer) within the oligonucleotide. To this end, we synthesised a free oligonucleotide that contained a PCR handle, a constant 12bp barcode region, succeeded by a 4bp VRFS, a 32 bp homodimer CMI, and finally a V base preceding the poly(dT) region (Fig. 3a). To pinpoint the CMI, we adopted two contrasting techniques. The first, known as the Positional Strategy, involved locating the end of the PCR handle through sequence alignment, then projecting the CMI's onset to be 16 base pairs distant from this terminus. Conversely, the ‘spacer’ Strategy discerned the start of the CMI by pattern-matching the VRFS sequence, thus identifying the CMI's initiation immediately post the VRFS’s conclusion. As used in the Examples and accompanying Figures, the VRFS is also referred to as a spacer sequence.

[0244] After capturing mRNA on a bulk scale and conducting ONT sequencing across three independent experiments (Fig. 3b), we observed a significant increase in the accurate identification of CMIs using the VRFS method compared to the positional one. This data indicates that incorporating a VRFS can boost the precision of determining the CMI's initial position. Such an advancement is probably attributed to effectively circumventing the inherent PCR, sequencing, and synthesis errors present within the barcode.

[0245] Example 4 - New bead design mitigates the issues seen with oligo truncation and coupling errors.

[0246] After demonstrating that incorporating a VRFS between the barcode and UMI enhances UMI identification, we integrated this feature into our Dropseq bead designs. Previously, we had developed an advanced Dropseq method named scCOLOR-seq (Philpott et al. Nat Biotechnol 39, 1517-1520 (2021); Philpott et al. Methods Mol Biol 2632, 259-267 (2023)), which employed error-correcting homodimer UMIs. By adding both a first VRFS and a V base to the capture oligonucleotide, we enhanced these beads (Fig. 4a). The first VRFS clearly demarked the boundary between the barcode and the UMI sequence (Fig. 4b). Our primary objective was to assess whether this modification increased the accuracy of UMI sequence detection.

[0247] Given the UMI's unique homodimer composition, we theorised that assessing perfect dimer nucleotide concordance throughout the UMI would be an effective validation metric for our improved strategy. If the homodimer concordance improved, it would suggest better UMI identification. We next compared the positional and ‘spacer’ approaches. The addition of the VRFS indeed increased the homodimer concordance rate (Fig. 4c), underscoring our ability to precisely determine the UMI's boundaries. This marked rise in the dimer concordance rate for the homodimer UMI bolstered our initial theory, highlighting that the VRFS’s inclusion genuinely augments the UMI sequence identification process. Example 5 - The inclusion of a VRFS improves UMI counts and number of transcripts detected.

[0248] We next tested the new bead design in a species mixing experiment. During the subsequent data analysis, we compared both the positional and ‘spacer’ techniques. Notably, implementing the VRFS-centric method (i.e. the ‘spacer’ technique) resulted in a significant increase in the number of detected UMIs (Fig. 5a) and transcripts (Fig. 5b) per cell relative to the positional strategy. Furthermore, more genes were identified within both human (Fig. 5c) and mouse (Fig. 5d, Fig. 5e and Fig. 5f) cells using the ‘spacer’ approach when compared to the positional approach. This suggests that integrating a VRFS in the bead oligonucleotide sequence effectively minimises artefacts, yielding a more precise tally of unique molecules across a broad spectrum of features. Additionally, the ‘spacer’ method revealed several genes that remained unnoticed with the positional approach.

[0249] In summary, our refined droplet-based single-cell oligonucleotide bead capture design, which integrates a VRFS between the barcode and UMI, as well as a V base between the UMI and the poly(dT) capture site, enhances UMI identification and optimises feature counting.

[0250] Discussion

[0251] Single-cell RNA-sequencing (scRNA-seq) technology is a rapidly advancing technology that is revealing new scientific findings. However, the data it generates can often contain numerous technical biases. Bead synthesis errors in droplet-based single-cell sequencing were examined. Several technical inaccuracies arising from imperfect synthesis of beadbound oligonucleotides were identified. To address these errors, an improved bead design was devised, effectively addressing the identified synthesis challenges. One of the pressing issues with scRNA-seq data is the amplification of technical errors, which introduces noise (Brennecke et al. Nat Methods 10, 1093-1095 (2013); Jia et al. Nucleic Acids Res 45, 10978-10988 (2017)), making gene counting unreliable and potentially skewing differential gene expression results. The ideal scenario would see scRNA-seq tools accurately gauging and correcting uncertainties from these biases and errors (Bai et al. Genomics 112, 346-355 (2020); Chu et al. Brief Bioinform 23 (2022)). However, discerning technical errors from genuine biological variations using computational techniques alone proves difficult. In earlier studies, sequencing and PCR errors were identified as sources of technical inconsistencies in single-cell transcriptomics, causing inaccurate feature counting (Sun et al. bioRxiv, 2023.2004.2006.535911 (2023); Philpott et al. Nat Biotechnol 39, 1517-1520 (2021)). Yet, the impact of oligonucleotide bead synthesis on droplet-based single-cell sequencing has been largely unexplored.

[0252] There are two primary bead synthesis strategies: on-bead chemical synthesis and enzymatic ligation. The on-bead chemical method incorporates a barcode generated using a split and pool technique, and then a random UMI synthesised using single random nucleosides, a concept first introduced by Drop-seq (Macosko et al. Cell 161, 1202-1214 (2015)). Its efficient creation of capture oligonucleotides on beads offers a diverse barcode size and ensures most cells align with a bead. Conversely, the enzymatic ligation method, initially introduced by InDrops2and adopted by other techniques (Zheng et al. Nat Commun 8, 14049 (2017); Delley et al. Sci Rep 11, 10857 (2021)), presents a more modular approach to bead synthesis. In this strategy, bead fabrication utilises combinations of a limited set of pre-synthesised oligonucleotides. In alignment with the chemical synthesis method's principles, the barcode is created using a split and pool approach using small stretches of pre-synthesised building blocks. However, a distinct feature here is that a pre-synthesised random UMI is enzymatically ligated to the barcode. While these methods differ in their approach to bead synthesis, both exhibit a notable drawback. Both are prone to significant truncation, likely due to challenges in purifying the oligonucleotides once affixed to the beads. The immobilised oligonucleotides can't be cleaved and purified due to the cell barcodes contained within each one.

[0253] We have shown that synthesis issues significantly influence the precision of single-cell sequencing. Oligonucleotide synthesis errors, with an average rate of 1 in 100 bases (Carr et al. Nucleic Acids Res 32, el62 (2004); Masaki et al. Sci Rep 12, 12095 (2022)), commonly result in truncated or elongated oligonucleotides. Although in standard practices these can be separated from solid bases and purified using HPLC or PAGE, it's not an option for single-cell sequencing. As a result, these length discrepancies can interfere with barcode and UMI sequence detection, a problem further amplified in long- read sequencing. This is because, in long-read sequencing, barcodes are discerned through computational alignment to the PCR primer, coupled with positional matching to pinpoint the cell barcode and UMI's beginning. Variability in oligonucleotide lengths on beads makes barcode and UMI detection more error prone. Our findings suggest that both dropseq and 10X beads struggle with barcode and UMI detection, more so during long- read sequencing. UMIs present a larger issue than barcodes; while whitelisting can correct most barcode errors, the UMIs' inherent randomness makes whitelisting impractical. To tackle synthesis-related challenges with UMI detection, we incorporated a VRFS and constructed our UMIs using homodimer nucleoside amidites. This strategy not only pinpoints the start of the UMI but also identifies and fixes issues in the UMI zone arising from PCR, sequencing, and synthesis.

[0254] Example 7 - Methods

[0255] Cell lines and reagents

[0256] Jurkat cells were purchased from ATCC. 5TGM1 were a kind gift from Prof Clair Edwards. Both Jurkat and 5TGM1 cells were cultured in complete RPMI medium supplemented with 10% Foetal Calf Serum. All parental cell lines were tested twice per year for mycoplasma contamination and authenticated by STR during this project.

[0257] Oligonucleotide synthesis

[0258] Single-cell oligonucleotide bead synthesis was executed in line with previously established methods (Philpott et al. (2023) and Philpott et al. (2021)), but with the following alterations to the bead design:

[0259] 5 ’ -B ead-HEG_Linker-

[0260] TCTCTCTCTACACGACGCTCTTCCGATCTJJJJJJJJJJJJBAGCNNNNNNNNVTTTTTTTTTTTTTTTTTTTTTTTTTTTTTT 3,

[0261] Here, "J" represents the monomer split-and-pool barcode, while "N" signifies the dimer amidite UMI. We procured CMI oligos from Sigma-Aldrich with desalting, designed as follows: Spacer oligo:

[0262] 5’-

[0263] TCTCTCTCTACACGACGCTCTTCCGATCTAGTGCGTAGCTGBAGCGGAACCTT GGCCTTAATTGGTTAAGGTTGGA4TTTTTTTTTTTTTTTT-3’

[0264] No spacer:

[0265] 5’-

[0266] TCTCTCTCTACACGACGCTCTTCCGATCTAGTGCGTAGCTGGGAACCTTGGCC TTAATTGGTTAAGGTTGGA4TTTTTTTTTTTTTTTT-3’

[0267] VRFS is underlined. V based is underlined and italicised.

[0268] Bead sequencing library preparation

[0269] 2,000 beads (in TE / TW storage buffer) were added to a PCR tube. Beads were washed in 200 pl of H2O, centrifuged (100 g, 1 minute) and supernatant removed. 50 pl of PCR mix (1.5 pl indexed polyA primer (100 pM), 1.5 pl new P5 primer (100 pM; for beads carrying a SMART PCR handle) or 1.5 pl NEBNext i50x primer (100 pM; for beads carrying P5 PCR handle) or 10 pl SI primer (from 10X kit; for 10X beads), 25 pl KAPA HiFi master mix and H2O to 50 pl) was added, mixed and immediately run in a thermocycler with the following conditions. 95°C for 3 minutes. 12 cycles of: 95°C for 20 seconds, 60°C for 15 seconds, 72°C for 15 seconds then 72°C for 1 minute and 4°C hold. After PCR, samples were cleaned up by adding 100 pl of SPRIselect beads (Beckman Coulter) and following the manufacturer’s instructions. Samples were eluted in 20 pl of H2O and run on a HS DI 000 tape (Agilent).

[0270] ONT Flongle libraries were prepared and loaded according to protocol sqk-lskl l4- ACDE_9163_vl 14_revJ_29Jun2022-flongle, starting with 100 fmol of PCR product (or pooled products where indexing was used). 10 fmol of library was loaded onto a flongle flow cell. Sequencing was run for 24 hours using super-accuracy basecalling, minimum fragment size of 20 bp and no filtering based on Q score. Primers used for the PCR and sequencing:

[0271]

[0272] 1 OX chromium library preparation

[0273] We prepared a single-cell suspension using Jurket and 5TGM1 cells using the standard 10X Genomics chromium protocol as per the manufacturer’s instructions. Briefly, cells were filtered into a single-cell suspension using a 40 pM Flomi cell strainer before being counted. We performed 10X Chromium library preparation following the manufacturers protocol. Briefly, we loaded 3,300 Jurkat:5TGMl cells at a 70:30 split into a single channel of the 10X Chromium instrument. Cells were barcoded and reverse transcribed into cDNA using the Chromium Single Cell 3’ library kit and get bead v3.1. We performed 10 cycles of PCR amplification before cleaning up the library using 0.6X SPRI Select beads. The library was split and a further 20 or 25 PCR cycles were performed using a biotin oligonucleotide (5-PCBioCTACACGACGCTCTTCCGATCT) and then cDNA was enriched using DynabeadsTM MyOneTM streptavidin T1 magnetic beads (Invitrogen). The beads were washed in 2X binding buffer (lOmM Tris-HCL pH7.5, ImM EDTA and 2M NaCl) then samples were added to an equi-volume amount of 2X binding buffer and incubated at room temperature for 10 mins. Beads were placed in a magnetic rack and then washed with twice with IX binding buffer. The beads were resuspended in H2O and incubated at room temperature and subjected to long-wave UV light (-366 nm) for 10 minutes. Magnetic beads were removed, and library was quantified using the QubitTM High sensitivity kit. Libraries were then prepared before sequencing.

[0274] Dropseq and scCOLOR-seqv2 library preparation

[0275] Single-cell capture and reverse transcription were performed as previously described. Briefly, Jurkat and 5TGM1 cells (20:80 ratio) were filtered into a single-cell suspension using a 40 pM Flomi cell strainer before being counted. Cells were loaded into the DolomiteBio Nadia Innovate system at a concentration of 310 cells per pL. Custom synthesised beads were loaded into the microfluidic cartridge at a concentration of 620,000 beads per mL. Cell capture was then performed using the standard Nadia Innovate protocol according to manufacturer’s instructions. The droplet emulsion was then incubated for 10 mins before being disrupted with lH,lH,2H,2H-perfluoro-l -octanol (Sigma) and beads were released into aqueous solution. After several washes, the beads were subjected to reverse transcription. Prior to PCR amplification, beads were treated with Exol exonuclease for 45 min. PCR amplification was then performed using the SMART PCR primer (AAGCAGTGGTATCAACGCAGAGT) and cDNA was subsequently purified using AMPure beads (Beckman Coulter). The library was split and a further 20 or 25 PCR cycles were performed using a biotin oligonucleotide (5 —

[0276] PCBioTACACGACGCTCTTCCGATCT) and then cDNA was enriched using DynabeadsTM MyOneTM streptavidin T1 magnetic beads (Invitrogen). The beads were washed in 2X binding buffer (lOmM Tris-HCL pH7.5, ImM EDTA and 2M NaCl) then samples were added to an equi-volume amount of 2X binding buffer and incubated at room temperature for 10 mins. Beads were placed in a magnetic rack and then washed with twice with IX binding buffer. The beads were resuspended in H2O and incubated at room temperature and subjected to long-wave UV light (-366 nm) for 10 minutes. Magnetic beads were removed, and library was quantified using the QubitTM High sensitivity kit. Libraries were then prepared for sequencing.

[0277] Single-cell Nanopore library preparation for sequencing

[0278] A total of 500 ng of single-cell PCR input was used as a template for ONT library preparation. Library preparation was performed using the SQK-LSK114 (kit V14) ligation sequencing kit, following the manufacturers protocol. Samples were then sequenced on either a Flongle™ device or a PromethlON™ device using RIO.4 (FLO-PRO114M) flow cells. The fast5 sequencing data was basecalled to fastq files using guppy basecaller (v6.5.7) using the Super-accuracy mode within the MinKnow software (v23.04.6).

[0279] / OX chromium short-read analysis workflow

[0280] The data was processed using a custom CGAT-core14pipeline ‘pipeline lOx shortread’, which is included within the TallyTriN Github repository (https: / / github.com / cribbslab / TallyTriN / blob / main / tallytrin / pipeline_10x_shortread.py). Briefly, the quality of each fastq file is evaluated using fastqc toolkit and summary statistics collated using Multiqc. We then identify putative barcodes using UMI-tools whitelist module and then extract the barcodes and UMIs from the read 1 fastq file and append them onto the read 2 file using umi tools extract module. Hisat2 is then used to map the reads to the hg38_ensembl98 genome and the resulting bam file is then sorted, and each read is assigned to a feature using featureCounts, with the alignment written to the XT flag of the output bam file. This is then indexed using samtools and then UMI counting is performed using the umi tools count module before being converted to a market matrix format. The resulting matrix files are then parsed into R / Bioconductor (v4.0.3) using the BUSpaRse (vl.14.1) package and downstream analysis was performed using Seurat (v 4.3.0.1). Transcript matrices were cell-level scaled and log-transformed. The top 2000 highly variable genes were then selected based on variance stabilising transformation which was used for principal component analysis (PCA). Clustering was performed within Seurat using the Louvain algorithm. To visualise the single-cell data, we projected data onto a Uniform Manifold Approximation and Projection (UMAP). / OX chromium long-read analysis workflow

[0281] To analyse the 10X chromium long -read data, we developed a custom cgatcore pipeline named ‘pipeline lOx’ in the TallyTriN Github repository (https: / / github.com / cribbslab / TallyTriN / blob / main / tallytrin / pipeline_10x.py). We split the fastq file into segments to optimize processing time. Each read's polyA tail was identified and reverse complemented for consistent orientation, discarding reads without a polyA tail. We identified the barcode and UMI by locating the PCR primer sequence ‘AGATCGGAAGAGCGT’ through pairwise alignment, and then extracted the 16 bp barcode and 12 bp UMI based on position. Barcodes were corrected using a whitelisting method similar to UMI-tools.

[0282] Mapping was done using minimap2 (v2.22) with settings: -ax splice -uf MD -sam-hit-only -junc-bed, referencing the human hg38 and mouse mm 10 transcriptomes. The resultant sam file was arranged and indexed via samtools. Read counting employed UMI-tools' count module, converting counts to a matrix format. We processed the raw expression matrices using R / Bioconductor (v4.0.3) and devised scripts to depict barnyard plots, displaying mouse and human cell proportions. Matrices were cell-level scaled and centre log ratio transformed. We selected the top 2000 variably expressed genes post-variance stabilising transformation for PC A. Clusters were identified in Seurat using the Louvain algorithm. For visualisation, we projected the single-cell data onto a Uniform Manifold Approximation and Projection (UMAP).

[0283] Dropseq analysis workflow

[0284] Dropseq data was processed using a CGAT-core workflow ‘pipeline macosko’, which is included within the TallyTriN repository

[0285] (https: / / github.com / cribbslab / TallyTriN / blob / main / tallytrin / pipeline_singlecell_macosko. py). Briefly, the fastq file was split into chunks so that analysis scripts can be processed on sections of the data to reduce the processing time. The polyA tail for each read was identified and then reverse complemented to keep all the reads within the same orientation, any reads not containing a polyA tail were discarded. Next the barcode and UMI was identified based on the identification of the PCR primer sequence using pairwise alignment and then selecting the 12 bp barcode and the 8bp UMI using positional matching. Next, the barcodes were corrected using a whitelisting approach like the one implemented by UMI-tools. The reads were then merged and then mapping was performed using minimap2 (v2.22). Mapping settings we as follows: -ax splice -uf MD -sam-hit-only -junc-bed and using the reference transcriptome for human hg38 and mouse mm 10. The resulting sam file was sorted and indexed using samtools. Mapping settings were as follows: -ax splice - uf MD -sam-hit-only -junc-bed and using the reference transcriptome for human hg38 and mouse mm 10. The resulting sam file was sorted and indexed using samtools. Counting was performed using UMI-tools count module before being converted to a market matrix format. Raw transcript expression matrices generated were processed using R / Bioconductor (v4.0.3) and custom scripts were used to generate barnyard plots showing the proportion of mouse and human cells. Transcript matrices were cell-level scaled and centre log ratio transformed. The top 2000 highly variable genes were then selected based on variance stabilising transformation which was used for principal component analysis (PCA). Clustering was performed within Seurat using the Louvain algorithm. To visualise the single-cell data, we projected data onto a Uniform Manifold Approximation and Projection (UMAP). scCOLOR-seqv2 analysis workflow

[0286] To process the drop-seq data, we wrote a custom cgatcore pipeline. We followed the workflow previously described for identifying barcodes and UMIs using scCOLOR-seq sequencing analysis. Briefly, to determine the orientation of our reads, we first searched for the presence of a polyA sequence or a polyT sequence. In cases were the polyT was identified, we reverse complemented the read. We next identified the barcode sequence by searching for the polyA region and flanking regions before and after the barcode. The dimer UMI was identified based upon the primer sequence TCTTCCGATCT at the TSO distal end of the read. Barcodes and UMIs that had a length less than 50 base pairs were discarded. Next, the barcodes were corrected using a whitelisting approach like the one implemented by UMI-tools. The reads were then merged and then mapping was performed using minimap2 (v2.22). Mapping settings we as follows: -ax splice -uf MD -sam-hit-only -junc-bed and using the reference transcriptome for human hg38 and mouse mm 10. The resulting sam file was sorted and indexed using samtools. Mapping settings we as follows: -ax splice -uf MD -sam-hit-only -junc-bed and using the reference transcriptome for human hg38 and mouse mm 10. The resulting sam file was sorted and indexed using samtools. The transcript name was then appended to the bam XT flag using the xt tag nano script before umi tools count module was used to count the features using the following settings: — per-gene — gene-tag=XT -per-cell — dual -nucleotide. The umi tools used to correct for the dimer UMIs is located in the AC-dual oligo in a fork at the repository: https: / / github.com / Acribbs / UMI-tools. Downstream analysis in R was then performed as described for the dropseq analysis workflow above.

[0287] Barcode, UMI and cDNA region extraction — protocol 1

[0288] Algorithm:

[0289] 1. A "reference" read to use for extraction of barcodes, UMIs and cDNA regions is specified by the user. This contains the sequences of various constant elements known to be in each read (e.g. the PCR handles and VRFSs) as well as the sequences of regions variable in each read (e.g. the barcode and UMI), the latter of which can be represented with N bases.

[0290] 2. Each sequenced read is aligned to both the reference read and the reference read in reverse complement. The qualities of these alignments are compared, with the better alignment predicting the complement of the sequenced read.

[0291] 3. The best of these alignments is taken and the bases of the sequenced read that align to the bases of the barcode and UMI in the reference read, or any other sequenced named when specifying the reference read, are extracted.

[0292] 4. The cDNA region is extracted as the sequence of bases after the polyT but before the TSO-end PCR handle.

[0293] 5. The sequences extracted are quality checked, for instance, by ensuring reasonable extracted sequence length, high dimer count in the UMI and thymine rich polyT sequences. Failing a quality check places the sequenced read in a failed bin.

[0294] 6. The cDNA regions of reads that passed QC are written to an output file annotated with the extracted barcodes and UMIs. Reads in the failed bin can optionally be written to a different file for further analysis. Barcode, L MI and cDNA region extraction — protocol 2

[0295] Algorithm:

[0296] 1. A "reference" read to use for extraction of barcodes, UMIs and cDNA regions is specified by the user. This contains the sequences of various constant elements known to be in each read (e.g. the PCR handles and VRFSs) as well as the sequences of regions variable in each read (e.g. the barcode and UMI).

[0297] 2. Each sequenced read is aligned to both the reference read and the reference read in reverse complement, using the Smith-Waterman algorithm. Each entire sequence is aligned to one another; however, the sequences are stored programmatically in segments, one for each feature.

[0298] 1. Specify Reference Read

[0299] -ATCT JJJJJJ CAGC NNNNNN ACGA TTTTTT-

[0300] Where ATCT is the bead-end PCT handle, Js are the nucleotides of the barcode, CAGC is VRFS1, Ns are the nucleotides of the UMI, ACGA is VRFS2, TTTTTT isthe analyecapture sequence.

[0301] 2. Load Sequenced Read from Input Fastq File -ATCTACTGACATCAAGGCCACGATTTTT-

[0302] 3. Align Against Sequenced Read

[0303] Reference: -ATCT JJJJJJ CAGC NNNNNN ACGA TTTTTT-

[0304] Sequenced: -ATCT ACTCGA CAGC AAGGCC ACGA TTTTTT-

[0305] 4. Extract Desired Features From Alignment Barcode: ACTCGA

[0306] UMI: AAGGCC

[0307] 3. The qualities of these alignments are compared, with the better alignment predicting the complement of the sequenced read. For each feature, an individual quality is calculated as the percentage of matches between the reference and sequenced reads in the feature. This linear approach can alternatively be replaced by an exponential method that increases quality exponentially with the number of matches. Overall alignment quality is then calculated by taking a weighted-average of these individual feature qualities, with these weights specified by the user.

[0308] An example of these weights is: bead_pcr_handle: 0.3; VRFS1 : 0.25; umi: 0.2; VRFS2: 0.25. The weights must add to 1.

[0309] 4. The best of these alignments is taken and the bases of the sequenced read that align to the bases of the barcode and UMI in the reference read, or any other sequenced named when specifying the reference read, are extracted.

[0310] 5. The cDNA region is extracted as the sequence of bases after the polyT but before the TSO-end PCR handle.

[0311] 6. The sequences extracted are quality checked. Failing a quality check places the sequenced read in a failed bin. The quality checks are as follows:

[0312] (a) Barcode Length: barcodes must be within a user-specified length range to allow a read to pass.

[0313] (b) UMI Length: the UMI must be within a user-specified length range to allow the read to pass.

[0314] (c) UMI Dimer Count: the UMI must contain more than a user-specified number of dimers for the read to pass.

[0315] (d) PolyT Length: the polyT region must be longer than a user-specified length for the read to pass.

[0316] (e) Alignment Quality: overall read alignment quality must be greater than a user- specified threshold for the read to pass.

[0317] (f) Corrupt PolyT : The polyT region must contain a percentage of T bases greater than a user-specified value for a read to pass.

[0318] (g) Ambiguous PolyT -When a V base is present before the polyT region in the reference read, if a valid V base is not found in the read tested, and the end of the UMI contains more T bases than a user-specified value then a read will be failed. 7. The cDNA regions of reads that passed quality checks are written to an output file annotated with the extracted barcodes and UMIs. Reads in the failed bin can optionally be written to a different file for further separate downstream analysis.

[0319] Further Embodiments

[0320] 1. An array of polynucleotides, wherein each polynucleotide of the array comprises:

[0321] (a) a PCR handle sequence;

[0322] (b) a barcode sequence (BC);

[0323] (c) a unique molecular identifier sequence (UMI); and

[0324] (d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI.

[0325] 2. The array of embodiment 1, wherein each polynucleotide of the array further comprises:

[0326] (e) an analyte capture sequence.

[0327] 3. The array of embodiment 1 or 2, wherein each polynucleotide of the array further comprises a second VRFS.

[0328] 4. The array of embodiment 3, wherein the second VRFS is (A) located 3’ of the portion of the polynucleotide comprising the BC and the UMI; and / or (B) between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI.

[0329] 5. The array of any one of the preceding embodiments, wherein the first VRFS and / or the second VRFS is 2 to 20 nucleotides in length.

[0330] 6. The array of any one of the preceding embodiments, wherein the first VRFS and / or the second VRFS is 2 to 8 nucleotides in length. 7. The array of any one of the preceding embodiments, wherein the first VRFS and / or the second VRFS is 4 to 8 nucleotides in length.

[0331] 8. The array of any one of the preceding embodiments, wherein the first VRFS and / or the second VRFS comprises 1 or more, 2 or more, 3 or more, or 4 or more constant nucleotides.

[0332] 9. The array of any one of the preceding embodiments, wherein the first VRFS and / or the second VRFS consists of constant nucleotides.

[0333] 10. The array of any one of embodiments 1 to 8, wherein the first VRFS and / or the second VRFS comprises 1 or more variable nucleotides.

[0334] 11. The array of embodiment 10, wherein the variable nucleotides are limited variable nucleotides.

[0335] 12. The array of any one of the preceding embodiments, wherein the first and / or the second VRFS does not comprise the same nucleotide in two consecutive positions.

[0336] 13. The array of any one of embodiments 3 to 12, wherein the first VRFS comprises a different sequence to the second VRFS.

[0337] 14. The array of any one of embodiments 3 to 13, wherein the first VRFS has a Levenshtein edit distance of at least two when compared to the second VRFS.

[0338] 15. The array of any one of embodiments 3 to 14, wherein the nucleotide at each position of the first VRFS is different to the nucleotide at each corresponding position of the second VRFS.

[0339] 16. The array of any one of embodiments 3 to 15, wherein the first VRFS is not the reverse complement of the second VRFS. 17. The array of any one of the preceding embodiments, wherein each polynucleotide of the array does not form a hairpin sequence and / or does not dimerise.

[0340] 18. The array of any one of the preceding embodiments, wherein the array comprises fewer than 10, fewer than 8, or fewer than 5 different first VRFSs.

[0341] 19. The array of any one of the preceding embodiments, wherein essentially each polynucleotide of the array comprises the same first VRFS.

[0342] 20. The array of any one of embodiments 3 to 19, wherein the array comprises fewer than 10, fewer than 8, or fewer than 5 different second VRFSs.

[0343] 21. The array of any one of embodiments 3 to 20, wherein essentially each polynucleotide of the array comprises the same second VRFS.

[0344] 22. The array of any one of the preceding embodiments, wherein the first VRFS is immediately 3’ of the BC, and the 5’ nucleotide of the first VRFS is different to the 3’ nucleotide of the BC.

[0345] 23. The array of any one of embodiments 1 to 21, wherein the first VRFS is immediately 3’ of the UMI and the 5’ nucleotide of the first VRFS is different to the 3’ nucleotide of the UMI.

[0346] 24. The array of any one of embodiments 3 to 22, wherein the second VRFS is immediately 3’ of the UMI, and the 5’ nucleotide of the second VRFS is different to the 3’ nucleotide of the UMI.

[0347] 25. The array of any one of embodiments 3 to 21 and 23, wherein the second VRFS is immediately 3’ of the BC, and the 5’ nucleotide of the second VRFS is different to the 3’ nucleotide of the BC. 26. The array of any one of embodiments 3 to 25, wherein the second VRFS is immediately 5’ of the analyte capture sequence, and the 3’ nucleotide of the second VRFS is different to the 5’ nucleotide of the analyte capture sequence.

[0348] 27. The array of one of embodiments 3 to 26, wherein the 3’ nucleotide of the second VRFS is not a T.

[0349] 28. The array of any one of the preceding embodiments, wherein the 5’ nucleotide of the BC and / or the UMI is different to the 3’ nucleotide of its 5’ flanking sequence, wherein its 5’ flanking sequence is the PCR handle or the first VRFS.

[0350] 29. The array of any one of the preceding embodiments, wherein the 3’ nucleotide of the BC and / or the UMI is different to the 5’ nucleotide of its 3’ flanking sequence, wherein its 3’ flanking sequence is the first VRFS, the second VRFS, or the analyte capture region.

[0351] 30. The array of any one of the preceding embodiments, wherein the nucleotide at the 5’ end and / or the 3’ end of the BC and / or the UMI is a limited variable base.

[0352] 31. The array of any one of the preceding embodiments, wherein the first nucleotide (5’ nucleotide) of the first VRFS is different to the second nucleotide of the first VRFS and / or the last nucleotide (3’ nucleotide) of the first VRFS is different from the penultimate nucleotide of the first VRFS.

[0353] 32. The array of any one of embodiments 3 to 31, wherein the first nucleotide (5’ nucleotide) of the second VRFS is different to the second nucleotide of the second VRFS and / or the last nucleotide (3’ nucleotide) of the second VRFS is different from the penultimate nucleotide of the second VRFS.

[0354] 33. The array of any one of embodiments 3 to 32, wherein each polynucleotide of the array comprises, in a 5’ to 3’ direction:

[0355] (a) the PCR handle sequence; (b) the BC or the UMI;

[0356] (c) the first VRFS;

[0357] (d) the BC or the UMI; and

[0358] (e) the second VRFS.

[0359] 34. The array of any one of embodiments 3 to 33, wherein each polynucleotide of the array comprises, in a 5’ to 3’ direction:

[0360] (a) the PCR handle sequence;

[0361] (b) the BC;

[0362] (c) the first VRFS;

[0363] (d) the UMI; and

[0364] (e) the second VRFS.

[0365] 35. The array of embodiment 33 or 34, wherein each polynucleotide of the array further comprises the analyte capture sequence 3’ of the second VRFS.

[0366] 36. The array of any one of the preceding embodiments, wherein each polynucleotide of the array is DNA.

[0367] 37. The array of any one of the preceding embodiments, wherein the analyte capture sequence is a polyT sequence.

[0368] 38. The array of any one of the preceding embodiments, wherein the BC and / or the UMI comprises a series of discrete nucleotide blocks, wherein the discrete nucleotide blocks are from a mixed pool of nucleotide blocks of known sequence, and wherein each nucleotide block sequence in the pool differs from each other nucleotide block in the pool by at least two nucleotide substitutions.

[0369] 39. The array of embodiment 38, wherein the discrete nucleotide blocks are dimers, preferably homodimers. 40. The array of any one of the preceding embodiments, wherein the BC comprises or consists of 4-20 nucleotides or nucleotide blocks, and / or wherein the UMI comprises or consists of 4-16 nucleotides or nucleotide blocks.

[0370] 41. The array of any one of the preceding embodiments, wherein the BC comprises or consists of 8-16 nucleotides or nucleotide blocks.

[0371] 42. The array of any one of the preceding embodiments, wherein the UMI comprises or consists of 6-10 nucleotides or nucleotide blocks.

[0372] 43. The array of any one of the preceding embodiments, wherein each polynucleotide of the array does not comprise a hairpin sequence.

[0373] 44. The array of any one of the preceding embodiments, wherein the array comprises at least 105polynucleotides.

[0374] 45. The array of any one of the preceding embodiments, wherein each polynucleotide of the array, is hybridised to, or further comprises, a sequence corresponding to an analyte.

[0375] 46. The array of any one of the preceding embodiments, wherein the array is divided into a plurality of sub-arrays, wherein the BC of each polynucleotide is the same as the BC of essentially each other polynucleotide of the same sub-array, but different from the BC of the polynucleotides of essentially every different sub-array.

[0376] 47. A micro-particle comprising a micro-bead and an array of polynucleotides according to any one of embodiments 1 to 45, wherein each polynucleotide is bound to the microbead.

[0377] 48. The micro-particle of embodiment 47, wherein the array comprises fewer than 10, fewer than 8, or fewer than 5 different BCs, optionally wherein the BC sequence of essentially each polynucleotide of the array is the same. 49. A plurality of micro-particles according to embodiment 47 or 48, wherein the BC sequence of each polynucleotide of each micro-particle is the same as the BC sequence of essentially each other polynucleotide of the same micro-particle, and different from the BC sequence of the polynucleotides of each other micro-particle.

[0378] 50. A method for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises:

[0379] (a) a PCR handle sequence;

[0380] (b) a barcode sequence (BC);

[0381] (c) a unique molecular identifier sequence (UMI); and

[0382] (d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI, and wherein the method comprises:

[0383] (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and

[0384] (ii) predicting the positions of the BC and / or the UMI of the read and / or its reverse complement based on the alignment.

[0385] 51. The method of embodiment 50, wherein the polynucleotide or its reverse complement further comprises an analyte capture sequence.

[0386] 52. The method of embodiment 50 or 51, wherein the polynucleotide or its reverse complement further comprises a second VRFS.

[0387] 53. The method of embodiment 52, wherein the second VRFS is (A) located 3’ of the portion of the polynucleotide comprising the BC and the UMI; and / or (B) between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI. 54. A method for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises:

[0388] (a) a PCR handle sequence;

[0389] (b) a barcode sequence (BC);

[0390] (c) a unique molecular identifier sequence (UMI);

[0391] (d) a second variable region flanking sequence (VRFS), wherein the second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI, and wherein the method comprises:

[0392] (i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and

[0393] (ii) predicting the positions of the BC and / or the UMI based on the alignment.

[0394] 55. The method of embodiment 54, wherein the polynucleotide or its reverse complement further comprises an analyte capture sequence.

[0395] 56. The method of embodiment 55, wherein the second VRFS is between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI.

[0396] 57. The method of any one of embodiments 50 to 56, wherein the polynucleotide is a polynucleotide as defined in any one of embodiments 1 to 46.

[0397] 58. The method of any one of embodiments 50 to 57, wherein the method is performed on reads produced by sequencing a polynucleotide or an array of polynucleotides, wherein the polynucleotide or the array is as defined in any one of embodiments 1 to 46. 59. The method of embodiment 58, wherein the sequencing is performed using long-read sequencing technology.

[0398] 60. The method of any one of embodiments 50 to 59, wherein the polynucleotide or each polynucleotide of the array of polynucleotides further comprises a captured analyte or its reverse complement and is generated by a method comprising: capturing analyte molecules on a polynucleotide, or an array of polynucleotides, as defined in any one of embodiments 1 to 44 and 46; and generating copies of the polynucleotide or the array of polynucleotides, including the captured analyte.

[0399] 61. The method of embodiment 60, wherein the analyte is RNA, preferably mRNA.

[0400] 62. The method of any one of embodiments 50 to 61, wherein the polynucleotide or each polynucleotide of the array of polynucleotides further comprises a captured analyte or its reverse complement and is generated by a method comprising: capturing RNA molecules on a polynucleotide, or an array of polynucleotides, as defined in any one of embodiments 1 to 44 and 46; performing reverse transcription of captured RNA molecules using a reverse transcriptase and primed using the captured RNA molecules to generate cDNA polynucleotides having 3’ end non-templated nucleotides; annealing a set of template switch oligonucleotides (TSOs) to the 3’ end non- templated nucleotides; and amplifying the cDNA.

[0401] 63. The method of any one of embodiments 50 to 62, wherein the reference sequence comprises two or more of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence.

[0402] 64. The method of embodiment 63, wherein the positions of the BC and / or the UMI are predicted based at least in part on the positions of two or more of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence in the reference sequence.

[0403] 65. The method of embodiment 63 or 64, wherein the reference sequence comprises the PCR handle, the first VRFS, the second VRFS, and the analyte capture sequence.

[0404] 66. The method of embodiment 65, wherein the positions of the BC and / or the UMI are predicted based at least in part on the positions of the PCR handle, the first VRFS, the second VRFS, and the analyte capture sequence in the reference sequence.

[0405] 67. The method of any one of embodiments 50 to 62, wherein the reference sequence comprises the BC, the UMI, and the VRFS(s).

[0406] 68. The method of embodiment 67, wherein the reference sequence further comprises the second VRFS, the PCR handle and / or the analyte capture sequence.

[0407] 69. The method of any one of embodiments 50 to 68, further comprising delimiting the positions of the PCR handle sequence, the BC, the UMI, the first VRFS, the second VRFS and / or the analyte capture sequence in the read and / or its reverse complement.

[0408] 70. The method of any one of embodiments 50 to 69, wherein the aligning step further comprises comparing the quality of the alignment of two or more of (A) the read to the reference read, (B) the read to the reverse complement of the reference read, (C) the reverse complement of the read to the reference read, and (D) the reverse complement of the read to the reverse complement of the reference read; and wherein the method further comprises selecting the alignment of higher quality.

[0409] 71. The method of embodiment 70, wherein the quality is determined by assigning a score for the alignment of the PCR handle sequence, a score for the alignment of the BC, a score for the alignment of the UMI, a score for the alignment of the first VRFS, a score for the alignment of the second VRFS and / or a score for the alignment of the analyte capture sequence. 72. The method of embodiment 71, wherein the score for the alignment is determined by:

[0410] (A) the percentage of matches in each aligned section; or

[0411] (B) an exponential of the percentage of matches in each aligned section.

[0412] 73. The method of embodiment 71 or 72, wherein the quality is determined by assigning a score for two or more of the alignment of the PCR handle sequence, the alignment of the BC, the alignment of the UMI, the alignment of the first VRFS, the alignment of the second VRFS, and the alignment of the analyte capture sequence, and taking a weighted average of the two or more scores.

[0413] 74. The method of any one of embodiments 50 to 73, wherein the method further comprises a step of determining a sequence for the BC and / or the UMI.

[0414] 75. The method of any one of embodiments 58 to 74, wherein the method further comprises a step of determining a sequence for the analyte.

[0415] 76. The method of embodiment 75, wherein the sequence of the analyte is determined as the sequence 3’ of the portion of the polynucleotide comprising the BC and the UMI, or is determined as the sequence 3’ of the analyte capture sequence; optionally wherein the sequence is further determined as the sequence 5’ of the sequence of a template-switch oligonucleotide (TSO) sequence.

[0416] 77. The method of embodiment 75 or 76, wherein the method further comprises annotating the sequence of the analyte with the BC and / or the UMI.

[0417] 78. The method of any one of embodiments 50 to 77, wherein the method further comprises comparing the predicted length of the BC and / or the UMI in the read to the length of the BC and / or the UMI in the reference sequence, and determining that there is an error in the read if the predicted length of the BC and / or the UMI differs from the length of the BC and / or the UMI in the reference sequence. 79. The method of any one of embodiments 50 to 78, wherein the BC and / or the UMI comprises nucleotide blocks and the method further comprises comparing the predicted number of nucleotide blocks in the BC and / or the UMI in the read to the number of nucleotide blocks of the BC and / or the UMI in the reference sequence, and determining that there is an error in the read if the number of nucleotide blocks in the predicted BC and / or the UMI differs from the number of nucleotide blocks in the BC and / or the UMI in the reference sequence.

[0418] 80. The method of any one of embodiments 50 to 79, wherein the method further comprises comparing the predicted length of the BC and / or the UMI in the read to the length of the BC and / or the UMI in the reference sequence and binning the read if the predicted length of the BC and / or the UMI differs by at least one nucleotide from the length of the BC and / or the UMI in the reference sequence.

[0419] 81. The method of any one of embodiments 50 to 80, wherein the BC and / or the UMI comprises nucleotide blocks and the method further comprises comparing the predicted number of nucleotide blocks in the BC and / or the UMI in the read to the number of nucleotide blocks in the BC and / or the UMI in the reference sequence and binning the read if the number of nucleotide blocks in the predicted BC and / or the UMI differs by at least one nucleotide block from the number of nucleotide blocks in the BC and / or the UMI in the reference sequence.

[0420] 82. The method of any one of embodiments 50 to 81, wherein the BC and / or the UMI comprise a series of discrete nucleotide blocks, wherein the discrete nucleotide blocks are from a mixed pool of nucleotide blocks of known sequence, and wherein each nucleotide block sequence in the pool differs from each other nucleotide block in the pool by at least two nucleotide substitutions, and the method further comprising comparing the sequence of each nucleotide block to the sequence of each nucleotide block in the pool, and determining that there is an error in the read if the sequence of a nucleotide block is not present in the pool. 83. The method of any one of embodiments 50 to 82, wherein the method further comprises comparing the sequence of at least two of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence in the read or its reverse complement to the reference sequence, and discarding the read if greater than a specified number of the nucleotides in the sequences that were compared are different to the corresponding portions of the reference sequence.

[0421] 84. The method of any one of embodiments 50 to 83, wherein the method further comprises comparing the sequence of the PCR handle, the first VRFS, the second VRFS, and the analyte capture sequence of the read or its reverse complement to the reference sequence, and discarding the read if greater than a specified number of the nucleotides in the sequences that were compared are different to the corresponding portions of the reference sequence.

[0422] 85. The method of any one of embodiments 50 to 84, wherein the method further comprises comparing the sequence of the reference sequence to the sequence of the read or its reverse complement, and discarding the sequence if greater than a specified number of the nucleotides in the sequence that was compared is different to the corresponding portion(s) of the reference sequence.

[0423] 86. The method of any one of embodiments 50 to 85 wherein the polynucleotide comprises a limited variable base between the analyte capture region and the portion of the polynucleotide comprising the UMI, BC, PCR handle and first VRFS, wherein the method comprises discarding the read if the limited variable base is not identified in the sequence of the read or its reverse complement.

[0424] 87. The method of any one of embodiments 50 to 86, further comprising comparing the barcode sequence to a whitelist of known barcode sequences.

[0425] 88. The method of any one of embodiments 50 to 87, further comprising generating a whitelist of barcode sequences using a knee-plot. 89. The method of any one of embodiments 50 to 88 further comprising generating a whitelist of barcode sequences in real time, comprising: associating barcodes of similar sequence, ranking barcodes based on the number of reads of each barcode sequence, and assigning the highest-ranking barcode of each group of associated barcodes to the whitelist.

[0426] 90. The method of any one of embodiments 87 to 89, wherein if the barcode sequence is not in the whitelist, it is corrected by assigning it a new sequence corresponding to the barcode sequence in the whitelist which is the nearest match.

[0427] 91. The method of embodiment 90, wherein the barcode sequence is assigned the new sequence if the barcode sequence in the whitelist matches with a Levenshtein edit distance of 4 or less.

[0428] 92. The method of embodiment 90 or 91, wherein the barcode sequence is assigned the new sequence if the barcode sequence in the whitelist matches with a Hamming distance of 6 or less.

[0429] 93. The method of any one of embodiments 50 to 92, further comprising identifying two or more reads with the same BC and preferably the same analyte sequence, and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with a Levenshtein edit distance of 4 or less and / or a Hamming distance of 6 or less.

[0430] 94. The method of embodiment 93, wherein the method comprises assigning the two or more reads as having the same UMI if the UMI sequences of each read match with a Hamming distance of 4 or less.

[0431] 95. A method for delimiting barcode sequences and / or unique molecular identifier sequences in a library of reads, each read corresponding to the sequence of at least a portion of a polynucleotide, wherein each polynucleotide or its reverse complement comprises: (a) a PCR handle sequence;

[0432] (b) a barcode sequence (BC);

[0433] (c) a unique molecular identifier sequence (UMI);

[0434] (d) a first variable region flanking sequence (VRFS) between the BC and the UMI and / or a second VRFS located 3’ of the portion of the polynucleotide comprising the BC and the UMI and / or between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI; and

[0435] (e) optionally, an analyte capture sequence, and wherein the method comprises, for each read, performing the method of any one of embodiments 50 to 94.

[0436] 96. The method of any one of embodiments 50 to 95, wherein the method is a computer- implemented method.

[0437] 97. A computer readable medium comprising instructions for performing the method of any one of embodiments 50 to 96.

[0438] 98. A computer comprising the computer readable medium of embodiment 97.

[0439] 99. A method of making the array of polynucleotides of any one of embodiments 1 to 46, wherein the method comprises using split-and-pool polynucleotide synthesis to make the BC and / or using degenerate polynucleotide synthesis to make the UMI.

[0440] 100. A method of producing a library of polynucleotides, wherein the polynucleotides are amplified from the polynucleotides of a sample and / or tag non-polynucleotide analytes of a sample, the method comprising:

[0441] (a) capturing analytes in the sample on an array of polynucleotides according to any one of embodiments 1 to 44 and 46, a micro-particle according to embodiment 47 or 48, or a plurality of micro-particles according to embodiment 49;

[0442] (b) generating copies of the array of polynucleotides, including any (I) polynucleotides of the sample or (II) polynucleotides that tag non-polynucleotide analytes of the sample, captured by the array polynucleotides; and (c) amplifying the number of copies of each polynucleotide to produce a library of polynucleotides amplified from or tagging analytes in the sample.

[0443] 101. A library of polynucleotides obtained or obtainable by the method of embodiment 100.

[0444] 102. A method of generating a library of reads, comprising sequencing the library of polynucleotide produced by method of embodiment 100, or the library of polynucleotides of embodiment 101.

Claims

Claims1. An array of polynucleotides, wherein each polynucleotide of the array comprises:(a) a PCR handle sequence;(b) a barcode sequence (BC);(c) a unique molecular identifier sequence (UMI); and(d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI.

2. The array of claim 1, wherein each polynucleotide of the array further comprises an analyte capture sequence and / or a second VRFS.

3. The array of claim 2, wherein the second VRFS is (A) located 3’ of the portion of the polynucleotide comprising the BC and the UMI; and / or (B) between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI.

4. The array of any one of the preceding claims, wherein:(a) the first VRFS and / or the second VRFS is 2 to 20 nucleotides in length;(b) the first VRFS and / or the second VRFS consists of constant nucleotides; and / or(c) the first VRFS comprises a different sequence to the second VRFS.

5. The array of any one of the preceding claims, wherein:(a) the 5’ nucleotide of the BC and / or the UMI is different to the 3’ nucleotide of its 5’ flanking sequence, wherein its 5’ flanking sequence is the PCR handle or the first VRFS; and / or(b) the 3’ nucleotide of the BC and / or the UMI is different to the 5’ nucleotide of its 3’ flanking sequence, wherein its 3’ flanking sequence is the first VRFS, the second VRFS, or the analyte capture region.

6. The array of any one of claims 2 to 5, wherein each polynucleotide of the array comprises, in a 5’ to 3’ direction:(a) the PCR handle sequence;(b) the BC;(c) the first VRFS;(d) the UMI;(e) the second VRFS; and(f) optionally, the analyte capture sequence.

7. The array of any one of the preceding claims, wherein:(a) the analyte capture sequence is a polyT sequence;(b) the BC and / or the UMI comprises a series of discrete nucleotide blocks, wherein the discrete nucleotide blocks are from a mixed pool of nucleotide blocks of known sequence, and wherein each nucleotide block sequence in the pool differs from each other nucleotide block in the pool by at least two nucleotide substitutions, optionally wherein the discrete nucleotide blocks are dimers, preferably homodimers; and / or(c) each polynucleotide of the array, is hybridised to, or further comprises a sequence corresponding to, an analyte.

8. The array of any one of the preceding claims, wherein the array is divided into a plurality of sub-arrays, wherein the BC of each polynucleotide is the same as the BC of essentially each other polynucleotide of the same sub-array, but different from the BC of the polynucleotides of essentially every different sub-array.

9. A micro-particle comprising a micro-bead and an array of polynucleotides according to any one of claims 1 to 8, wherein each polynucleotide is bound to the micro-bead.

10. A plurality of micro-particles according to claim 9, wherein the BC sequence of each polynucleotide of each micro-particle is the same as the BC sequence of essentially each other polynucleotide of the same micro-bead, and different from the BC sequence of the polynucleotides of each other micro-particle.

11. A method for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises:(a) a PCR handle sequence;(b) a barcode sequence (BC);(c) a unique molecular identifier sequence (UMI); and(d) a first variable region flanking sequence (VRFS), wherein the first VRFS is between the BC and the UMI, and wherein the method comprises:(i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and(ii) predicting the positions of the BC and / or the UMI of the read and / or its reverse complement based on the alignment.

12. A method for delimiting a barcode sequence (BC) and / or a unique molecular identifier sequence (UMI) in a read corresponding to the sequence of at least a portion of a polynucleotide, wherein the polynucleotide or its reverse complement comprises:(a) a PCR handle sequence;(b) a barcode sequence (BC);(c) a unique molecular identifier sequence (UMI);(d) a second variable region flanking sequence (VRFS), wherein the second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI, and wherein the method comprises:(i) aligning one or more of (A) the read to a reference sequence, (B) the read to the reverse complement of the reference sequence, (C) the reverse complement of the read to the reference sequence, and (D) the reverse complement of the read to the reverse complement of the reference sequence; wherein the reference sequence comprises the sequence of the first VRFS; and(ii) predicting the positions of the BC and / or the UMI based on the alignment; optionally wherein the polynucleotide or its reverse complement further comprises an analyte capture sequence.

13. The method of claim 11 or 12, wherein the polynucleotide is a polynucleotide as defined in any one of claims 1 to 8.

14. The method of any one of claims 11 to 13, wherein the polynucleotide or each polynucleotide of the array of polynucleotides further comprises a captured analyte or its reverse complement and is generated by a method comprising: capturing analyte molecules on a polynucleotide, or an array of polynucleotides, as defined in any one of claims 1 to 8; and generating copies of the polynucleotide or the array of polynucleotides, including the captured analyte; optionally wherein the analyte is RNA, preferably mRNA.

15. The method of any one of claims 11 to 14, wherein the reference sequence comprises two or more of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence, and wherein the positions of the BC and / or the UMI are predicted based at least in part on the positions of two or more of (A) the PCR handle, (B) the first VRFS, (C) the second VRFS, and (D) the analyte capture sequence in the reference sequence.

16. The method of any one of claims 11 to 15, wherein the reference sequence comprises the BC, the UMI, and the VRFS(s), and optionally further comprises the second VRFS, the PCR handle and / or the analyte capture sequence.

17. The method of any one of claims 11 to 16, wherein the method further comprises comparing the predicted length of the BC and / or the UMI in the read to the length of the BC and / or the UMI in the reference sequence, and:(a) determining that there is an error in the read if the predicted length of the BC and / or the UMI differs from the length of the BC and / or the UMI in the reference sequence; and / or(b) binning the read if the predicted length of the BC and / or the UMI differs by at least one nucleotide from the length of the BC and / or the UMI in the reference sequence.

18. The method of any one of claims 11 to 17, further comprising:(a) comparing the barcode sequence to a whitelist of known barcode sequences;(b) generating a whitelist of barcode sequences using a knee-plot; and / or(c) generating a whitelist of barcode sequences in real time, comprising: associating barcodes of similar sequence, ranking barcodes based on the number of reads of each barcode sequence, and assigning the highest-ranking barcode of each group of associated barcodes to the whitelist; wherein if the barcode sequence is not in the whitelist, it is corrected by assigning it a new sequence corresponding to the barcode sequence in the whitelist which is the nearest match, optionally wherein the barcode sequence is assigned the new sequence if the barcode sequence in the whitelist matches with a Levenshtein edit distance of 4 or less or a Hamming distance of 6 or less.

19. The method of any one of claims 11 to 18, further comprising identifying two or more reads with the same BC and preferably the same analyte sequence, and assigning the two or more reads as having the same UMI if the UMI sequences of each read match with a Levenshtein edit distance of 4 or less and / or a Hamming distance of 6 or less.

20. A method for delimiting barcode sequences and / or unique molecular identifier sequences in a library of reads, each read corresponding to the sequence of at least a portion of a polynucleotide, wherein each polynucleotide or its reverse complement comprises:(a) a PCR handle sequence;(b) a barcode sequence (BC);(c) a unique molecular identifier sequence (UMI);(d) a first variable region flanking sequence (VRFS) and / or a second VRFS, wherein the first VRFS is between the BC and the UMI, wherein the second VRFS is located 3’ of the portion of the polynucleotide comprising the BC and the UMI and / or between the analyte capture sequence and the portion of the polynucleotide comprising the BC and the UMI; and(e) optionally, an analyte capture sequence, and wherein the method comprises, for each read, performing the method of any one of claims 11 to 19.

21. The method of any one of claims 11 to 20, wherein the method is a computer- implemented method.

22. A computer readable medium comprising instructions for performing the method of any one of claims 11 to 21, or a computer comprising the computer readable medium.

23. A method of producing a library of polynucleotides, wherein the polynucleotides are amplified from the polynucleotides of a sample and / or tag non-polynucleotide analytes of a sample, the method comprising(a) capturing analytes in the sample on an array of polynucleotides according to any one of claims 1 to 8, a micro-particle according to claim 9, or a plurality of micro-particles according to claim 10;(b) generating copies of the array of polynucleotides, including any sample polynucleotides captured by the array polynucleotides; and(c) amplifying the number of copies of each polynucleotide to produce a library of polynucleotides amplified from or tagging analytes in the sample.

24. A library of polynucleotides obtained or obtainable by the method of claim 23.

25. A method of generating a library of reads, comprising sequencing the library of polynucleotide produced by method of claim 23, or the library of polynucleotides of claim 24.

Citation Information

Patent Citations

  • Polynucleotide arrays

    WO2021229230A1

  • oligonucleotides

    WO2022118027A1

  • Chimeric artefact detection method

    WO2023194714A1

  • Bead-hashing

    WO2024028589A1

  • Methods and systems for processing polynucleotides

    US20200002764A1