Method of transposon insertion site sequencing capable of uniquely identifying insertions sites within repeated genetic elements

Long-read sequencing technologies overcome the limitations of short-read TIS by uniquely identifying insertion sites within repeated genetic elements, enhancing accessibility and reducing costs for scientists.

GB2617524BActive Publication Date: 2026-05-21QUADRAM INSITUTE BIOSCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
QUADRAM INSITUTE BIOSCI
Filing Date
2022-02-04
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Current transposon insertion site sequencing (TIS) methods are limited by short read lengths, which prevent unambiguous assignment of insertion sites within repeated genetic elements, and are costly and require specialized laboratory equipment.

Method used

Adopting long-read sequencing technologies, such as Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) single-molecule real-time (SMRT) sequencing, to generate reads exceeding 300 bp, particularly over 600 bp, enabling unique identification of insertion sites within repeated genetic elements.

Benefits of technology

Long-read sequencing allows for precise assignment of transposon insertion sites within repeated genetic elements, reducing costs and increasing accessibility to scientists with limited resources by using portable, cost-effective equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000003_0000
    Figure 00000003_0000
Patent Text Reader

Abstract

A method of transposon insertion sequencing (TIS) capable of resolving insertion sites in or around repeat elements, said method including the steps of; preparing a pool or library of mutant cells by
Need to check novelty before this filing date? Find Prior Art

Description

The present invention relates to methods of transposon (Tn) insertion site sequencing (TIS) which enables the unique identification of Tn insertions sites within repeating genetic elements of target genomes, reduces costs and improves accessibility of the technology to scientists with limited resources by broadening the range of sequencing platforms capable of producing suitable data. Although the following description refers largely to transposon-directed insertion site sequencing (TraDIS), the person skilled in the art will appreciate that other TIS methods could be employed, such as transposon sequencing (Tn-seq), high-throughput insertion tracking by deep sequencing (HITS), insertion sequencing (INSeq) and the like. Tn- mutagenesis with insertion site sequencing is a robust tool used to identify the genetic loci associated with the phenotype of an organism following any selective procedure. The genome-wide TIS methods were developed in 2009 and included a number of approaches to capture the Tn-insertion site sequence, (e.g. transposon-directed insertion site sequencing (TraDIS), transposon sequencing (Tn-seq), high-throughput insertion tracking by deep sequencing (HITS) and insertion sequencing (INSeq)). The methods exploit a range of techniques and demonstrate advantages and disadvantages. However, the most important limitation for all approaches is the length of the nucleotide sequence reads used to resolve the Tn-insertion site. The longest sequence reads that can be generated currently by these methods is up to 300 bases using single end sequencing. Due to this limitation, insertion sites within genes or other loci which represent repeated elements in the target genome cannot be assayed effectively. This is because insertion site sequences generated from within such repeat regions may all be identical, making unambiguous assignment of the real site (between repeated possible candidates) impossible, so the insertion sites within repeated genetic elements cannot be assigned uniquely. The Illumina(RTM) platform is currently the most widely used DNA sequencing technology. One of the advantages of using Illumina (RTM)-based approaches is that all the TIS protocols are optimised for this platform. Over the last few years, nucleotide sequencing costs on Illumina (RTM) machines have come down in price, and substantial data is available to compare the results from experiments carried out from a range of laboratories. However, there are some significant drawbacks inherent with current TIS protocols: 1) DNA sequencing instruments (often installed in a climate-controlled laboratory) can be expensive to maintain and operate. 2) While current TIS approaches have the capacity to multiplex many samples, this capacity may be a limitation if only one or two samples are required to be analysed as a whole sequencing flow cell still needs to be used. 3) Another challenge for some of the TIS approaches is that all the sequencing reads start with the same set of nucleotides (derived from the transposon). As a consequence, this identity is challenging for some of the sequencing platforms to resolve. To overcome this limitation for the Illumina (RTM) platform, nucleotide sequence diversity may be introduced in the sample to be sequenced (by including PhiX174 DNA or running dark cycles). However, this solution is inefficient as the additional sequence reads are present only to allow the unambiguous capture of the transposon insertion site junctions, after which they are discarded. 4) TIS methods may be limited by the sequence read length. For example, the longest read length on the Illumina (RTM) platform is 300 nucleotides. This may become a limitation in that insertion sites will not be able to be unambiguously identified in specific repeat sequence motifs in target genomes. It is therefore the aim of the present invention to provide a method of transposon mutagenesis with insertion site sequencing of target genomes that addresses the abovementioned drawbacks. In a first aspect of the invention there is provided a method of transposon insertion sequencing (TIS) capable of resolving insertion sites in or around repeat elements according to claim 1. Typically the DNA fragmentation is by physical sheering methods or enzymatic methods. The DNA amplification is performed by PCR. As such, this Long-Read method of Transposon Insertion Site Sequencing (LoRTIS) allows all the transposon insertion sites to be assigned uniquely, even those within the repeating genetic elements in target genomes Typically, the library of mutant cells can be prokaryotic or eukaryotic cells. In one embodiment the library of mutant cells is created by transformation or conjugation. In one embodiment the mutant cell library is subject to one or more conditions or environmental changes which can separate the mutants according to phenotypic differences. In one embodiment the mutant cell library is subject to a selective condition which can be any condition where bacterial survival or growth can be measured. Typically, the library of mutants is subject to a physical or chemical stress. Further typically the physical stresses include temperature stress, restraint environment such as a biofilm environment, reducing atmosphere or similar. In one embodiment of the invention a chemical stress includes a growth medium that includes one or more drug compounds, acids, oxidising agents or disinfectants. The DNA fragment for long-read sequencing is more than 300 bp. Further typically the long-read fragments are over 600 bp. In one embodiment the DNA fragments for long-read are over 600 bp. Typically, the long-read fragments are more than 300 10,000 bp. Further typically the long-read fragments are 600-15,000 bp. The person skilled in the art will appreciate that long-reads can easily be generated that are several lOOkb long. Typically, the long-read fragments in this technique are around 15,000 bp. In one embodiment the Oxford Nanopore Technologies’ (ONT) (RTM) nanopore long read sequencing technique is used. In one embodiment Pacific Biosciences’ (PacBio) (RTM) single-molecule real-time (SMRT)sequencing is used. The person skilled in the art will appreciate that other long read sequencing techniques can be used. In particular, incorporation or modification of the DNA fragments with adapters can render the same suitable for various long read techniques. In one embodiment the cells are prokaryotic cells. Typically, the cells are bacterial cells. In one embodiment the cells are eukaryotic cells. As such, the present method can be used to uniquely assign insertion sites within repeated genetic elements in bacteria that is not possible with short reads. In addition, the present method can resolve transposon insertion sites in eukaryotic cells. In a second embodiment of the invention there is provided a method of TIS, said method including the steps of preparing a pool of cells each containing a randomly inserted transposon. Specific embodiments of the invention are now described with reference to the following figures (all data produced from the same transposon mutant library, made in E. z-o / z BW25113) wherein: Figure 1 is a scatter chart where the number of sequence-reads per gene is plotted for replicate experiments using TraDIS-A^rm and LoRTIS demonstrating the experimental reproducibility for both technologies; Figure 2 shows a genetic map of the relative gene positions compared to TraDIS-Xpress and LoRTIS data. This demonstrates that the distribution of mapped reads for LoRTIS is similar to those generated by TraDIS-Ap / w; Figure 3 shows long reads that map uniquely within a single repeat of over 5 kb in length (one of seven ribosomal RNA gene clusters). A genetic map of the relative gene positions is shown at the bottom of the panel; Figure 4 shows a distribution map of the long reads that are uniquely aligned to ribosomal repeat regions; and Figure 5 shows the reproducibility of different replicates analysed on one nanopore flow cell and demultiplexed. We have overcome the limitations of conventional transposon insertion site sequencing by developing a method using nanopore flow cells and have produced long reads of more than 10,000 nucleotides. This new method has several advantages over conventional methods of transposon insertion site sequencing including: (i) increased read-length, (ii) lower cost and flexibility of choice of sequencing machines; (iii), when the flow cell has been used for a limited number of samples it can be re-used, as unused pores remain active, thereby reducing costs further; this is not the case for the Illumina flow cell which requires renewal after each sequencing. Our results show that LoRTIS is as reproducible as the Illumina (RTM)-based TraDIS-A / >7w, and is more cost effective, particularly for a small number of experiments. The new method can process multiple samples simultaneously (in “multiplex”) and generates sequence reads of long or short length, depending on the DNA fragment preparation and size selection. LoRTIS is a novel method that is robust and can uniquely assign transposon insert sites within the repeated genetic elements of candidate genomes. Furthermore, the lower cost and small size of the nucleotide sequencing machines (the ONT MinlON (RTM) sequencer is approximately the size of an office stapler) simplifies implementation of the approach where access to sophisticated laboratory set-ups is lacking. Transposon mutagenesis with sequencing has been used since 2009 to identify the repertoire of essential genes in many bacterial and eukaryotic organisms. Furthermore, the general approach has been used to identify the genes important for survival and growth when candidate transposon-mutant libraries are subjected to stress conditions. Transposon mutagenesis with sequencing provides unprecedented and unparalleled coverage of genome mutations at base-pair level resolution that other molecular techniques have not been able to achieve so far. Transposon insertion sequence methods may be used to examine the mechanism of resistance to a particular stress condition, for these experiments, transposon mutant libraries are grown under the stress condition and compared to a control (which may be the same library not exposed to the stress condition). The DNA of mutant libraries is extracted and, using a customised protocol and customised oligonucleotides, nucleotide sequence reads are generated from the known nucleotide sequences of the mutating transposon into the adjacent genome. Transposon insertion sites are identified precisely from where the nucleotide sequences of the generated reads match those of a reference genome. Since 2009, this method has helped the scientific community to unravel novel genotypephenotype associations and enhanced scientific understanding. All genome-wide TIS methods developed so far utilise either Illumina (RTM) or Ion torrent sequencing platforms and may be limited by the nucleotide sequence read-length of these respective technologies. To overcome these limitations, we have developed a method that utilises the ONT (RTM) platform to generate long-reads from the transposon insertion sites in a target genome (although other recent nucleotide sequencing methods could be used). The ONT MinlON (RTM) sequencers are small (one could fit inside a clothes pocket), portable, cost-effective and are easy to operate in any laboratory or in the field. We have used these Oxford Nanopore devices to conduct TIS experiments and show that the new approach produces a range of reads from 300bp to 14000 bp which can overcome the limitations associated with short reads approaches. Methods DNA extraction for LoRTISLibrary preparation The E. coli strain BW25113 transposon mutant library consisting of over 300 000 mutants which was used in this work has been described (Yasir et al., 2020). DNA was extracted from this transposon mutant library using the Zymogen DNA extraction kit according to manufacturer’s instructions. DNA fragment preparation for LoRTIS One microgram of DNA was tagmented using 1 pl of MuSeek Library Preparation Kit™ enzyme mix at 30°C for 5 minutes. The tagmented DNA was cleaned using Agencourt® AMPure ® XP (Beckman-Coulter (RTM)). LongAmp® Taq DNA Polymerase (New England BioLabs Inc. (RTM)) was used to amplify the transposon specific DNA fragments with biotinylated oligonucleotides specific to the transposon. DNA fragments incorporating the transposon were enriched using the streptavidin-coupled Dynabeads™ kilobaseBINDER™ Kit (Thermo-Fisher (RTM)). Another PCR, this time nested, was performed using LongAmp® Taq DNA Polymerase (New England BioLabs Inc. (RTM)) to generate replicate DNA fragments that are not attached to the Dynabeads™. The PCR product was cleaned using Agencourt® AMPure ® XP and a Ligation Sequencing Kit (SQK-LSK109; ONT) was used to ligate sequencing adapters. These DNA fragments with sequencing adapters were loaded on to Oxford Nanopore MinlON (RTM) sequencer to generate sequencing reads. Bioinformatics Results from short reads data from Illumina (RTM) were analysed using Bio-TraDIS (version 1.4.1) (Barquist et al., 2016). FAST5 format data files from the MinlON (RTM) sequencer were processed using the Guppy Basecalling software (version 3.6.0) running in High Accuracy Calling (HAC) mode to generate FASTQ files. Using QCat (version 1.1.0), the resulting FASTQ format data files were demultiplexed and transposon-specific nucleotide sequences were trimmed. The trimmed reads were mapped to Uscherichia coll BW25113 reference genome (CP009273) using Minimap2 (version 2.17-r941) to generate plot files output similar to short reads data (Lott et al. 2022). The insertion patterns at candidate loci were inspected visually using Artemis (version 18.1.0) (Carver et al. 2012) which was also used to capture images for figures. Gene essentiality was determined using the tradis_essentiality.R script from BioTradis and a correlation chart was plotted for the replicates after normalising the reads using Excel (Microsoft Inc). The correlation between short-read T1S and LoRTIS replicates was calculated using Spearman's rank correlation coefficient. Results LoRTIS data generation is reproducible The number of nucleotide sequence reads generated by different preparations of LoRTIS that match to specific genes shows a strong correlation, confirming the reproducibility of the methodology (Figure 1). Figure 1 is a scatter chart comparing duplicate data sets for LoRTIS and for short-read TIS. The short-read TIS data show a higher correlation between data sets, but these included many more sequence reads. The Spearman correlation coefficient for the LoRTIS data shown in Figi is 0.93, which can be improved by including more long-read sequences. The figure 1 scatter chart compares the number of nucleotide sequence-reads that match each gene in the reference genome for duplicate samples of short-read TIS data (blue) and for LoRTIS data (red). Each point is the number of sequence reads that matched a gene for duplicate data set 1 (x-axis) against duplicate data set 2 (y-axis). The nearer the points are to the diagonal the closer the correlation. The two short-read TIS data sets included more nucleotide sequence reads than for LoRTIS, and consequently these larger data sets are more closely correlated. Correlation between the LoRTIS data sets can be improved by including more sequence reads. Candidate Essential genes and those important for growth To confirm that the nucleotide sequence reads from LoRTIS are indeed generated from transposon insertion sites, very few, if any reads, should map within genes that are likely to be essential or that are important for growth, as transposon insertions into such genes results in the loss of these mutants from the transposon mutant library. Of 390 genes in this category that were identified by TraDIS-A^rm, 331 (85%) were also identified by LoRTIS, thus, confirming that the LoRTIS reads are generated primarily from the transposon copies present in the transposon mutant library. This is exemplified by Figure 2 which shows that the distribution of matched nucleotide sequence reads generated by LoRTIS is very similar to that generated using TraDIS-J^r^ across this relatively short section of the genome. From this it is apparent that no sequence reads matched within the candidate essential genes groS and groL, whereas an abundance of reads matched within the dcuA,fxsA,jjeH and jyeJ genes when using either LoRTIS or TraDIS-A^rm, confirming that LoRTIS is at least equal to TraDIS-J^ftrm in this respect. The genetic map in Figure 2 shows the relative gene positions at the bottom of the panel. White arrowhead boxes represent the position of genes and blue arrowhead boxes represent encoded proteins from the genetic code. Above this, each row of vertical red or blue lines indicates the position of mapped reads and the height of the bar represents the relative number of reads mapped. Red and blue differentiate the orientation of the transposon insertions. The top row shows TraDIS-J^n?^ data, the with LoRTIS data below, demonstrating the similarity of the data generated by LoRTIS to that generated by TraDIS-Apr^r. The advantages of LoRTIS compared to short-reads TIS The long nucleotide sequence reads generated using LoRTIS are particularly helpful when the genome size of the organisms is large (for example — as is found in a eukaryotic organism) and / or when there are repeating genetic elements present in the genome that need to be analysed. The long sequence reads generated by LoRTIS can identify transposon insertions uniquely within a particular repeated genetic element within a target genome. This resolution is achieved either by the long read extending into the unique flanking sequences, or by spanning sufficient polymorphisms within the repeated sequence to allow a specific match. To exemplify this special feature of LoRTIS, sequence reads were found that map uniquely within a single copy of the ribosomal RNA gene clusters in the E. coll BW25113 genome. There are seven of these gene clusters in the E. colt BW25113 genome, each over 5 kilobases in length and each containing two ribosomal RNA genes that are extremely similar between the seven clusters. Therefore, this is an excellent test of the performance of LoRTIS. Figure 3 illustrates how the long sequence reads generated by LoRTIS matched uniquely to a single rRNA gene cluster. For this, within the Bio-Tradis workflow, Minimap2 was used to match sequence reads to the reference genome, then Samtools was used to remove any matches that had more than one location with similar scores. The location of each uniquely matched read is then provided by the results from BioTradis. Figure 3 illustrates many long sequence reads that have been matched uniquely to the rrsA-rrlh rRNA genes. While most of these long reads are only around 2 kilobases in length (Figure 4) this is sufficient to match them uniquely to the reference genome because they span across polymorphic regions that exist between repeated clusters or extend into unique flanking nucleotide sequences (Figure 3). Regarding Figure 3, in the E. coll genome, the longest repeating nucleotide sequences are those of the ribosomal RNA genes, of which there are seven copies of over 5 kb in length. These are, therefore, a good test of the ability of LoRTIS to identify insertion sites uniquely within long repeating sequences. A genetic map of the rRNA genes is shown at the bottom of the panel as blue arrowhead boxes. Above this, each panel depicts the long reads, each shown as a fine horizontal line, that map uniquely within this repeat element. Underneath each depiction of the long-reads, the line graph indicates the depth of coverage. Figure 4 shows a bar chart where nucleotide sequence reads generated usingLoRTIS range in size from a few hundred to 10, 000 base pairs. Most of the reads are less than 2000 base pairs but nevertheless can be matched uniquely to a single copy of the repeated ribosomal RNA gene cluster due to polymorphisms within the repeat or because the reads extend into unique nucleotide sequences that flank the repeat. Multiplexing of LoRTIS experiments Up to 96 sequence adapters are available to multiplex DNA samples on an Oxford Nanopore MinlON (RTM) sequencer. Four different replicates were prepared, and nucleotide sequence reads generated using a MinlON (RTM) sequencer. The results show that each of the four samples were demultiplexed correctly, and the data obtained from each replicate was reproducible (Figure 5). This confirms that, using LoRTIS, several different experimental samples may be sequenced concurrently as a single experiment on one sequencer. Regarding Figure 5 the plots showing the relative locations of nucleotide sequence reads that matched the nucleotide sequences of the E. coll BW25113 reference genome. Sequence reads were generated concurrently using LoRTIS incorporating multiplexing on an Oxford Nanopore MinlON (RTM) sequencer. The outer circle indicates the genome coordinates and the four inner circles each represent data generated from a different replicate. Each dot locates a transposon insertion site, and the distance from the respective baseline represents the number of sequence reads that matched at each site. In this figure the reference genome is shown circularised to reflect the natural state of the bacterial chromosome. The data confirm the reproducibility between replicates and the feasibility of incorporating nucleotide sequence indices to enable sample multiplexing. The utility of the LoRTIS technology A direct comparison of the costs of TraDIS-A^rm sequencing (using the Illumina (RTM) sequencing platform) with LoRTIS (using the Oxford Nanopore Technologies (RTM) platform) sequencing is not possible because the cost of both vary depending on a number of variables in the design of any individual experiment. A comparison of other factors has been summarised in Table 1. However, Oxford Nanopore sequencing has the added advantage that a LoRTIS sequencing run may be performed even on a single sample without the cluster registration problems encountered with Illumina (RTM) technologies. In addition, the cost and maintenance of any sequencing machine needs to be taken into account and can be a substantial initial outlay for Illumina (RTM) -based sequencing approaches. Furthermore, there can be additional cost / constraints in transporting samples to specific DNA sequencing facilities from the field — specifically from remote regions, or in developing nations that may have limited access to specialised sequencing units. In comparison, an Oxford Nanopore MinlON (RTM) sequencer is highly portable (being pocket-sized, it can be transported to remote regions), can be powered from a laptop computer and is highly cost-effective. . Figure 6 shows an overview of the LoRTIS technique according to one embodiment using Oxford Nanopore Technologies (RTM) long read sequencing. As illustrated, genomic DNA is extracted from a transposon mutant library consisting of a pool of many copies of different mutants. Genomic DNA is fragmented, and sequencing adapters added simultaneously in a process termed “tagmentation.” A PCR is then performed using biotin-modified oligonucleotides that incorporate all nucleotide sequences necessary for sequence read generation (for flow-cell binding and indices) at their 5’-ends and nucleotide sequences that specifically bind to the transposon at their 3-end. These PCR products generated from the transposon can then be enriched by capture of the biotin moiety to magnetic streptavidin-coupled Dynabeads™. A nested PCR is then performed to generate DNA fragment replicates that are no longer attached to the Dynabeads™. These fragments are then applied to a nanopore sequencer, such as the Oxford Nanopore MinlON (RTM), to generate sequence reads from the captured transposon into the adjacent insertion site. The transposon insertion sites are then mapped to a reference genome sequence by identifying where the nucleotide sequence of the reads matches that of the reference genome. This final mapping step is performed using the biotradis software with modifications described above. Summary 5 TIS technologies provide a very powerful and robust genomics methodology allowing the simultaneous assay of every gene in a target genome for associating genotype with phenotype. Adapting nanopore, long-read nucleotide sequencing to TIS, has improved the technology by enabling the unique identification of transposon insertion sites within repeating genetic elements of target genomes. Such 10 unique identification of insertion sites is not possible using short-read sequencing technologies that cannot span the length of a repeated genomic sequence. In addition, the portability and relative low cost of Oxford Nanopore-based sequencing technologies increases the utility and accessibility of this powerful genomics technology (Table 1). 15 Table 1. Comparison between Illumina and Nanopore costs Illumina (RTM) Nanopore Instrument price £80.000-£500,000 £500 Laboratory time (for generation of sequences) 20-56 h 48 h Input DNA required 1-50 ng 100-1000 ng Library preparation ~£50 ~£120 Long reads No Yes Requirement for diversity of sequencing library Yes No Potential for single sample run No (challenging) Yes Figures 100000 10000 1000 1 1 10 100 1000 10000 100000 Figure 1: Scatter chart comparing short-read TIS and LoRTIS dataof Scatter chart comparing the number of nucleotide sequence-reads that matched each gene in the reference genome for duplicate samples of short-read TIS data (blue) and for LoRTIS data (red). Each point is the number of sequence reads that matched a gene for duplicate data set 1 (x-axis) against duplicate data set 2 (y-axis). The nearer the points are to the diagonal the closer the correlation. The two shortread TIS data sets included more nucleotide sequence reads than for LoRTIS, and, consequently these larger data sets are more closely correlated. Correlation between the LoRTIS data sets can be improved by including more sequence reads. Figure 2: Comparison of Transposon insertions sites of TraDIS-Jewess and LoRTIS A genetic map of the relative gene positions is shown at the bottom of the panel. White arrowhead boxes represent the position of genes and blue arrowhead boxes represent encoded proteins from the genetic code. Above this, each row of vertical red or blue lines indicates the position of mapped reads and the height of the bar represents the relative number of reads mapped. Red and blue differentiate the orientation of the transposon insertions. The top row shows TraDIS-A^rm data, the with LoRTIS data below, demonstrating the similarity of the data generated by LoRTIS to that generated by TraDIS-Aprm. Figure 3. Long reads mapped uniquely to a single element of repeating ribosomal RNA gene cluster of over 5 kb 5 In the E. coli genome, the longest repeating nucleotide sequences are those of the ribosomal RNA genes, of which there are seven copies of over 5 kb in length. These are, therefore, a good test of the ability of LoRTIS to identify insertion sites uniquely within long repeating sequences. A genetic map of the rRNA genes is shown at the bottom of the panel as blue arrowhead boxes. Above this, each panel depicts the 10 long reads, each shown as a fine horizontal line, that map uniquely within this repeat element. Underneath each depiction of the long-reads, the line graph indicates the depth of coverage. 15 Distribution of length for unique reads Figure 4: Bar chart showing the number of nucleotide sequence reads distributed according to their length that match uniquely within one copy of 5 the repeating ribosomal RNA gene clusters. Nucleotide sequence reads generated using LoRTIS ranged in size from a few hundred to 10, 000 base pairs. Most of the reads are less than 2000 base pairs but nevertheless can be matched uniquely to a single copy of the repeated ribosomal RNA gene cluster due to polymorphisms within the repeat or because the reads 10 extend into unique nucleotide sequences that flank the repeat. Figure 5 Reproducibility of different replicates after generation of sequence reads concurrently using an Oxford Nanopore MinlON sequencer followed 5 by demultiplexing. Plots showing the relative locations of nucleotide sequence reads that matched the nucleotide sequences of the E. co / / BW25113 reference genome. Sequence reads were generated concurrently using LoRTIS incorporating multiplexing on an Oxford Nanopore MinlON (RTM) sequencer. The outer circle indicates the genome 10 coordinates and the four inner circles each represent data generated from a different replicate. Each dot locates a transposon insertion site, and the distance from the respective baseline represents the number of sequence reads that matched at each site. In this figure the reference genome is shown circularised to reflect the natural state of the bacterial chromosome. The data confirm the reproducibility between replicates and the feasibility of incorporating nucleotide sequence indices to enable sample multiplexing. LoRTIS Overview Pool of 1000$ Transposon mutants Matching of nucleotide sequence reads with reference genome maps the precise location of transposon insertion sites Genomic ONA from pool of mutants PCR products generated from transposon nucleotide sequences incorporate 5’biotin adapter oligonucleotide biotinylated "Nested" PCR Extraction of genomic DMA sequencing adapters □ligo hybridises within captured transposon end. Nanopore sequencing / adapters oligonucleotide binding to transposon end replicates of DNA fragments enriched for transposon nucleotide sequences that are unattached to streptavidin bead DNA fragments enriched for transposon nucleotide sequences streptavidin bead enrichment Narsopero nucleotide sequence determination and analysis nanopore adapter sequences streptavidin bead — — Reference genome Figure 6. Overview of the LoRTIS method. Genomic DNA is extracted from a transposon mutant library consisting of a pool of many copies of different mutants. Genomic DNA is fragmented, and sequencing adapters added simultaneously using “tagmentation.” A PCR is then performed using biotin-modified oligonucleotides that incorporate sequencing adapters at their 5’-ends and nucleotide sequences that specifically bind to the transposon at their 3-end. These PCR products generated from the transposon are enriched using streptavidin-coupled Dynabeads™. A nested PCR is then performed to generate DNA fragment replicates unattached to the Dynabeads™. These fragments are applied to a nanopore sequencer to generate sequence reads from the captured transposon into the adjacent insertion site. The transposon insertion sites are then mapped to a reference genome sequence by identifying where the nucleotide sequence of the reads matches that of the reference genome. This final mapping step is performed using the “biotradis” software with modifications (as described above).

Claims

1. A method of transposon insertion sequencing (TIS) capable of resolving insertion sites in or around repeat elements, said method including the steps of:a) preparing a library of mutant cells by insertion of one or more transposon sequences;b) extracting genomic DNA from said library of mutant cells;c) fragmenting the extracted genomic DNA and adding sequencing adapters using tagmentation;d) amplifying transposon specific DNA fragments via PCR using biotin-modified oligonucleotides that include sequencing adapters at a 5’ end and nucleotide sequences that specifically bind to the transposon at a 3’ end and that are paired with second oligonucleotides specific to the sequencing adapters added in step c);e) enriching the amplified DNA fragments incorporating the transposon by capturing biotin moieties on magnetic streptavidin-coupled beads;f) performing nested PCR using oligonucleotides that include nanopore adapter sequences at their 5’ end and nucleotide sequences that specifically bind to the transposon at their 3’ end and oligonucleotides that include nucleotide sequences that specifically bind to the sequencing adaptors added in step c) at their 3’ end and nanopore adapter sequences at their 5’ end, to generate DNA fragment replicates that are no longer attached to the magnetic streptavidin-coupled beads;g) nanopore sequencing the amplified DNA fragments using a long-read sequencing method, wherein the DNA fragments for long read sequencing are more than 300bp; h) mapping transposon insertion sites in the sequenced DNA fragments to a reference genome sequence by identifying nucleotide sequences of the long read sequences that match that of the reference genome sequence.

2. A method according to claim 1 wherein the Long-Read method of Transposon Insertion Site Sequencing (LoRTIS) enables unique mapping of transposon insertion sites within repeated nucleotide sequences in target genomes.

3. A method according to claim 1 wherein the library of mutant cells can be prokaryotic or eukaryotic cells.

4. A method according to claim 3 wherein the library of mutant cells is created by transformation or conjugation.

5. A method according to claim 1 wherein the library of mutant cells can be sorted according to a phenotype.

6. A method according to claim 1 wherein the library of mutant cells is subject to a selective condition.

7. A method according to claim 5 wherein the library of mutant cells is subject to a physical and / or chemical stress.

8. A method according to claim 7 wherein the chemical stress includes a growth medium that includes one or more drug compounds.

9. A method according to claim 1 wherein the DNA fragments for long-read sequencing are over 600 bp.

10. A method according to claim 9 wherein the DNA fragments for long-read sequencing are 600-15,000 bp.

11. A method according to claim 9 wherein the DNA fragments for long-read sequencing are around 15,000 bp.