High-throughput methods for identifying and recovering gene clusters of interest

The method of nucleic acid fractionation and sequencing with low depth sequencing in clone libraries addresses the inefficiencies of existing gene cluster identification, enabling efficient and cost-effective recovery of complete gene clusters.

WO2025163168A1PCT designated stage Publication Date: 2025-08-07GENERARE BIOSCIENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/052574
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2025-01-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing methods for identifying and isolating biosynthetic gene clusters are inefficient, costly, and biased towards predetermined sequences, often yielding truncated clusters and requiring extensive resource consumption.

Method used

A method involving nucleic acid fractionation and sequencing of clone libraries, using a divide-and-conquer pooling strategy with low sequencing depth to identify and recover gene clusters without the need for degenerate primers, allowing for unbiased detection of diverse gene clusters.

Benefits of technology

Enables efficient, cost-effective identification and recovery of complete gene clusters, reducing resource consumption and overcoming primer bias, suitable for large clone libraries and diverse gene clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025052574_07082025_PF_FP_ABST
    Figure EP2025052574_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for identifying and recovering a gene cluster of interest.
Need to check novelty before this filing date? Find Prior Art

Description

HIGH-THROUGHPUT METHODS FOR IDENTIFYING AND RECOVERINGGENE CEUSTERS OF INTERESTFIEED OF INVENTION

[0001] The present invention relates to methods for identifying / localizing gene clusters of interest. In particular, the invention relates to the discovery of novel biosynthetic gene clusters (BGC) in prokaryotic and eukaryotic cells.

[0002] The present invention is thus applicable to the screening of clone libraries.

[0003] Accordingly, the present invention relates to methods which can be compatible with high-throughput studies, and which may be compatible with the detection of a plurality of gene clusters, in an affordable and efficient manner.BACKGROUND OF INVENTION

[0004] There is a general need for identifying novel products, in particular secondary metabolites, from host organisms.

[0005] Such metabolites may include clinically valuable molecules, such as antibacterial and antifungal products, in particular in the context of antimicrobial resistance, as reported in Eewis (“The Science of Antibiotic Discovery”; Cell 181, 2020).

[0006] However, such metabolites are also challenging to characterize and isolate.

[0007] Such metabolites tend to be enzvmatically biosynthesized by the products of more than one gene, often grouped into gene clusters.

[0008] Moreover, such gene clusters may not be readily expressed under laboratory conditions. Alternatively, such gene clusters may be differentially expressed when cloned in microorganisms belonging to different species, or even to a different genus.

[0009] Such gene clusters may also be differentially expressed when corresponding hosts are cultured in different conditions.

[0010] Also, such gene clusters include a plurality of genes, thus leading to high- molecular weight nucleic acids which can be difficult to clone or costly to synthetize or costly to sequence in an unbiased manner.

[0011] There is thus a need for novel methods which are capable of identifying, cloning and recovering such gene clusters, especially in microorganisms, in an efficient and cost- friendly manner.

[0012] One strategy consists in cloning the entirety of the genome which is susceptible to contain the gene cluster in the form of a random genomic fragment library, and then to screen the clones to localize and recover the fragment of interest.

[0013] Screening of a clone library to find the localization of a few clones of interest is still a time and resource consuming step. This is typically done for instance by using a PCR strategy and primers directed against conserved regions of such clusters. Amplicons may then be sequenced to provide additional insight at the nature of the amplified material, which may be used for dereplication.

[0014] On the other hand, performing PCR on each individual clone would be impractical. Hence, it has been suggested to reduce the number of PCR steps required to localize a clone in a given clone library.

[0015] Owen et al. (“Mapping gene clusters within arrayed metagenomic libraries to expand the structural diversity of biomedically relevant natural products”; PNAS, vol. 110, no. 29, 2013) teaches multiplex sequencing of barcoded PCR amplicons from cosmid clone libraries.

[0016] Lam et al. (“Evaluation of a Pooled Strategy for High-Throughput Sequencing of Cosmid Clones from Metagenomic Libraries”; 2014 PLoS ONE 9(6): e98968.doi: 10.1371 / journal. pone.0098968) teaches a pooled sequencing strategy involving combining clones into one sample for sequencing and assembly, andsubsequently using previously obtained Sanger end-tags to retrieve specific clone sequences.

[0017] Libis et al. (“Uncovering the biosynthetic potential of rare metagenomic DNA using co-occurrence network analysis of targeted sequences”; Nat Commun., 10( 1 ):3848, 2019) and WO2021041397A2 teach the amplification of clone pools using degenerate PCR primers targeting conserved regions, combined with a prediction algorithm requiring the identification of domain pairs showing non-random occurrence.

[0018] Libis et al. (“Multiplexed mobilization and expression of biosynthetic gene clusters”; 13:5256, 2022) teaches a variant of the previous method, relying also on degenerate primer amplification, but further including a step of cloning directly large and intact gene clusters into a final shuttle vector.

[0019] Crits-Christoph et al. (‘ ‘Novel soil bacteria possess diverse genes for secondary metabolite biosynthesis”; Nature, 558, 440-444, 2018) illustrates a case wherein the nearcomplete genome of microorganisms from the soil ecosystem is reconstituted de novo using genome-resolved metagenomic methods. The performance of degenerate primers is also assessed to detect gene clusters of interest; thus leading to the conclusion that such primers may fail to amplify genetically divergent sequences.

[0020] Hence, although such methods remain efficient in identifying gene clusters coding for clinically valuable metabolites, they are still highly dependent upon the choice of primers, and therefore tend to be limited toward predetermined clusters; in particular certain biosynthetic gene cluster (BGC) classes, and other conserved sequences on which the primers can be bound.

[0021] Moreover, PCR screens aiming to physically recover the clones corresponding to a BGC from a pool are time consuming and often yield truncated BGCs containing the primer binding sites but not all of the sequence of interest.

[0022] Accordingly, there still remains a need for novel methods for identifying, recovering and isolating, gene clusters of interest in cells and other microorganisms, in an unbiased, efficient and complete manner.

[0023] There still remains a need to compensate for the bias induced by the requirement for degenerate primers in amplicon-based sequencing.

[0024] The invention has for purpose to meet the above-mentioned needs.SUMMARY

[0025] This invention thus relates to a method for identifying at least one gene cluster of interest, comprising steps of: a) providing a clone library comprising a plurality of clones with high- molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; b) partitioning the clone library into at least a first and second plurality of separate partitions: the first plurality of separate partitions comprising a plurality of clones, the plurality of separate partitions thereby forming a plurality of clone pools; the second plurality of separate partitions comprising a plurality of clones from clone pools, thereby forming a plurality of clone stacks; wherein the high-molecular weight nucleic acids from the clone pools and the clone stacks are fractionated; c) determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acids in the clone pools and in the clone stacks, thereby providing a plurality of sequence reads, the sequencing depth corresponding to said plurality of sequence reads being of less than 30x, for example less than 20x; thereby identifying the at least one gene cluster of interest.

[0026] The method is thus suitable for identifying gene cluster(s) of interest in the clone library, and in the cell(s) from which the clone library was previously obtained.

[0027] The method is thus suitable for localizing gene cluster(s) of interest in the clone library, and in the cell(s) from which the clone library was previously obtained.

[0028] The method is thus suitable for recovering gene cluster(s) of interest in the clone library, and in the cell(s) from which the clone library was previously obtained.DEFINITIONS

[0029] In the present invention, the following terms have the following meanings:

[0030] The terms “comprising” , “consisting essentially of ” and “consisting of’ may be replaced with either of the other two terms. The terms and expressions which have been employed are used as terms of description and not of limitation, and use of such terms and expressions do not exclude any equivalents of the features shown and described or portions thereof, and various modifications are possible within the scope of the claimed technology.

[0031] The term “a” or “an” can refer to one of or a plurality of the elements it modifies (e.g., “a reagent” can mean one or more reagents) unless it is contextually clear either one of the elements or more than one of the elements is described. As used herein, the expression “a” or “at least one” encompasses “one”, or “more than one”; which encompasses a “plurality”, such as two, or more than two, which may encompass, three, four, five, six or even more than six.

[0032] The term “about” as used herein refers to a value within 10% of the underlying parameter (i.e., plus or minus 10%), and use of the term “about” at the beginning of a string of values modifies each of the values (i.e., “about 1, 2 and 3” refers to about 1, about 2 and about 3).

[0033] The term ' nucleic acid fractionation” refers to the process of breaking a high- molecular nucleic acid (e.g. generally a large target DNA) randomly into smaller fragments. Nucleic acid fractionation may be used prior to cloning nucleic acid fragments into a vector or prior to sequencing the fragments.

[0034] The term ' Shotgun sequencing” refers to the sequencing of individual fractionated nucleic acids, for example DNA or RNA nucleic acids. The term is not meant to be limitative toward one specific type of shotgun sequencing, although the fractionation step is generally achieved using either enzymes, or through mechanical means. This process thus entails breaking a high-molecular nucleic acid (e.g. generally a large target DNA) randomly into smaller fragments and end-sequencing these smaller fragments.

[0035] The term ' Sequence depth” or “read depth” refers to the number of times a given base or nucleotide in a nucleic acid is read during a given sequencing process. Accordingly, the term may correspond to the number of unique sequence reads that align to a region in a reference sequence, or reference genome, or de novo assembly. In a non- limitative manner, sequence depth may be equal to or superior to O.Olx, O.lx, 0.5x, lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, lOOx or more. In a non-limitative manner, sequence depth may also be equal to or inferior to lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, lOOx. Sequence depth values which may be considered as standard will depend upon the type of sequencing technique. For example, a sequence depth considered as “standard” or “high” may thus correspond to 20x or more; for example 30x or more. For example, a sequence depth considered as “low” may thus correspond to lower than 20x; for example lower than lOx, lower than 5x or even lower than lx.

[0036] The term “completeness” or “ genomic coverage” refers to the percentage of base pairs corresponding to the reference sequence (e.g. the gene or gene cluster) which is covered by sequencing. The completeness of sequencing, may thus be expressed as a percentage of the entire length of a reference sequence, or a portion of said reference sequence which is covered by the nucleic acid sequence reads. The completeness of sequencing may further be quantified for a given length of nucleic acid sequence portions;for example as a percentage of said reference sequence which is covered by 1-kb nucleic acid sequence portions.

[0037] The term “gene” refers to gene the biologic unit of heredity, self-reproducing and located at a definite position (locus) on a particular chromosome. In one embodiment the particular chromosome is a bacterial chromosome. The term bacterial chromosome is used interchangeably herein with the term bacterial genome. The term “gene cluster” as used herein refers to a group of genes located closely together on the same chromosome whose products play a coordinated role in a specific aspect of cellular primary or secondary metabolism.

[0038] The term “Gene clusters” refers to sets of genes having a common function or function product. Genes are typically found within physical proximity to each other within genomic DNA (e.g., within one centiMorgan (cM)). Gene clusters can occur in prokaryotic or eukaryotic cells.

[0039] The term “biosynthetic gene cluster” or “BGC” or “metabolite gene cluster” refers to clusters (i.e two or more) of physically adjacent genes coding for metabolites, such as secondary metabolites, which hence comprise polynucleotide sequences encoding the functions required for synthesis and activity of such metabolites. In a non-limitative manner, this term may thus encompass those BGCs which comprise polynucleotide sequences encoding for one metabolite, or more than one (i.e. a plurality of) metabolites. In particular, BGCs which are considered herein include those standardized by the “Minimal Information about a Biosynthetic Gene cluster” (MIBiG) as defined by the genomic standards consortium, and / or antiSMASH. Accordingly, and unless stated otherwise, the term may include, in a non-limitative manner, BGCs for the synthesis of metabolites selected from the group consisting of: alkaloids, polyketides, saccharide, Non-Ribosomal peptides (NRPs), post-translationally modified peptides (RiPP), terpenes.

[0040] The terms “secondary metabolite” or “specialized metabolite” refer to compounds that are not involved in primary metabolism, and therefore differ from the more prevalent macromolecules such as proteins and nucleic acids that make up the basicmachinery of life. Many metabolites find important biotechnological applications in biomedical and drug discovery research, and in the agricultural, aquaculture and chemical industries. Examples include: antibacterial products such as penicillin, and daptomycin; antifungal products such as amphotericin; cholesterol-lowering products such as lovastatin; anticancer products such as bleomycin; and immune-modulating products including rapamycin, and cyclosporine.

[0041] The term "library ", as in "clone library " or "genomic fragment library " or “multigenomic fragment library " or “high molecular weight DNA inserts library ", as used herein refers to any collection of nucleic acids, in particular any collection of DNA”, that can stored and propagated in a population of micro-organisms (or clones) through the process of molecular cloning; which may thus include cDNA libraries, genomic or multigenomic libraries and / or any synthetic mutant library which may, for example, include variants of nucleic acid sequences, for example any mutated nucleic acid sequences derived from an endogenous sequence. Accordingly, and unless stated otherwise, the term “clone library " may thus indifferently apply to either a cloning vector comprising the one or more nucleic acids, or to the host cell comprising the said cloning vector. For example, the term “clone library " may refer to cloning vector replicating autonomously in a host cell, the cell being for example in the form of a prokaryotic cell, for example gram-positive or gram-negative bacteria.

[0042] The term “cloning vector ", as used herein, a nucleic acid molecule, such as a: cosmid, fosmid, phage, bacterial artificial chromosome (BAC), Pl -derived artificial chromosome (PAC), yeast artificial chromosome (YAC), fungal artificial chromosome (FAC), that has the capability of replicating autonomously in a host cell. Cloning vectors typically contain one or a small number of restriction endonuclease recognition sites that allow insertion of a nucleic acid molecule in a determinable fashion without loss of an essential biological function of the vector, as well as nucleotide sequences encoding a marker gene that is suitable for use in the identification and selection of cells transformed with the cloning vector. Marker genes typically include genes that provide chloramphenicol resistance or apramycin resistance.

[0043] As used herein, the term “host cell” refers to any cell capable of replicating and / or transcribing and / or translating a heterologous gene. Thus, a “host cell” refers to any prokaryotic cell (including but not limited to E. coll) or eukaryotic cell (including but not limited to yeast cells, fungal cells, mammalian cells, avian cells, amphibian cells, plant cells, fish cells, and insect cells).

[0044] The terms, “DNA library”, “genomic DNA library”, “multigenomic library”, “cDNA library” and “environmental DNA library” as used herein refer to such libraries that comprise a plurality of genetic constructs wherein each genetic construct comprises an insert polynucleotide sequence in the form of DNA. Each of these terms takes their common meaning as known and used in the art.

[0045] The term “polynucleotide(s)” or “nucleic acid” as used herein, means a single or double- stranded deoxyribonucleotide or ribonucleotide polymer of any length, and include as non-limiting examples, coding and non-coding sequences of a gene, sense and antisense sequences, exons, introns, genomic DNA, cDNA, pre-mRNA, in RNA, rRNA, siRNA, miRNA, tRNA, ribozymes, recombinant polynucleotides, isolated and purified naturally occurring DNA or RNA sequences, synthetic RNA and DNA sequences, nucleic acid probes, primers, fragments, genetic constructs, vectors and modified polynucleotides. Reference to nucleic acids, nucleic acid molecules, nucleotide sequences and polynucleotide sequences is to be similarly understood.DETAILED DESCRIPTION

[0046] The inventors are of the opinion that screening of a whole clone library to find the localization of a few clones of interest is a time and resource consuming step. The inventors originally speculated that the presence of gene clusters, in particular biosynthetic gene clusters, could be assessed with minimal effort, through a fractionation step of nucleic acids following by nucleic acid sequence determination; in particular shotgun sequencing methods.

[0047] The inventors thus reasoned that PCR and amplicon sequencing strategies, previously used to assess the presence of genes of interest, could be simply replaced altogether by a fractionation and sequencing step, even for the sequencing of clone libraries such as BACs, while remaining efficient in identifying novel and diverse gene clusters of interest.

[0048] Accordingly, the inventors now propose to combine a primer-free technique relying on nucleic acid fractionation and pooling-based approaches; in particular a multiplexed divide-and-conquer pooling strategy. This strategy exhibits some key advantages in comparison to alternative PCR screens described in the prior art.

[0049] As a primer-free technique, it is not limited to predetermined biosynthetic gene cluster (BGCs) classes, nor to the need for the targeted classes to harbor conserved sequences on which primers can be bound. In particular, all potential BGCs classes, or any other arbitrary DNA sequence, can be detected by shotgun sequencing whenever a related reference sequence is used for the alignment.

[0050] Advantageously, the methods of the invention can thus be applied to large clone libraries, comprising high-molecular weight nucleic acid, which may contain such gene clusters.

[0051] As opposed to PCR screens which often yields truncated gene clusters containing the primer binding sites but not all of the sequence of interest, the completeness of the region of interest captured on an insert can be assessed by shotgun sequencing without the need to first physically recover the corresponding clone from a pool. This step saves significant efforts wasted in recovering truncated sequences, which typically represent more than half of the clones recovered by PCR screens for large gene clusters.

[0052] The complete reference gene cluster sequence allows for selection and prioritization of gene clusters (e.g. BGCs) of interest as it may be used to assess the putative novelty and gene composition of a captured BGC. Furthermore, a more finegrained analysis of read data may also reveal the presence of novel gene inserts within captured BGCs that were absent from the reference.

[0053] According to a first main embodiment, the invention thus relates to a method for identifying at least one gene cluster of interest, comprising steps of: a) providing a clone library comprising a plurality of clones with high-molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; b) partitioning the clone library into at least a first and second plurality of separate partitions: the first plurality of separate partitions comprising a plurality of clones, the plurality of separate partitions thereby forming a plurality of clone pools; the second plurality of separate partitions comprising a clone from clone pools, thereby forming a plurality of clone stacks; wherein the high-molecular weight nucleic acids from the clone pools and the clone stacks are fractionated; c) determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acids in the clone pools and in the clone stacks; thereby identifying the at least one gene cluster of interest.

[0054] According to particular embodiments, the invention relates to a method for identifying at least one gene cluster of interest, comprising steps of: a) providing a clone library comprising a plurality of clones with high-molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise biosynthetic gene clusters of interest; b) partitioning the clone library into at least a first and second plurality of separate partitions:the first plurality of separate partitions comprising a plurality of clones, the plurality of separate partitions thereby forming a plurality of clone pools; the second plurality of separate partitions comprising a plurality of clones from clone pools, thereby forming a plurality of clone stacks; wherein the high-molecular weight nucleic acids from the clone pools and the clone stacks are fractionated; c) determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acids in the clone pools and in the clone stacks, thereby providing a plurality of sequence reads, the sequencing depth corresponding to said plurality of sequence reads being of less than 30x, for example less than 20x; thereby identifying the at least one gene cluster of interest.

[0055] Advantageously, the method of the invention is thus suitable for localizing said gene cluster(s) of interest in the clone library, and / or for localizing the gene cluster(s) of interest in a cell, or plurality of cells, from which the clone library was previously obtained.

[0056] Advantageously, the method of the invention is thus suitable for recovering said gene cluster(s) of interest in the clone library, and / or for recovering the gene cluster(s) of interest in a cell, or plurality of cells, from which the clone library was previously obtained.

[0057] According to a particular embodiment, the invention thus relates to method for identifying at least one gene cluster of interest, in particular a plurality of gene clusters of interest, in cell(s) or a clone library thereof, preferably a plurality of BGCs in eukaryotic or prokaryotic cell(s) or a clone library thereof, comprising steps of: a) providing a clone library comprising a plurality of clones with high- molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest;bl) partitioning the clone library into a first plurality of separate partitions comprising a plurality of clones, thereby forming a plurality of clone pools; b2) fractionating the high-molecular weight nucleic acid from the clone pools, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone pools, thereby determining the plurality of clone pools comprising at least one occurrence of said at least one gene cluster of interest; cl) partitioning the clone library into a second plurality of separate partitions, comprising a plurality of clones from clone pools, thereby forming a plurality of clone stacks; c2) fractionating the high-molecular weight nucleic acid from the clone stacks, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone stacks; d) determining the clones comprising the gene cluster(s) of interest in the library, based on the nucleic acid sequences from the clone stacks, thereby identifying the at least one gene cluster of interest.

[0058] According to said embodiment, determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid (e.g. in the clone pools and / or the clone stack) may comprise providing a plurality of sequence reads, the sequencing depth corresponding to said plurality of sequence reads being of less than 30x, for example less than 20x.

[0059] The second plurality of separate partitions may comprise a plurality of clones from clone pools, thereby forming a plurality of clone stacks. Advantageously, the second plurality of separate partitions (i.e. the second partition) does not necessarily comprise all the clones for which a nucleic acid sequence is determined in the first plurality (i.e. the first partition). Ideally, it may only comprise the clones from the first clone pools in which gene clusters of interest (e.g. BGCs) are statistically present. One advantage is thus to limit the sequencing to part of the clone library in the other partitions (e.g. the second partition, and optionally additional further partitions).

[0060] In view of the above, based on the nucleic acid sequences determined in the first plurality, the separate partitions which form the second plurality may thus be characterized in that they comprise all or only a part of the clones of the first plurality. For example those separate partitions may comprise all, or substantially all or only a part of the clones of first plurality.

[0061] According to a particular embodiment, the clone library is partitioned into at least a second plurality of separate partitions, comprising a clone from each clone pool, thereby forming a plurality of clone stacks.

[0062] When the method requires further partition steps, for example a third and fourth partition (thereby forming a third plurality and a fourth plurality), they may thus be optionally characterized in the same manner; that is, the separate partitions from each further partition step may comprise all or only a part of the clones of the previous partitions (e.g. of the first plurality, or alternatively of the second plurality, and / or any other partition in a preceding order).

[0063] According to a particular embodiment, the invention relates to method for identifying at least one gene cluster of interest in a cell, in particular a plurality of gene clusters of interest in cell(s) or a clone library thereof, preferably a plurality of BGCs in eukaryotic cell(s) or a clone library thereof, comprising steps of: a) providing a clone library comprising a plurality of clones with high-molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; bl) partitioning the clone library into a first plurality of separate partitions comprising a plurality of clones, thereby forming a plurality of clone pools; b2) fractionating the high-molecular weight nucleic acid from the clone pools; b3) determining the nucleic acid sequence of all or part of the fractionated high- molecular weight nucleic acid in the clone pools, thereby determining the plurality of clone pools comprising at least one occurrence of said at least one gene cluster of interest;cl) partitioning the clone library into a second plurality of separate partitions, comprising a clone from each clone pool, thereby forming a plurality of clone stacks; c2) fractionating the high-molecular weight nucleic acid from the clone stacks; c3) determining the nucleic acid sequence of all or part of the fractionated high- molecular weight nucleic acid in the clone stacks; d) determining the clones comprising the gene cluster(s) of interest in the library, based on the nucleic acid sequences from the clone stacks, thereby identifying the at least one gene cluster of interest.

[0064] According to said embodiment, steps b3) and / or c3) of determining nucleic acid sequences may comprise providing a plurality of sequence reads, the sequencing depth corresponding to said plurality of sequence reads being of less than 30x, for example less than 20x.

[0065] According to a particular embodiment, the method of the invention is for identifying a plurality of gene clusters in an eukaryotic cell or clone library thereof; said eukaryotic cell being in particular fungi, preferably filamentous fungi, for example belonging to the genus Aspergillus.

[0066] According to a particular embodiment, the method of the invention is for identifying a plurality of gene clusters in a clone library; said clone library being preferably derived from eukaryotic cells, the cells being in particular fungi, preferably filamentous fungi, for example belonging to the genus Aspergillus.

[0067] According to a particular embodiment, the method of the invention is for localizing a plurality of gene clusters in a clone library; said clone library being preferably obtained from eukaryotic cells, the cells being in particular fungi, preferably filamentous fungi, for example belonging to the genus Aspergillus.

[0068] According to some embodiments, the gene cluster(s) is / are biosynthetic gene cluster(s) (BGC). According to some other embodiments, the gene cluster(s) is / are not biosynthetic gene cluster(s) (BGC).

[0069] According to a particular embodiment, the method of the invention is for identifying a plurality of biosynthetic gene clusters (BGC).

[0070] According to a particular embodiment, the method of the invention is for identifying a plurality of biosynthetic gene clusters (BGC) in an eukaryotic cell.

[0071] According to a particular embodiment, the plurality of biosynthetic gene clusters (BGC) is for the synthesis of terpenes, NRPS, RiPP-like, T1 PKS, NRPS-like, melanin, butyrolactone, T3 PKS, T2 PKS, ectoine, NAPA A, betalactone, lanthipeptide class III, PKS-like, lassopeptide, LPA, CDPS, RRE-containing proteins, redox-cofactors, thiopeptides, lanthipeptide class I, phenazine, hglE-KS, lanthipeptide class II, indoles, ladderane, linaridin, arylpolyene, lanthipeptide class IV, blactam, nucleosides, oligosaccharides, lanthipeptide class V, transAT-PKS, guanidinotides, siderophores, aminoglycoside / aminocyclitol (“amglyccycl”), thioamide-NRP, prodigiosin, furan, cyanobactin, phosphoglycolipids, phosphonate, transAT-PKS-like, bottromycin, thioamitides.

[0072] According to a particular embodiment, the plurality of biosynthetic gene clusters (BGC) is for the synthesis of non-ribosomal peptides (NRPS), polyketides (PKS), terpenes, aminoglycosides, bacteriocins, lassopeptides, lantipeptides, or combinations thereof.

[0073] According to a particular embodiment, the plurality of biosynthetic gene clusters (BGC) is for the synthesis of non-ribosomal peptides (NRPS), polyketides (PKS), terpenes, aminoglycosides, bacteriocins, lassopeptides, lantipeptides, nucleosides, phosphonates or combinations thereof.

[0074] According to a particular embodiment, the at least one gene cluster of interest is a BGC for the synthesis of NRPS or PKS.

[0075] According to a particular embodiment, the at least one gene cluster of interest is a BGC for the synthesis of combinations of NRPS and PKS.

[0076] According to one embodiment, the at least one gene cluster of interest (e.g. a biosynthetic gene cluster) has a size of less than 300 kbp; for example less than 200 kbp,less than 100 kbp, less than 20 kbp or even less than 10 kbp.

[0077] According to a particular embodiment, the at least one gene cluster of interest (e.g. a biosynthetic gene cluster) has a size of at least 10 kbp; for example at least 20 kbp, for example at least 100 kbp.

[0078] According to a particular embodiment, the at least one gene cluster of interest has a size of at least 20 kbp, for example ranging from 20 kbp to 300 kbp, for example ranging from 20 kbp to 200 kbp.

[0079] According to a particular embodiment, the at least one gene cluster of interest has a size of at least 20 kbp; which may thus include 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300 kbp.

[0080] According to a particular embodiment, the clone library is characterized in that it consists of, or comprises, a biosynthetic gene cluster.

[0081] According to a particular embodiment, the clone library is characterized in that the size of the high-molecular weight nucleic acid is of at least 10 kbp; for example at least 20 kbp, for example at least 100 kbp. According to a particular embodiment, the clone library is characterized in that the size of the high-molecular weight nucleic acid is of at least 20 kbp; which may thus include 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300 kbp.

[0082] According to a particular embodiment, the clone library is characterized in that the size of the high-molecular weight nucleic acid is of at least 10 kbp; the clone library being further characterized in the clone library consists of, or comprises, a biosynthetic gene cluster.

[0083] According to a particular embodiment, the clone library is characterized in that the size of the high-molecular weight nucleic acid is of at least 20 kbp, for example ranging from 20 kbp to 300 kbp for example ranging from 20 kbp to 200 kbp.

[0084] According to a particular embodiment, the clone library is characterized in that the size of the high-molecular weight nucleic acid is of at least 20 kbp; which may thusinclude 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300 kbp

[0085] According to a particular embodiment, the clone library is characterized in that it consists of a cloning vector library selected from a group consisting of: cosmid, fosmid, phage, bacterial artificial chromosome (BAC), Pl-derived artificial chromosome (PAC), yeast artificial chromosome (YAC), fungal artificial chromosome (FAC).

[0086] According to a particular embodiment, the clone library comprises or consists of a shuttle vector; in particular a shuttle vector which is suitable for replication in a plurality of cells belonging to distinct genus or kingdoms, more particularly one belonging to the genus Streptomyces and the other to the genus Escherichia', preferably a shuttle vector which is suitable for replication in E. coli and Streptomyces spp.

[0087] According to a particular embodiment, the shuttle vector is suitable for replication in a plurality of prokaryotic cells belonging to distinct genus, and genomic integration, for example genomic integration in Streptomyces spp.

[0088] According to a particular embodiment, the clone library is provided in the form of a prokaryotic cell; for example gram-positive or gram-negative bacteria, for example belonging to the genus Escherichia, preferably Escherichia coli.

[0089] According to a particular embodiment, the method is for identifying a plurality of gene clusters of interest in a cell or library thereof; wherein the method further includes a step of recovering the high-molecular weight nucleic acid(s) from the library in which the said cluster(s) of interest(s) is / are detected.

[0090] It is noteworthy that establishing completeness of the gene cluster of interest (e.g. the BGC) on a given clone after nucleic acid fractionation and determination of nucleic acid sequences (e.g. through shotgun sequencing) is greatly facilitated if the gene cluster of interest is assumed to be statistically present on only one clone within a pool.

[0091] Indeed, if the gene cluster of interest (e.g. the BGC) was present in more than one clone in a pool, reads originating from overlapping clones aligning to the full sequence of the gene cluster might be observed even when none of them individuallycontain the complete sequence.

[0092] Consequently, it is very advantageous during the first partition of the library to constitute a number of pools greater than the expected number of clones containing partial or complete gene clusters of interests. This number can be estimated statistically by taking into account the ratio between the number of nucleotides captured in the clone library and the number of nucleotide in the genomes of interest from which the library was created.

[0093] Also, while localizing clones by triangulation necessarily involves resources and efforts allocated to analyze multiple dimensions and partitions (e.g. multiple partitions of pools), the proposed sequential pooling strategy minimizes them by localizing only a selection of clones which contain complete gene clusters of interest.

[0094] Indeed, one further advantage of partitioning the clone library into a first plurality of clone pools, and sequencing (e.g. shotgun sequencing) of this first partition, are three-fold: to obtain definitive information regarding the nature, the completeness of the clusters, and gain partial information regarding their location in the clone library.

[0095] From this step onward, the allocation of further resources to obtain the exact clone location of gene clusters of interest can be deliberately decided based on the level of interest of each cluster (which can be assigned priority scores) as well as on the economy of scale attached to the simultaneous detection of other gene clusters of interest in the same pools.

[0096] Typically, deciding to localize too few gene clusters will lead to higher costs percluster as little economies of scale can be harnessed and sequencing of large number of pools, or pools with large numbers of clones is thus required.

[0097] Conversely, attempting to localize too many or all of the gene clusters of interest contained in the clone library will likely face diminishing returns and also lead to a high costs per-cluster.

[0098] The possibility to allocate or not resources to localize given clusters after the initial sequencing step, offers the option to aim for various optima, such as minimizing the cost per-cluster or the total number of rounds of pooling during the procedure.

[0099] A read coverage (or “sequencing depth”) analysis is generally conducted on the resulting alignments and reveals which references can be assigned enough reads regularly distributed over the full reference sequence.

[0100] Advantageously, the method according to the invention is thus suitable for localizing gene clusters even under low sequencing depth, or very low sequencing depth, while remaining applicable toward a large set of clone libraries and a large selection of gene clusters; in particular a large selection of BGCs.

[0101] Advantageously, the method as described herein, may be thus further characterized in that the step of determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone pools and / or the clone stacks, thereby providing a plurality of sequence reads, is achieved with varying levels of sequencing depth; for example with low sequencing depth, while retaining its ability to efficiently identify or recover gene cluster(s) of interest.

[0102] According to a particular embodiment, the method is characterized in that the sequencing depth corresponding to the plurality of sequence reads is of less than 20x, for example less than 19x, 18x, 17x, 16x, 15x, 14x, 13x, 12x, l lx, lOx, 9x, 8x, 7x, 6x, 5x, 4x, 3x, 2x, or less than lx. According to a particular embodiment, the method is characterized in that the sequencing depth corresponding to the plurality of sequence reads is of less than 20x, for example less than 19x, 18x, 17x, 16x, 15x, 14x, 13x, 12x, 1 lx, lOx, 9x, 8x, 7x, 6x, 5x, 4x, 3x, 2x, or less than lx; and the method is for identifying at least one gene cluster of at least 10 kbp; for example of at least 20 kbp, or more. According to a particular embodiment, the method is characterized in that the sequencing depth corresponding to the plurality of sequence reads is of less than 20x, for example less than 19x, 18x, 17x, 16x, 15x, 14x, 13x, 12x, l lx, lOx, 9x, 8x, 7x, 6x, 5x, 4x, 3x, 2x, or less than lx; and the method is for identifying at least one biosynthetic gene cluster (BGC).

[0103] According to exemplified embodiments, when applied to the identification of gene clusters (e.g. BGCs), it is shown that the method is capable of recovering said clusters by determining a plurality of sequence reads with low to very low sequencingdepth; in particular with a sequencing depth of lOx or less than lOx; for example of 2x or less than 2X, for example of 0. lx, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, 1.Ox, 1. lx, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x.

[0104] According to some embodiments, the sequencing depth is equal or superior to O.Olx, 0.02x, 0.03x, 0.04x, 0.05x, 0.06x, 0.07x, 0.08x, 0.09x, O.lx, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, l.Ox, l.lx, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, or 2. Ox.

[0105] According to some embodiments, the sequencing depth is ranging from O.lx, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, l.Ox, l.lx, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, or 2. Ox, to lOx, 20x or 30x. According to some embodiments, the sequencing depth is ranging from O.lx, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, l.Ox, l.lx, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, or 2. Ox, to lOx or 20x. According to some embodiments, the sequencing depth is ranging from O.lx, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, l.Ox, l.lx, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, or 2.0x, to lOx.

[0106] According to a particular embodiment, the method is characterized in that the completeness of sequencing, as a percentage of the reference sequence (e.g. the gene cluster or a portion thereof) covered by nucleic acid sequence reads, is above a predetermined threshold; for example above at least 10%.

[0107] According to a particular embodiment, the method is characterized in that the completeness of sequencing, as a percentage of the reference sequence (e.g. the gene cluster or a portion thereof) covered by 1-kb nucleic acid sequence portions, is above a pre-determined threshold; for example above at least 10%.

[0108] According to a particular embodiment, the clone pools forming the clone stacks, statistically comprise at least one occurrence of the gene cluster of interest.

[0109] According to a particular embodiment, each clone pool from the plurality of clone pools statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest.

[0110] According to a particular embodiment, each clone stack from the plurality of clone stacks statistically comprises less than two, in particular at most one, occurrence ofsaid at least one gene cluster of interest.

[0111] The above-mentioned embodiments may advantageously be combined altogether. Hence, according to a particular embodiment, the clone pools forming the clone stacks, statistically comprise at least one occurrence of the gene cluster of interest; and / or each clone pool from the plurality of clone pools statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest; and / or each clone stack from the plurality of clone stacks statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest.

[0112] According to a particular embodiment, the clone pools forming the clone stacks, statistically comprise at least one occurrence of the gene cluster of interest; and each clone pool from the plurality of clone pools statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest; and each clone stack from the plurality of clone stacks statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest.

[0113] Advantageously, the step of fractionating the high-molecular weight nucleic acid from the clone pools may thus give hindsight on how the partitioning of the second plurality of separate portions could be achieved.

[0114] In view of the above, it is foreseen that the step of partitioning the clone library into a second plurality of separate partitions, comprising a clone from each clone pool, is guided by the nucleic acid sequences which were previously determined from all or part of the fractionated high-molecular weight nucleic acid in the clone pools.

[0115] In particular, the clone pools forming the clone stacks, could be re-arranged in that each clone pool from the clone stack statistically comprise at least one occurrence of the gene cluster of interest, based on the first determination of nucleic acid sequences.

[0116] Hence, according to a particular embodiment, the invention relates to method for identifying at least one gene cluster of interest in cell(s) or a clone library thereof, in particular a plurality of gene clusters of interest in cell(s), preferably a plurality of BGCs in eukaryotic cell(s) or a clone library thereof, comprising steps of:a) providing a clone library comprising a plurality of clones with high- molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; bl) partitioning the clone library into a first plurality of separate partitions comprising a plurality of clones, thereby forming a plurality of clone pools; b2) fractionating the high-molecular weight nucleic acid from the clone pools, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone pools, thereby determining the plurality of clone pools comprising at least one occurrence of said at least one gene cluster of interest; cl) partitioning the clone library into a second plurality of separate partitions, comprising a clone from each clone pool, thereby forming a plurality of clone stacks; c2) fractionating the high-molecular weight nucleic acid from the clone stacks, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone stacks; d) determining the clones comprising the gene cluster(s) of interest in the library, based on the nucleic acid sequences from the clone stacks, thereby identifying the at least one gene cluster of interest; optionally wherein the clone pools forming the clone stacks, statistically comprise at least one occurrence of the gene cluster of interest; optionally wherein each clone pool from the plurality of clone pools statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest; optionally wherein each clone stack from the plurality of clone stacks statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest.

[0117] According to a preferred embodiment, the method comprises steps:a) providing a clone library comprising a plurality of clones with high- molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; b) partitioning the clone library into a first plurality of separate partitions comprising a plurality of clones, thereby forming a plurality of clone pools, fractionating the high-molecular weight nucleic acid from the clone pools, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone pools, thereby determining the plurality of clone pools comprising at least one occurrence of said at least one gene cluster of interest; c) partitioning the clone library into a second plurality of separate partitions, comprising a clone from each clone pool, thereby forming a plurality of clone stacks, fractionating the high-molecular weight nucleic acid from the clone stacks, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone stacks; d) determining the clones comprising the gene cluster(s) of interest in the library, based on the nucleic acid sequences from the clone stacks, thereby identifying the at least one gene cluster of interest; wherein the clone pools forming the clone stacks, statistically comprise at least one occurrence of the gene cluster of interest. optionally wherein each clone pool from the plurality of clone pools statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest; optionally wherein each clone stack from the plurality of clone stacks statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest.

[0118] According to a particular embodiment, the method includes a plurality of steps of partitioning the clone library; thereby forming more than two partitions, for example more than two clone pools and more than two clone stacks.

[0119] According to a particular embodiment, the step of providing a clone library may comprise the sub-steps of:(i) providing a cell or plurality of cells susceptible to encode or comprise gene clusters of interest (e.g. BGCs and / or gene clusters having a size of at least 10 kbp);(ii) isolating nucleic acids from the cell, or plurality thereof;(iii) fractionating the nucleic acids into high-molecular nucleic acids susceptible to encode or comprise gene clusters of interest; and(iv) preparing a plurality of clones with the high-molecular weight nucleic acid, thereby preparing the clone library.

[0120] According to some embodiments, the method according to the invention comprises a step of aligning the sequence reads with all or part of a reference sequence.

[0121] According to a particular embodiment, the step of determining the sequence of high-molecular weight nucleic acids in the clone pool and the clone stack may comprise the substeps of:(i) determining the nucleic acid sequence of all or part of the corresponding fractionated high-molecular weight nucleic acids, thereby providing a plurality of sequence reads;(ii) aligning the sequence reads with all or part of a reference sequence, thereby determining the sequence of the high-molecular weight nucleic acids.

[0122] According to a particular embodiment, the sequencing depth corresponding to said plurality of sequence reads is of less than 30x. According to a particular embodiment, the sequencing depth corresponding to said plurality of sequence reads is of less than 20x; for example less than lx.

[0123] According to a particular embodiment, the invention thus relates to a method for identifying at least one gene cluster of interest (e.g. BGC), preferably a plurality of geneclusters, comprising steps of: a) providing a clone library comprising a plurality of clones with high-molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; b) partitioning the clone library into at least a first and second plurality of separate partitions: the first plurality of separate partitions comprising a plurality of clones, the plurality of separate partitions thereby forming a plurality of clone pools; the second plurality of separate partitions comprising a plurality of clones from clone pools, thereby forming a plurality of clone stacks; wherein the high-molecular weight nucleic acids from the clone pools and the clone stacks are fractionated; c) determining the nucleic acid sequence of all or part of the corresponding fractionated high-molecular weight nucleic acids, thereby providing a plurality of sequence reads, the sequencing depth corresponding to said plurality of sequence reads being of less than 30x, for example less than 20x; d) aligning the sequence reads with all or part of a reference sequence, thereby determining the sequence of the high-molecular weight nucleic acids, thereby identifying the at least one gene cluster of interest.

[0124] The step of determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone pools and the clone stacks, provides a plurality of sequence reads; including:- a plurality of sequence reads from fractionated high-molecular weight nucleic acids in the first partition (e.g. clone pools); and- a plurality of sequence reads from fractionated high-molecular weight nucleicacids in the second partition (e.g. clone stacks).

[0125] According to some embodiments, both sets of sequence reads (i.e. from the first and second partition) are provided at a sequencing depth of less than 30x, for example less than 20x; the sequencing depth of each plurality being the same or different. For example, the sequencing depth corresponding to the first partition may be higher than the sequencing depth corresponding to the second partition. Alternatively, the sequencing depth corresponding to the first partition may be lower than the sequencing depth corresponding to the second partition.

[0126] According to some embodiments, the sequencing depth corresponding to the plurality of reads of the first partition may be of less than 30x, and the sequencing depth corresponding to the plurality of reads of the second partition may be of less than 30x. According to some embodiments, the sequencing depth corresponding to the plurality of reads of the first partition may be of less than 20x, and the sequencing depth corresponding to the plurality of reads of the second partition may be of less than 20x. According to some embodiments, the sequencing depth corresponding to the plurality of reads of the first partition may be of less than 30x, and the sequencing depth corresponding to the plurality of reads of the second partition may be of less than 20x. According to some embodiments, the sequencing depth corresponding to the plurality of reads of the first partition may be of less than 20x, and the sequencing depth corresponding to the plurality of reads of the second partition may be of less than 30x. According to some embodiments, the sequencing depth corresponding to the plurality of reads of the first partition may be of less than lOx, and the sequencing depth corresponding to the plurality of reads of the second partition may be of less than lOx.

[0127] According to some embodiments, more than one set of sequence reads may be obtained from a given (e.g. first or second) partition. This may be particularly advantageous when one set of sequence reads is below a reference threshold of completeness. The reference threshold of completeness may be modulated, depending on the number of sets of sequence reads for a given partition, and / or for a given set of partitions. For example, the reference threshold may be at or below a pre-determined percentage of X-kb nucleic acid sequence reads covering the entire length of saidpartition / reference sequence; wherein X may be 20-kb; 15-kb, 10-kb, 9-kb, 8-kb, 7-kb, 6- kb, 5-kb, 4-kb, 3-kb, 2-kb, 1-kb, 900 bp, 800 bp, 700 bp, 600 bp, 500b, or less than 500 bp (e.g. below a pre-determined percentage of 1-kb nucleic acid sequence reads covering the entire length of said partition / reference sequence).

[0128] According to some embodiments, the step of determining the sequence of high- molecular weight nucleic acids in the clone pool and the clone stack may comprise the sub steps of:(i) determining the nucleic acid sequence of part of the corresponding fractionated high-molecular weight nucleic acids, thereby providing a first plurality of sequence reads;(ii) aligning the first plurality of sequence reads with all or part of a reference sequence; and(iii) optionally, when the first plurality of sequence reads is at, or below, a reference threshold of completeness, determining the nucleic acid sequence of a second part of the corresponding fractionated high-molecular weight nucleic acids, thereby providing a second plurality of sequence reads.

[0129] According to said embodiments, the plurality of sequence reads may comprise:- a first plurality of sequence reads from fractionated high-molecular weight nucleic acids in the first partition (e.g. clone pools); and- at least a second plurality of sequence reads from fractionated high-molecular weight nucleic acids in the first partition (e.g. clone stacks).

[0130] According to said embodiments, the plurality of sequence reads may comprise:- a first plurality of sequence reads from fractionated high-molecular weight nucleic acids in the second partition (e.g. clone pools); and- at least a second plurality of sequence reads from fractionated high-molecular weight nucleic acids in the second partition (e.g. clone stacks).

[0131] According to some embodiments, the number of sequence reads is sufficient for aligning at least one read to the reference sequence; for example for aligning for example at least about 10 reads to the reference sequence; for example at least about 100 reads to the reference sequence; for example at least 1000 about reads to the reference sequence.

[0132] According to some embodiments, the method according to the invention comprises a step of assembling all or part of the previously aligned sequence reads. According to some other embodiments, the method according to the invention does not comprise a step of assembling the previously aligned sequence reads.

[0133] According to some embodiments the sequencing depth is of less than lOOOx; for example less than 500x; for example less than lOOx; for example less than 20x; for example less than lOx; for example less than lx.

[0134] According to some embodiments the sequencing depth is of less than lx for at least one step of determining the nucleic acid sequence of fractionated high-molecular weight nucleic acids; for example for determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acids in the first partition.

[0135] Hence, according to some embodiments, the step of determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the first partition and the second partition, provides a plurality of sequence reads; including:- a plurality of sequence reads from fractionated high-molecular weight nucleic acids in the first partition (e.g. clone pools), wherein the sequencing depth is of less than lOOOx; for example less than lOOx; for example less than lOx; for example less than lx; and- a plurality of sequence reads from fractionated high-molecular weight nucleic acids in the second partition (e.g. clone stacks), wherein the sequencing depth is of less than lOOOx; for example less than lOOx; for example less than lOx; for example less than lx.

[0136] According to some embodiments, each clone from the plurality of clones is attributed coordinates in the corresponding clone pool.

[0137] According to some embodiments, each clone from the plurality of clones is attributed coordinates in the corresponding clone pool and clone stack.

[0138] According to some embodiments, the method comprises a step of determining the coordinates of a clone comprising said at least one gene cluster of interest in the library.

[0139] According to some embodiments, the method comprises a step of localizing the coordinates of a clone comprising said at least one gene cluster of interest in the library

[0140] According to a particular embodiment, the method is a computer- implemented method. According to said embodiment, one or more steps are computer-implemented; for example steps of determining the clones comprising the gene cluster(s) of interest in the library, based on the nucleic acid sequences from the clone pools and / or clone stacks.

[0141] According to a particular embodiment, the method comprises a step of recovering, preferably isolating, the clone comprising the at least one gene cluster of interest from the library.

[0142] According to a particular embodiment, the method comprises a step of replicating the recovered clone, and / or cultivating a cell comprising the recovered clone in a suitable medium.BRIEF DESCRIPTION OF THE DRAWINGS

[0143] Figure 1 is a schema showing the general concept of localization of a BGC of interest from a clone library through rounds of partitioning followed by shotgun sequencing.

[0144] Figure 2 is a schema showing a multiplexed divide and conquer strategy to efficiently localize multiple BGCs of interest in a clone library. After sequencing of the pools that constitute the first partition, only fractions of the library predicted to contain clones harboring a selection of complete BGCs of interest are retained to form newpartitions (a second partition, and optionally one or more further partitions). The selection of the clones which will form the subsequent partitions can be optimized computationally to efficiently triangulate multiple BGCs of interest in parallel in a cost and / or time effective manner.

[0145] Figure 3 shows the coverage of a reference BGC from Example 1 using sequencing reads from a single sample. Different lines represent varying sequencing depths, achieved by random down- sampling at the mapping step. A coverage of 2X is sufficient to determine that the variant of the target reference BGC is complete in this sample. In X-axis, the position across the target reference BGC is expressed in bases, from 0 to 100000. In Y-axis, the coverage is expressed as the mean sequencing depth over 1-kb bins.

[0146] Figure 4 displays a multipanel histogram illustrating the observed coverage (percentage of 1-kb bins covered) for targets from Example 1. Each panel represents a different sequencing depth, achieved through random down-sampling during the mapping step, with 4A corresponding to a 20X sequencing depth (left panel) and 10X sequencing depth (right panel), and 4B corresponding to a 2X sequencing depth (left panel) and 0.2X sequencing depth (right panel). In X-axis, the number of BGCs. In Y-axis, the observed coverage is provided, expressed as the percent of 1-kb bins across the BGCs having at least one read. Targets predicted to be complete using the full dataset are highlighted in light grey. Notably, even at 2X coverage, a threshold can be established to distinguish between complete and incomplete targets.

[0147] Figure 5 is a graph showing the number and diversity of localized NRPS from example 1 in two sequencing steps. Each node is a reference BGC of type NRPS predicted complete in at least one plate of our clone library according to the first sequencing step. Edges link BGCs that are structurally similar according to Big-SCAPE (default parameters). Dark- filled nodes are reference BGCs localized at the well-level after the second sequencing step. Dark-bordered nodes are those deemed interesting for their novelty. All complete NRPS are represented, including those dimmed less interesting (grey border). The sequential nature of this workflow allows to aim for specific BGCs and thus to aim for the most interesting and diverse set of BGCs while the localization isongoing. Importantly, interesting BGCs currently not completely localized (shown in white with dark border) could be localized at the price of additional sequencing steps.EXAMPLES

[0148] The present invention is further illustrated by the following examples.Example 1:Materials and Methods

[0149] Streptomyces strain collection

[0150] Around 1g of top soil were collected in 96 locations in the Ile-de-France region in France, in compliance with the local regiementation related to the Nagoya protocol. Microbial colonies with morphological features associated with Streptomyces were isolated from each sample by using a standard isolation procedure. PCR was performed on each strain with barcoded universal 16S primers (27F: AGAGTTTGATCCTGGCTCAG (SEQ ID N°l), 1492R:GGTTACCTTGTTACGACTT (SEQ ID N°2)) and the amplicons were sequenced on a MinlON sequencer (Oxford Nanopore Technologies) with a R10.4.1 flow cell. Reads associated with each strain were aligned by Blastn to a database composed of the genomes present on RefSeq and the strains with a 16S nucleotide identity above 98% to a genome annotated as Streptomyces were selected for long term storage in the form of glycerol stocks.

[0151] High molecular weight DNA extraction and BAC libraries construction

[0152] Spore stocks of 140 Streptomyces species were grown individually in TSB medium. Mycelia were harvested just after reaching stationary phase. Mycelia were centrifuged, and 20 pF of each pellet was pooled together. The pooled mycelia were washed in 2 volumes of buffer (200 mM NaCl; 10 mM TrisHCl pH8; 100 mM EDTA) and centrifuged. The pooled mycelia was incorporated in agarose plugs and treated with lysozyme and proteinase-K according to the instructions of the CHEF Bacterial Genomic DNA Plug Kit (Bio-Rad). BAC library construction was performed by a service provider(CNRGV-INRAE, Toulouse) according to standard methods. Briefly, the high-molecular weight DNA present in the plugs was partially digested with BamHI and ligated into the pAGIBAC vector modified to include fragment of pOJ436 containing the Streptomyces shuttle elements (OriT, Apramycin resistance, 0C31 integrase gene). The ligation product was transformed into E. coli DH10B and the resulting transformants were picked into 384-well plates containing LB and chloramphenicol.

[0153] Initial partition and shotgun sequencing

[0154] Each 384-wells plate of the BAC library was replicated in new 384-wells plates with fresh media to grow sufficient biomass for each clone. After 24h at 37°C, for each plate, all wells were combined in a single pool. The 384 BAC present in each pool were extracted together using a standard BAC extraction protocol. The concentration of the resulting purified DNA was normalized prior to their sequencing by a service provider. Each extracted pool was sequenced on a Novaseq sequencer (Illumina), using 2*150bp paired-end reads, to generate around 1 Gigabase per pool. Quality control of the reads was done using fastp for low quality reads removal and trimming. Reads were aligned with bbmap (BBtools may be cited using the primary website: BBMap - Bushnell B. - sourceforge.net / projects / bbmap / ) on the E. coli DH10B genome as well as the sequence of the BAC vector and only non-aligned reads were analyzed further.

[0155] Presence and completeness prediction

[0156] Reads were aligned with bbmap to the database of reference BGCs using conservative parameters to generate per-nucleotide coverage and Ikb binned coverage (covbinsize=1000, maxsites2=5000, delcoverage=f). An arbitrary per-nucleotide coverage of 50% was used to create a whitelist of reference BGCs that may partially be in the library. The binned coverage of whitelisted references was analyzed more in detail to predict which BGCs were likely complete (or truncated, or absent) and in which pools. Specifically, we considered that references having 98% or more of their mean binned coverage above 50% were likely receiving reads from a closely related BGC, and were predicted complete. Last bin was ignored from the analysis as it may be shorter than Ikb and its coverage may bias the statistics. Similarly, we considered that references wereincomplete and susceptible to be difficult to distinguish from a complete reference if they had between 66% to 98% of their bins covered (i.e. per-bin mean coverage above 50%). Smaller truncated regions were not considered to be problematic for localization.

[0157] Database of reference BGCs

[0158] A database of reference BGCs was built by analysing all the actinomycetota genomes present in RefSeq with the antiSMASH software (Blin et al. Nucleic Acids Research, 51, Wl, 2023). Regions detected by antiSMASH were used as a proxy to define BGCs and their associated boundaries. BGCs were deduplicated with dedupe. sh (from the BBTools suite) using conservative parameters (minidentity=99 e=1000 ac=f). The resulting FASTA contained more than 300000 loosely deduplicated antiSMASH regions that were considered as “reference BGCs sequences” for the purpose of the experiment. A bbmap index was built using this file.

[0159] BGC interest assessment and priorisation

[0160] The antiSMASH output of reference BGCs predicted complete were manually reviewed to evaluate the relative interest of the cluster. Each BGC was attributed an integer score between 0 and 5. References likely to produce a known natural product reveived a score of 0 and were removed from the analysis. All scored references above or equal to 1 were considered valid targets. Higher scores aimed at favoring the localization of some BGCs over others if resources allocated to localization were to be constrained.

[0161] Pool design for the subsequent shotgun sequencing rounds

[0162] A linear programming algorithm was used to find combinations of plates that have strictly one predicted complete occurrence of a maximum of high-score references without any incomplete one. Each reference of score n was arbitrarily considered worth as much as 3 references of the score n-1. The algorithm was run iteratively, removing references found at a previous iteration, to find combinations of plates allowing the localization of all the references.

[0163] Subsequent shotgun sequencing rounds and localization

[0164] This step was performed identically to the first shotgun sequencing step except as to how the pools were constituted. Briefly, once grown, cells were transferred well-to- well from selected plates to a receiving 384-wells plate (all Al to Al, etc.). Then, sequencing samples were formed by pooling all wells for each row (resp. columns) resulting in 16 (resp. 24) pools. Sequencing results were analyzed through the same workflow as described before to identify which pools contained which references at which coverage values. For each target BGC, a row and column were predicted based on the highest per-nucleotide coverage in each corresponding pool.

[0165] Long-read sequencing and confirmation of the localized BGCs

[0166] Additional confirmation of the localization was conducted by randomly selecting 10 clones predicted to harbor one of the non-ambiguously localized BGCs. Sequencing of the corresponding BACs was done on a MinlON sequencer (Oxford Nanopore Technologies) and confirmed the presence of sequences closely related to the expected reference sequences in all cases.Results

[0167] In order to prove the concept of the method presented herein, multiple biosynthetic gene clusters (BGCs) originating from bacteria of the genus Streptomyces were successfully isolated on cloning vectors maintained in E. coli. Streptomyces are advantageous candidates, as they include many BGCs encoding for yet unknown natural products. Soil samples were collected and Streptomyces strains were isolated with standard isolation techniques. The 16S locus of these strains were amplified by PCR and sequenced to ensure that they belong to Streptomyces. 140 Streptomyces strains with distinct 16S sequences were retained for the demonstration.

[0168] Selected Streptomyces strains were grown and harvested to build a bacterial artificial chromosome (BAC) library using a standard technique. The first step was to grow each strain individually to build-up a sufficient DNA amount. Then, the genomic DNA of all strains was extracted and pooled to form a high-molecular weight (BMW)library. These HMW DNA fragments were then electroporated into E. coli and maintained as bacterial artificial chromosomes (BAC). E. coli transformants were finally plated and picked individually into 384- well plates to form a BAC library of Streptomyces inserts in which each clone was stored in an individual well. Consequently, each well harbors a clone that is expected to harbor a BAC containing a random DNA insert from one of the selected Streptomyces strains. This step led to the creation of a BAC library composed of 18432 clones. The mean size of these inserts was estimated around lOOkb based on the analysis of a subset of 10 clones chosen at random.

[0169] Only a small proportion of clones were expected to be of interest: the subset of those harboring a complete BGC that was never reported to be associated with a known secondary metabolite, i.e. not present on the MIBiG database (Terlouw et al. Nucleic Acids Research, Vol. 51, Issue DI, 2023).

[0170] The purpose is to localize interesting clones without wasting resources on others. A sequential, multi-round pooling strategy is provided, to analyze intermediate partitions and guide the design of subsequent partitions in order to localize as many BGCs of interest as possible, while minimizing the time or the resources required

[0171] For each 384-well plate of the library, all the clones of each plate were pooled and shallowly sequenced. Reads were aligned to a database of dereplicated known BGCs sequences preferences '). As captured BGCs are expected to have closely related but different sequences to the references, alignment parameters were set to allow for errors between raw reads and reference sequences.

[0172] A read coverage analysis was conducted on the resulting alignments, to predict which references were likely complete in which pools. As we are only interested in the completeness of the BGC we only have to check that reads are roughly distributed over the full length of the reference sequence. For each sequencing round, reads from each pool are mapped against the reference BGC sequence. If their completeness passes a predetermined threshold, the detailed coverage is examined to determine if the core BGC is indeed complete. Advantageously, when a target is deemed interesting and complete, one or more sequencing steps may be introduced / provided, to seek more localizationinformation. Also advantageously, distinct pools may be designed for each sequencing step, in order to gain further information on a plurality of target nucleic acid sequences / targets .

[0173] This use of shotgun sequencing technology as a screening tool only requires a small sequencing depth (figure 3 and 4A, B), which in turn allows to pool much more samples. Advantageously, the assembly step, that is notoriously difficult for sequences having repeated domains like BGCs, may be by-passed.

[0174] The list of reference BGCs expected to be complete at the first sequencing step is used to prioritize a selection of BGCs. A large number of BGCs in all kinds of BGCs families were detected, including those BGCs encoding genes for the synthesis of one or more of the following compounds and / or polypeptides: terpenes, NRPS, RiPP-like, T1 PKS, NRPS-like, melanin, butyrolactone, T3 PKS, T2 PKS, ectoine, NAPAA, betalactone, lanthipeptide class III, PKS-like, lassopeptide, LPA, CDPS, RRE-containing proteins, redox-cofactors, thiopeptides, lanthipeptide class I, phenazine, hglE-KS, lanthipeptide class II, indoles, ladderane, linaridin, arylpolyene, lanthipeptide class IV, blactam, nucleosides, oligosaccharides, lanthipeptide class V, transAT-PKS, guanidinotides, siderophores, aminoglyco side / aminocyclitol (“amglyccycl”), thioamide- NRP, prodigiosin, furan, cyanobactin, phosphoglycolipids, phosphonate, transAT-PKS- like, bottromycin, thioamitides. Some BGCs belong to multiple classes. The BGC reference sequence’s features were used to specifically target BGCs for their novelty (figure 5). 161 BGCs were considered valuables, novel, and thus worth retrieving. Some of these BGCs occurred only once in the whole library, while others had complete or incomplete copies across multiple plates. At this point of the analysis the platelocalization of each of those BGCs is known, but not yet their well-localization.

[0175] First, the PCR and amplicon sequencing strategies were replaced altogether by a shotgun sequencing strategy. Samples were shotgun sequenced with a shallow coverage and reads were aligned to a database of dereplicated known BGCs sequences (“references”). As the captured BGCs are expected to have closely related but different sequences to the references, alignment parameters are set to allow for errors between raw reads and reference sequences. A read coverage analysis is conducted on the resultingalignments and reveals which references can be assigned enough reads regularly distributed over the full reference sequence, in order to determine if a reference (or a close variant) is likely captured in one or several clones.

[0176] Secondly, a sequential, multi-round pooling strategy allows to analyze intermediate partitions and guide the design of subsequent partitions in order to localize as many BGCs of interest as possible while minimizing the time or the resources required. One implementation of this concept can be seen in this example. For each 384-wells plates of the library, all the clones of the plate were pooled and sequenced as a single sample. The analysis of this first dimension using read coverage data revealed which reference BGC variants were expected to be in each pools, and hence in the corresponding plates.

[0177] In particular, the analysis detected a large number of BGCs, 161 of which were considered novel and interesting. Some of these BGCs occurred only once in the whole library, while others had complete or incomplete copies accross multiple plates.

[0178] New sets of pools were designed to maximize the number of localized BGCs in a minimum number of sequencing steps. To do so, subsets of various numbers of plates were determined to maximize the number of complete BGCs of interest that would be in strictly one copy accross each subset.

[0179] As a demonstration, a subset of 24 plates was selected after it was predicted to contain 63 complete BGCs of interest, each present complete on exactly 1 BAC across the whole subset. While several subsets of plates could be constituted for subsequent analysis to recover more BGCs, this example focused only on this first subset of 24 plates. In order to obtain the exact location of each BGC of interest, pools were constituted by aggregating the clones belonging to the same row accross the 24 plates (16 rows in total, thus yielding a total of 16 pools), as well as pools aggregating the clones belonging to the same column accross the 24 plates (24 columns in total, thus yielding a total of 24 pools).

[0180] Because the BGCs of interest are expected to be present in only one clone in exactly one of the 24 plates, they are expected to be present in exactly one of the 16 pools corresponding to rows and one of the 24 pools corresponding to columns, thus allowing the simultaneous triangulation of the exact location of the clones containing each of theBGCs of interest.

[0181] Analysis of the read coverage data from the plates, rows and columns dimensions, demonstrated the unambiguous localization of 55 complete BGCs of interest. These BGCs were later individually transferred into a heterologous host to assess for the production of their associated natural products using standard analytical chemistry methods.

[0182] This experimental demonstration showcase a key feature of this approach: the presence and localization of a large number of BGCs, potentially belonging to a large number of different classes, can be assessed in parallel, reducing significantly the per- BGC cost of the procedure. In addition, the example also shows that shotgun sequencing can be used even with shallow sequencing depth which requires much less resources than for the full assembly into contigs of the exact sequences present in the library.Example 2:

[0183] This example focuses on one reference BGC (GCF_000716685.1.contig0007.region001) that is a T2PKS for which a related BGC was localized following a similar procedure as the one detailed in example 1.

[0184] In comparison to the procedure described in example 1, here we used three sequencing steps (instead of two), and we worked at a lower sequencing depth.

[0185] For the first sequencing step (PE), each sample is equivalent to the wells of 8 pooled 384-well plates. Assuming, BACs with lOOkb insert, lOkb backbone and no contamination whatsoever (including from the hosting strain), this gives a theoretical sequencing depth of 6X for 2Gb of data (2e6 / (8*384*110).

[0186] For the second sequencing step (MS), each sample is equivalent to at most 21 pooled 384-well plates. The list of plates in each pool is chosen such that the combination of PE and MS data gives the plate-localization.

[0187] Finally, a last partition (S) of plates has their rows and columns sequenced in distinct samples like we previously described to resolve the well-localization.

[0188] Relevant metrics, including binned coverage at each sequencing step are provided. Specifically, for the partition PL, the BGC is seen complete in three pools (internal reference numbers: PL-GEN- 12-8 and PL-GEN- 12-7) of the first sequencing step. The BGC is then seen as complete in pool MPST-240823-S0M6 of the second sequencing step. Knowing which plates constitute these pools, it can be inferred that only plate 55 of library 12 can explain the observed signal (data not shown). This plate was included in the last sequencing step to finalize the well-localization of the BGC (among others). Using pooled rows and columns (partition S), the full localization of the BGC in well K17 of plate 55 of library 12 is recovered. Note that all three sequencing steps used a very low sequencing depth that would be improper for de novo assembly. In this particular example, the mean read coverage is inferior to 3X. Despite this fact, the target localization could still be resolved.

[0189] The localization of this BGC was later validated in a long-read sequencing quality check experiment. Read mapping showed successful retrieval of a variant of GCF_000716685.1.contig0007.region001, i.e. a sequence close but different from the reference, with an intact core between 35-40kb. Specifically, the consensus sequence of the retrieved BGC has a percent of identity of 92%

Claims

CLAIMS1. A method for identifying at least one gene cluster of interest, comprising steps of: a) providing a clone library comprising a plurality of clones with high-molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest; b) partitioning the clone library into at least a first and second plurality of separate partitions: the first plurality of separate partitions comprising a plurality of clones, the plurality of separate partitions thereby forming a plurality of clone pools; the second plurality of separate partitions comprising a plurality of clones from clone pools, thereby forming a plurality of clone stacks; wherein the high-molecular weight nucleic acids from the clone pools and the clone stacks are fractionated; c) determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acids in the clone pools and in the clone stacks, thereby providing a plurality of sequence reads, the sequencing depth corresponding to said plurality of sequence reads being of less than 30x; thereby identifying the at least one gene cluster of interest.

2. The method according to claim 1, for identifying a plurality of gene clusters in eukaryotic cell(s) or a clone library thereof; said eukaryotic cell being in particular a fungus, preferably a filamentous fungus, for example belonging to the genus Aspergillus.

3. The method according to claim 1 or 2, for identifying a plurality of biosynthetic gene clusters (BGC) in a fungus or a clone library thereof.

4. The method according to claim 3, wherein the plurality of biosynthetic gene clusters (BGC) is for the synthesis of non-ribosomal peptides (NRPS), polyketides (PKS), terpenes, aminoglycosides, bacteriocins, lassopeptides, lantipeptides, nucleosides, phosphonates, or combinations thereof.

5. The method according to any of claims 1 to 4, wherein at least one gene cluster of interest is a BGC for the synthesis of NRPS or PKS.

6. The method according to any one of claims 1 to 5, wherein said at least one gene cluster of interest has a size of at least 20 kbp, for example ranging from 20 kbp to 200 kbp.

7. The method according to any one of claims 1 to 6, the clone library is characterized in that the size of the high-molecular weight nucleic acid is of at least 20 kbp, for example ranging from 20 kbp to 200 kbp.

8. The method according to any one of claims 1 to 7, wherein the clone library is characterized in that it consists of a cloning vector library selected from a group consisting of: cosmid, fosmid, phage, bacterial artificial chromosome (BAC), Pl- derived artificial chromosome (PAC), yeast artificial chromosome (YAC), fungal artificial chromosome (FAC).

9. The method according to any one of claims 1 to 8, wherein the clone library comprises or consists of a shuttle vector which is suitable for replication in a plurality of cells belonging to distinct genus or kingdoms, more particularly one belonging to the genus Streptomyces and the other to the genus Escherichia', preferably a shuttle vector which is suitable for replication in E. coli and mobilization into Streptomyces spp.

10. The method according to any of claims 1 to 9, comprising steps of: a) providing a clone library comprising a plurality of clones with high-molecular weight nucleic acid, said nucleic acid being susceptible to encode or comprise gene clusters of interest;b) partitioning the clone library into a first plurality of separate partitions comprising a plurality of clones, thereby forming a plurality of clone pools, fractionating the high-molecular weight nucleic acid from the clone pools, and determining the nucleic acid sequence of all or part of the fractionated high- molecular weight nucleic acid in the clone pools, thereby determining the plurality of clone pools comprising at least one occurrence of said at least one gene cluster of interest; c) partitioning the clone library into a second plurality of separate partitions, comprising a clone from each clone pool, thereby forming a plurality of clone stacks, fractionating the high-molecular weight nucleic acid from the clone stacks, and determining the nucleic acid sequence of all or part of the fractionated high-molecular weight nucleic acid in the clone stacks; d) determining the clones comprising the gene cluster(s) of interest in the library, based on the nucleic acid sequences from the clone stacks, thereby identifying the at least one gene cluster of interest; wherein the clone pools forming the clone stacks, statistically comprise at least one occurrence of the gene cluster of interest.

11. The method according to any of claims 1 to 10, which includes a plurality of steps of partitioning the clone library; thereby forming more than two partitions.

12. The method according to any one of claims 1 to 11; the step of determining the sequence of high-molecular weight nucleic acids in the clone pool and the clone stack comprising:(i) determining the nucleic acid sequence of all or part of the corresponding fractionated high-molecular weight nucleic acids, thereby providing a plurality of sequence reads;(ii) aligning the sequence reads with a reference sequence, thereby determining all or part of the sequence of the high-molecular weight nucleic acids.

13. The method according to claim 12, wherein the sequencing depth corresponding to the plurality of sequence reads is of less than 20x; for example less than lx.

14. The method according to any one of claims 1 to 13, wherein each clone pool from the plurality of clone pools statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest; and / or each clone stack from the plurality of clone stacks statistically comprises less than two, in particular at most one, occurrence of said at least one gene cluster of interest.

15. The method according to any one of claims 1 to 14, which comprises a step of recovering, preferably isolating, the clone comprising the at least one gene cluster of interest from the library.

Citation Information

Patent Citations

  • Methods and apparatus for efficient and accurate assembly of long-read genomic sequences

    US20230049048A1

  • Compositions and methods for identifying gene clusters

    WO2021041397A2