Method for target metagenomic sequencing of multiple microbial gene families with full phylogenetic coverage in a single experiment

The novel metagenomics sequencing method using in-solution DNA hybridization and a custom enrichment panel addresses the limitations of current methods by achieving strain-level resolution and detecting low-abundance microorganisms with enhanced phylogenetic coverage and sensitivity, significantly improving species detection.

WO2026082975A1PCT designated stage Publication Date: 2026-04-23CONSEJO SUPERIOR DE INVESTIGACIONES CIENTIFICAS (CSIC) +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CONSEJO SUPERIOR DE INVESTIGACIONES CIENTIFICAS (CSIC)
Filing Date
2025-10-17
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current metagenomic sequencing methods struggle to provide full phylogenetic coverage and strain-level resolution for low-abundance microorganisms, particularly those in the rare biosphere, due to biases in existing approaches that rely on known reference genomes and are limited by the genetic material of abundant taxa.

Method used

A novel non-PCR-based target metagenomics sequencing method using in-solution DNA hybridization techniques with a custom enrichment panel designed through an iterative, data-driven probe optimization workflow, incorporating genomic and metagenomic sequences, and applying multiple-threshold sequence clustering and off-target filtering to maximize phylogenetic coverage and sensitivity.

Benefits of technology

Enables the detection and phylogenomic characterization of unknown microorganisms with strain-level resolution, uncovering a large amount of low-abundance microorganisms and increasing species detection by up to five times compared to conventional methods, while preserving accurate relative abundance estimates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000026_0001
    Figure IMGF000026_0001
  • Figure 00000034_0000
    Figure 00000034_0000
  • Figure 00000034_0001
    Figure 00000034_0001
Patent Text Reader

Abstract

The vast majority of prokaryotic microorganisms remains uncultured and are not abundant enough to be effectively studied using metagenomic sequencing methods. Here, we present a novel PCR-free target capture sequencing approach that enables targeting specific genes or functions with full phylogenetic coverage in a single sequencing experiment, proving also a higher sensitivity and precision than regular shotgun and amplicon-based approaches.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for target metagenomic sequencing of multiple microbial gene families with full phylogenetic coverage in a single experiment.

[0002] TECHNICAL FIELD

[0003] The present invention refers to a method for targeting, in a single sequencing experiment of an environmental or biological sample, multiple microbial gene families, providing full phylogenetic coverage, high-sensitivity to detect low-abundance organisms and strain-level precision in the microorganism’s identification.

[0004] BACKGROUND ART

[0005] Microorganisms dominate the earth and have shaped the planet as we know it. However, despite their importance, our current knowledge of the microbial world remains very limited. Today, we know that microbial biodiversity, both at the phylogenetic and functional levels, is dominated by a large number of species that we have barely started to understand and even detect. The vast majority of them have not been cultured in the lab, and many more are not abundant enough to be studied with current genomic approaches.

[0006] This huge fraction of elusive microbes, usually referred to as the rare biosphere, is considered of utmost importance to understand and predict microbiome responses to environmental change, especially in highly biodiverse settings such as free-living habitats. It is known that the rare biosphere harbors keystone species that, even at very low abundance, can modulate nitrogen fixation, phosphorus cycling, and sulfur and carbon flows. The rare members of microbial communities are also hypothesized to represent a large reservoir of xenobiotic metabolism (i.e., involved in biodegradation), therefore determining the resilience of microbial ecosystems to pollutants, pesticides, and antibiotics. Likewise, it is believed that latent organisms in the rare biosphere can heavily affect the tolerance of natural habitats to climate change.

[0007] However, uncovering low abundant microorganisms using metagenomics and metataxonomic techniques is extremely challenging. Amplicon-based sequencing of marker genes (e.g., 16S, 18S, ITS) provides predictions with limited phylogenetic resolution, entails important technical biases associated with their amplification, produces inaccurate abundance profiles due to copy-number variation, and misses the detection of phylogenetically distant unknown lineages (Bahram et al. 2019). Full length 16S sequencing mitigates some of these problems, achieving nearly specieslevel resolution, but it cannot reach strain resolution and does not allow for phylogenomic approaches. Similarly, PCR amplification of universal single-copy markers has been previously tested as an unbiased and higher resolution alternative to 16S and whole-genome sequencing. However, PCR approaches are only feasible for targeting specific genes from specific species, therefore incapable of providing wide phylogenetic coverage.

[0008] On the other hand, shotgun metagenomic sequencing offers a comprehensive view of the genomes present across entire microbial communities, but the sequenced data is typically saturated with the genetic material of the most-abundant members. To date, even the highest-depth sequencing efforts in metagenomics can only provide random glimpses of the genetic variability present in the rare biosphere.

[0009] Existing computational and experimental approaches have attempted to address these limitations. For instance, tools such as MetaPhlAn and mOTUs enable taxonomic profiling of metagenomic data by each fragment cluster is mapped back to the same sources and scored per OG using a positive / negative match ratio; any cluster with a ratio below 1 is excluded, which removes probe candidates that tend to cross-hit other OGs mapping reads against databases of precomputed marker genes. However, these methods rely entirely on previously known reference genomes and therefore fail to detect unknown or uncharacterized species not represented in such databases. Furthermore, they depend on shotgun sequencing data, which remains biased toward abundant taxa.

[0010] Other strategies, such as probe-based target capture systems (e.g., MetCap and CATCH), have demonstrated the potential of in-solution hybridization to enrich metagenomic DNA for specific genes or taxonomic groups. MetCap, for example, provides a general pipeline to design capture probes for functional metagenomics, while CATCH optimizes probe sets to cover known viral or bacterial diversity using compact oligonucleotide collections. Nevertheless, both approaches rely primarily on the diversity represented in existing genome repositories and lack explicit mechanisms to include metagenomic sequences from uncultivated organisms or to optimize probe selection based on quantitative and qualitative (i.e. maximizing accuracy) phylogenetic coverage metrics. In addition, they are not designed to produce a single comprehensive panel capable of covering entire prokaryotic diversity across multiple marker gene families in one capture experiment.

[0011] In summary, there are currently no methods capable of uncovering all the genetic variability of one or multiple specific gene families in a metagenomic sample, with full phylogenetic coverage, quantitative optimization, and sensitivity sufficient to detect both high- and low-abundance microorganisms in a single experiment, preferably at strain level resolution.

[0012] In the present invention, we propose a novel non PCR-based target metagenomics sequencing method based on the use of in-solution DNA hybridization techniques, prior to shotgun sequencing, to capture and enrich all DNA fragments in a sample corresponding to the desired pool of gene families. In-solution hybridization strategies are known to deliver much higher resolution in the detection of novel genetic variants than traditional PCR and Sanger sequencing, but their use for taxonomic and functional profiling of metagenomic samples is currently limited by the high number of sequence probes needed to cover all microbial biodiversity at once, and the lack of sufficiently accurate probes to target particular gene families across all microorganisms.

[0013] A key component of the invention, therefore, lies in our methodology to design custom enrichment panels composed of an optimal set of nucleotide sequence probes that (1 ) provide high sensitivity and precision in targeting the desired gene families, reaching DNA material in very low abundance; (2) cover all known sequence variability of the desired pool of gene families; (3) incorporate potential sequence diversity derived from uncultivated or uncharacterized microorganisms present in public metagenomic datasets; and (4) can be manufactured in a single capture hybridization panel using current technology.

[0014] Unlike prior approaches such as MetCap or CATCH, the present method introduces an iterative, data-driven probe optimization workflow that integrates genomic and metagenomic sequences, performs multiple-threshold sequence clustering (NSI- based), applies off-target filtering by positive / negative hit ratios, and uses explicit quantitative metrics (RecovT, RecovLen, and PhyloC) to maximize phylogenetic coverage while minimizing the total number of probes. This allows the generation of a single hybridization panel with full prokaryotic coverage and strain-level taxonomic resolution, enabling the discovery of previously undescribed microbial taxa. Other existing hybridization capture methods, such as Gasc & Peyret (https: / / pubmed.ncbi.nlm.nih.gov / 29587880 / ), focus primarily on single-gene enrichment, often 16S rRNA, and cannot simultaneously target multiple independent marker genes or integrate phylogenetic information across genes. As a result, these methods are limited in their ability to resolve strain-level diversity or to accurately place unknown taxa, especially at basal branches (see Jhonson et al. 2019).

[0015] The present invention addresses these limitations by providing a method for designing custom capture panels targeting multiple gene families in parallel, covering the full sequence variability present in both cultivated and uncultivated microbes. The multigene capture approach provides sufficient phylogenetic signal to accurately place previously undescribed or unknown species within the prokaryotic tree, including basal branches, while simultaneously detecting rare and low-abundance organisms.

[0016] One example of our custom capture panel uses a set of 580,369 probes optimized to cover all known biodiversity of ten universally conserved marker genes commonly used for high-precision taxonomic profiling, allowing the identification of low-abundance species across the entire bacterial phylogeny and reaching strain-level resolution in the taxonomic identification (Figure 1 and 2). By validating our method on 57 samples from different habitats, including human gut microbiome, soil and extreme environments, we show that the method provides unprecedented taxonomic coverage and phylogenetic resolution on samples of diverse environments, outperforming standard shotgun metagenomic sequencing, amplicon-based methods.

[0017] BRIEF DESCRIPTION OF THE FIGURE

[0018] Figure 1. Capture efficiency and taxonomic profiling from DNA capture of 10 singlecopy universal marker genes in samples from human gut (1st column), soil (2nd column) and microbial mats (3rd column). A) Proportion of on-target reads (on-target reads I total clean reads, Iog10 scale), arranged by intervals of % of identity computed from alignments of on-target reads to the original MGs used to design the probes, for capture (red) and shotgun (blue) data. B) mOTUs (-g3) relative abundance capture vs shotgun, with the same number of 2x100 pair reads (standardized set) for each capture (red) I shotgun (blue) sample pair (standardized set). C) Number of observed taxa across all taxonomic ranks with mOTUs (-g3; “motus_c”, “motus_s”; capture and shotgun data, respectively)), DECIPHER (dpher_c, dpher_s; capture and shotgun data, respectively), MetaPhlAn 4 (“mpan”) and Kraken2 (“kraken”), with the same number of 2x100 pair reads (standardized set) for each capture I shotgun sample pair (standardized set). D) Number of observed taxa across all taxonomic ranks with 16S (Silva DB; “16S_silva”), 16S (GTDB DB, “16S_GTDB”), mOTUs (capture data, -g 3, “motus_c”), DECIPHER (capture data, “dpher_c”).

[0019] Figure 2 Comparative of the performance of targeted capture metagenomic sequencing and standard metagenomics methods in taxonomic profiling. (A-B) Number of bacterial species detected using targeted capture metagenomic sequencing (analyzed with mOTUs) compared to shotgun metagenomics (mOTUs, Kraken2, MetaPhlAn4) and 16S rRNA amplicon sequencing across human gut (A) and soil (B) samples. (C) Number of bacterial strains detected per sample using SameStr (Podlesny et al, 2022) based on single-nucleotide variant (SNV) profiles. Comparisons were made using MetaPhlAn marker genes for shotgun data, and the ten mOTUs marker genes for shotgun and capture-based datasets. (D) Correlation between relative abundances between shotgun and capture-based metagenomic sequencing in a representative human gut sample. (E) Number of unclassified or novel bacterial species detected at different taxonomic ranks (order, family, genus).

[0020] DESCRIPTION OF EMBODIMENTS OF THE INVENTION

[0021] The vast majority of prokaryotic microorganisms remains uncultured and are not abundant enough to be effectively studied using metagenomic sequencing methods. Here, we present a novel PCR-free target capture sequencing approach that enables targeting specific genes or functions with full phylogenetic coverage in a single sequencing experiment, proving also a higher sensitivity and precision than regular shotgun and amplicon-based approaches. In contrast to existing profiling tools such as MetaPhlAn or mOTUs, which rely on the detection of known reference markers in shotgun datasets, and in contrast to capture-design systems such as MetCap and CATCH, which are limited to pre-defined genomic diversity, the present method integrates both genomic and metagenomic sequence data to design capture probes that represent the full known and potential diversity of target gene families.

[0022] This integrated workflow achieves, for the first time, a single, comprehensive capture panel capable of simultaneously enriching low- and high-abundance taxa across the full prokaryotic tree of life, with strain-level resolution. It allows detection and phylogenomic characterization of unknown microorganisms not present in current databases, while preserving accurate relative abundance estimates of known species.

[0023] We demonstrate the validity and practical application of our approach by using the method to design a capture-kit panel targeting ten universally-conserved marker genes, enabling the accurate identification of low abundant microorganisms with broad phylogenetic coverage. When applied on human gut, soil and extreme environments samples, our method identifies up to 5 times more species than conventional shotgun or 16S sequencing, achieving strain level resolution and enabling population metagenomic analysis. Furthermore, it unveils a large amount of unknown low- abundance microorganisms present in the rare microbial biosphere. Therefore, one aspect of the present invention refers to an in vitro or computer implemented method for providing or designing a collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions, said method comprising the steps of: a. Step 1 . Gene family selection: Compiling a list of marker gene families or functional orthologous groups (OG) by, preferably, identifying genes from a database within or associated to a specified group of molecular interactions and reaction networks; b. Step 2 Data mining: Identifying and providing from one or more databases, a data corpus or dataset of nucleotide sequences having a percentage of nucleotide identity with the genes and OGs identified in step 1 of at least 50%, 60%, 70%, 80%, 90%, 95%, 99% or more; c. Step 3. Design of capture-kits: Analyzing the data corpus or dataset provided in Step 2 in order to reduce its dimensionality to make it accessible for a single targeted sequencing experiment, preferably while maintaining the full phylogenetic coverage, high sensibility and precision; and d. Providing the design of the panel or collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions selected in Step 1 ; and e. Optionally, manufacturing or producing the panel of step d.

[0024] In certain embodiments, step d) is performed by a computer-implemented design process that, based on the dimensionality-reduced set from Step 3, computes and outputs a machine-readable specification (e.g., a file or data structure) of a collection of sequencing probes targeted to the OGs selected in Step 1 and / or 2. The collection is algorithmically constrained such that it enables targeting with full phylogenetic coverage and high sensitivity, as determined by pre-set Step-xi performance metrics (PhyloC > C_min, RecovT > T_min, RecovLen > L_min). Accordingly, the technical contribution of Step d is the generation of this actionable, machine-readable probe design that transforms heterogeneous sequence data into a concrete, experimentready specification meeting explicit coverage / sensitivity thresholds and singleexperiment feasibility constraints. The method does not directly obtain, manufacture, or produce a physical panel at Step d; rather, the output of Step d can be used in turn by any suitable downstream synthesis, fabrication, or assembly techniques to create the physical panel referenced in optional Step e.

[0025] Alternatively step d) can be drafted as outputting, by the computer, a machine-readable specification of a collection of sequencing probes for the OGs of Step 1 and 2, the collection enabling targeting with full phylogenetic coverage and high sensitivity as determined by the Step xi metrics (PhyloC > C_min, RecovT > T_min, RecovLen > L_min).

[0026] In certain embodiments, Step (d) comprises, by operation of one or more processors, automatically generating and storing a machine-readable specification defining a collection of sequencing probes targeted to the orthologous groups selected in Steps (1 ) and / or (2), the collection being computed from the dimensionality-reduced set of Step (3) such that the predefined performance constraints PhyloC > C_min, RecovT > T_min, and RecovLen > L_min are satisfied for a single targeted sequencing experiment.

[0027] In certain embodiments, Step (d) comprises producing, by the computer, a non- transitory computer-readable file comprising: for each probe, a sequence string, a probe identifier, an intended modification tag, an orthologous-group mapping, an intended pool identifier, and a recommended representation value, wherein the file encodes only those probes that together meet PhyloC, RecovT, and RecovLen metrics under single-experiment feasibility limits.

[0028] In yet other embodiments, Step (d) comprises generating, by the computer, a probe design package consisting of a manifest, a pooling map, and validation metadata, the validation metadata comprising machine-computed estimates of per-OG coverage breadth, predicted recovery length distribution, and expected on-target fraction, each meeting or exceeding C_min, L_min, and T_min, respectively.

[0029] In some embodiments, Step (d) comprises outputting, by the computer, a probe collection constrained by a maximum aggregate probe length and a maximum probe count derived from single-experiment instrument capacity, while guaranteeing that for each OG the fraction of representative sequences captured at or above a defined hybridization identity threshold yields PhyloC > C_min.

[0030] In additional embodiments, Step (d) comprises computing, by the computer, a set of probes partitioned into one or more sub-pools, each sub-pool satisfying the Step-xi metrics for a designated subset of OGs, and outputting a pool-aware specification enabling modular use in single or multiple targeted experiments without re-design.

[0031] In other embodiments, Step (d) comprises applying, by the computer, a learned selection model trained on historical capture outcomes to rank candidate probes from Step (3), selecting top-ranked probes subject to the Step-xi constraints, and exporting a manifest and accompanying model-derived confidence scores per OG that attest to meeting C_min, T_min, and L_min.

[0032] In certain embodiments, Step (d) comprises computing, by the computer, a deterministic hash of the dimensionality-reduced input of Step (3) and binding said hash to the emitted probe manifest as a design-integrity token, whereby the probe specification is cryptographically linked to the precise input and the Step-xi thresholds achieved.

[0033] In further embodiments, Step (d) comprises emitting, by the computer, an interchange specification conforming to a predefined schema, the schema requiring fields sufficient for automated synthesis and quality control planning, and the emitted instance being validated at design time to demonstrate conformance to PhyloC > C_min, RecovT > T_min, and RecovLen > L_min.

[0034] In some embodiments, Step (d) comprises outputting, by the computer, both a primary probe set that satisfies the Step-xi metrics and an auxiliary replacement set for probes predicted to underperform in specific GC or variant contexts, together with a deterministic rule set describing how auxiliary probes may be swapped without violating PhyloC, RecovT, or RecovLen.

[0035] In a further embodiment of the invention, upon receipt of the machine-readable probe specification generated in Step (d), a physical capture panel can be produced by synthesizing the listed oligonucleotide probes according to the sequences and perprobe attributes encoded therein. In certain embodiments, the probes are DNA oligonucleotides synthesized by array or column methods; in alternative embodiments, the probes are RNA baits enzymatically transcribed from DNA masters. Each probe may include prescribed chemical features, e.g., a 5' affinity tag and optional nonhybridizing handles, consistent with the design output and intended capture workflow. Following synthesis, probes are released, deprotected, and purified to vendor-grade quality appropriate for hybrid-capture applications, producing full-length species suitable for pooling. The probes are then quantified and normalized to establish a defined molar representation per the design’s pooling map, which may specify uniform or intentionally weighted abundance across probes, gene families, GC bins, or subpanels. Normalized probes are combined into one or more pools to yield a panel that, as a whole, targets the orthologous groups selected in Step (1 ) / (2) and constrained by the dimensionality reduction for single-experiment feasibility.

[0036] Accordingly, Step (e) yields a physical panel preferably comprising a normalized, quality-controlled collection of capture probes manufactured directly from the Step-(d) output and preferably packaged for shipment and use, the panel being configured to provide, in a single targeted sequencing experiment, full phylogenetic coverage and high sensitivity across the selected gene families while meeting explicit feasibility and performance constraints.

[0037] In a preferred embodiment, in Step 2 the data corpus of nucleotide sequences is provided by retrieving for each of the selected OGs of step 1 , the sequences from (i) global genomic repositories and (ii) public metagenomic datasets.

[0038] In another preferred embodiment, Step 3 is carried out by a method comprising the following steps i) to iv) for each targeted OG (tOG) included in a kit: i. Decomposing into non-overlapping chunks of from 10 - 200 pair bases, preferably from 50 to 150 pair bases, more preferably 80 or 120 pair bases, all or part of the nucleotide sequences comprising the dataset obtained in Step 2 for each of the tOG, herein referred to as probes; ii. Aligning the nucleotide sequences comprising the dataset obtained in step 2 into a multiple sequence alignment (MSA), where the conservation level of each aligned column is calculated; iii. optionally the multiple sequence alignment produced is used to generate a phylogenetic tree of all or part of the non-redundant sequences; iv. All or part of the generated probes in step i) are compared in an all-against-all fashion to calculate their pairwise sequence similarity, computed as the % of nucleotides identities observed between two probes and provide an all-against-all probe comparison matrix; wherein the method further comprises the following steps: v. Using the all-against-all probe comparison matrix of step iv to generate clusters of probes at the NSI (Nucleotide Sequence Identity) threshold, wherein the NSI threshold is selected from one or more values in the range of from 30% to 100%; vi. Mapping all or part of the probe clusters of step v back to the same original genomic and metagenomic databases used in Step 2 using BLASTN searches; vii. Analyzing to compute the ratio of positive matches (i.e. hits against a gene belonging to the expected OG) over negative matches (i.e. hits against a gene belonging to an unwanted OG) of the hits produced in step vi; viii. Discarding all probe clusters with a ratio of positive matches lower than 1.0 (i.e. probes producing at least one mismatch) from the set; ix. Mapping all or part of the remaining selected probe clusters to the MSA generated in step ii), preferably by recording the exact position of matching residues of the original probe on the MSA; x. Discarding probes matching columns of the alignment produced in ii with a residue conservation score, preferably of below 0.75; and xi. Mapping back the final set of selected probes to the original sequences selected in Step 2, preferably by using BLASTN, and preferably calculating: number of original sequences recovered by the probes (RecovT), the average length of the original sequence covered by the probe-selection (RecovLen) and the phylogenetic coverage based on the the generated in substep C (PhyloC).

[0039] Preferably, in connection to the above embodiment, for each tOG the total number of probes selected (Nprobes), together with their RecovT, RecovL and PhyloC values obtained for each clustering experiment under each NSI are automatically evaluated to select the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing NProbes.

[0040] More preferably, the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing Nprobes, for each tOG, is combined into a single dataset, and providing the final design of the panel. Still, more preferably, the method further comprises the manufacturing or production of the said panel of probes.

[0041] It is noted that although probe-design tools such as MetCap and CATCH provide strategies to design capture probes from collections of sequences, the present method is distinct in that it (i) integrates both genomic and public metagenomic sequences as primary input to capture uncultivated diversity, (ii) decomposes target OG sequences into non-overlapping probe candidates and applies iterative NSI clustering across multiple thresholds, (iii) maps probe clusters back to sequence repositories and discards clusters with any off-target hits (ratio positive / negative < 1.0), (iv) maps surviving clusters to the MSA and discards probes matching low-conservation columns (conservation < 0.75), and (v) selects final probe sets by simultaneously maximizing RecovT, RecovLen and PhyloC while minimizing Nprobes. This combination of steps yields a single capture panel that achieves (and we demonstrate) substantially greater species recovery, rare-taxa detection, and strain resolution than prior capture designs. A further aspect of the invention refers to a physical panel obtained or manufactured or produced as described above, wherein the probes are polynucleotides, preferably DNA or RNA, probes. Preferably, this panel is used for capturing and target enrichment assays to identify organisms, either known or unknown organisms.

[0042] A further embodiment of the invention refers to the“panel” referenced above and obtained in step e) for capture and target enrichment assays directly on environmental or biological DNA samples prior to shotgun sequencing, for obtaining a comprehensive taxonomic profile of a microbial community with full phylogenetic coverage, including simultaneous detection of high- and low-abundance taxa, with strain-level resolution, and comprising the detection and characterization of previously undescribed or unknown microbial species or strains not present in current databases. Preferably, said use, can be done in combination with standard taxonomic profiling or metagenomic sequencing methods, wherein said combined use significantly increases the number of microbial species and strains identified compared to untargeted shotgun sequencing alone. More preferably, said can be carry out in combination with standard taxonomic profiling methods, wherein said combined use does not significantly alter the relative abundance estimations of the identified microorganisms compared to untargeted shotgun sequencing, thereby preserving quantitative accuracy while enhancing taxonomic resolution. Also preferably, said panel can be used for capturing and enriching multiple gene families from environmental or biological DNA samples, wherein the captured genes are used in combination to: a) identify microbial species and strains, including previously undescribed or unknown organisms; which are undetectable by standard shotgun metagenomic sequencing; b) place the identified organisms more accurately within the prokaryotic phylogeny, by deriving phylogenetic placement support from multiple independent marker genes, thereby increasing the robustness and resolution of phylogenetic inference.

[0043] The combined use of multiple marker genes enables not only detection of novel species but also accurate phylogenetic placement, including basal and deep-branch lineages, which is not achievable using 16S rRNA alone due to its limited phylogenetic signal for basal branches.

[0044] Further aspects of the present invention refer to a kit comprising the panel obtained or manufactured or produced by the above method.

[0045] Said kit or panel can be useful for: a. Hybridizing custom capture-kits with environmental DNA samples prior to sequencing to specifically select the targeted genes b. Shotgun sequencing of the enriched DNA sample DEFINITIONS

[0046] As used herein, “functional orthologous groups (OG)” shall be understood as a comprehensive group of nucleotide sequences representing genes from all known species with a common evolutionary origin and molecular and cellular function.

[0047] As used herein “marker gene families”, shall be understood as a set of functional orthologous groups (OGs) that are representative for a given microbial function or characteristic. For instance, phylogenetic marker gene families refer to OGs that are particularly suitable to establish the phylogenetic position and evolutionary relationships among species, as they are composed of highly coserved genes involved in the molecular functions present in all living organisms. Similarly, marker genes families for the microbial function associated with nitrogen fixation would refer to one or various OGs where genes encode for molecular functions highly specific for the process of nitrogen fixation.

[0048] As used herein, “group of molecular functions” shall be understood as the molecular activities performed by a genes.

[0049] As used herein, “full phylogenetic coverage” shall be understood as the fact of sequencing, detecting or using genes from all microbial species known to date.

[0050] As used herein, when referring to “the conservation level of each column is calculated” shall be understood as the assessment of the probability that an specific nucleotide base (i.e. A, C, T or G) is the same in all genes from the same OG.

[0051] As used herein, when referring to “decomposing into non-overlapping chunks” shall be understood as the fragmentation of long DNA strings into smaller consecutive strings.

[0052] As used herein, “a phylogenetic tree” shall be understood as the tree-like data structure that represents the evolu tionary relationships among all genes / sequences.

[0053] As used herein, “non-redundant sequences” shall be understood as unique sequences that differ from the rest in at least one nucleotide.

[0054] As used herein, “compared in an all-against-all fashion” shall be understood as the comparison of all possible pairs of sequences.

[0055] As used herein, “calculate their pairwise sequence similarity” shall be understood as the calculation of identical nucleotide positions between two sequences, expressed as percentage of the total length of the shortest sequence. As used herein, “all-against-all probe comparison matrix” shall be understood as the matrix of nucleotide identity values obtained for all pairwise sequence comparisons.

[0056] As used herein, “clusters of probes at the NSI threshold” shall be understood as the process of grouping probes based on their nucleotide sequence identity (NSI). For instance, clustering at NSI=80% would mean that all probes with a 80% NSI or higher are grouped together into a single cluster.

[0057] As used herein, “mapping back to the same original genomic and metagenomic databases using BLASTN searches” shall be understood as the comparison of nucleotide probe sequences against a sequence database using the blast-n software. As used herein, the phrase “maximizes RecovT, RecovL and PhyloC, while minimizing Nprobes, for each tOG” shall be understood as the automated process by which the methods select s the set of probes that contains the least num ber of probes, while maintaining nearly full phylogenetic coverage and high sensitivity and precision.

[0058] DETALED DESCRIPTION

[0059] As indicated previously, in the present invention, we present a novel PCR-free target metagenomics sequencing approach based on a custom capture enrichment panel covering the entire prokaryotic phylogeny. Such custom capture enrichment panel (comprising capture sequence probes that cover an entire prokaryotic phylogeny) can be obtained by an in vitro or computer implemented method in accordance with the first aspect of the present invention,

[0060] In the present invention, we postulate that full phylogenetic coverage for many key gene families could be implemented in relatively inexpensive capture kits, increasing the precision and coverage of current functional and taxonomic profiling methodologies. This is supported by (i) the great level of customization presently available for the design of sequence probes, (ii) the available size range of current capture panels (admitting up to hundreds of thousands of different probes), and (iii) the estimated extent of the genetic variability for most important marker genes in ecology. Although there is no record in the literature for such a broad implementation, our tests performed during the last years indicate that high-depth exploration of key functional and phylogenetic markers genes from rare genomes are technically feasible, provided the manufacturing of capture kits and target enrichment protocols are combined with phylogenetically-driven design of target probes. In this sense, in the present invention we provide a study showing that, indeed, a single capture kit based on ten phylogenetic marker gene families suffices to cover all known prokaryotic diversity: ~125,000 references prokaryotic species and over 1 million strains , detecting microbial species at levels of abundance two orders of magnitude below of what regular shotgun sequencing can reach , and uncovering three times as much biodiversity (Figure 1 and Figure 2A and 2B). In addition, our capture-based sequencing approach identified significantly more strains than two independent methodologies based on standard shotgun sequencing (see Figure 2C). These species and strain detections based on capture were further predicted with their relative abundance, which has also been proven to be unbiased compared to regular shotgun approaches (Figure 2D), proving that our method can also be used to estimate relative abundances for rare taxa. Furthermore, our results uncover hundreds of unclassified on-target sequences lacking close phylogenetic relatives, which might represent detections for never-seen low abundance organisms and that are undetectable by current shotgun approaches (Figure 2E). This is a desirable effect derived from the fact that template probes can be designed to capture, not only the diversity of all known genes, but also a certain degree (up to 30%) of unknown genetic variability. The above results were obtained in our laboratory by testing a pilot custom capture kit and applying it on a variety of environmental and human gut samples, using the same sequencing depth for both shotgun and capture plus shotgun experiments. These results prove that it is technically possible to design broad hypothesis-driven capture panels (i.e., spanning vast phylogenetic distances) targeting, not only phylogenetic marker genes, but also dozens of ecologically relevant gene functions.

[0061] In particular, the procedure, in particular the computer implemented method, for designing and manufacturing broad hypothesis-driven capture panels provided by the present invention comprises the following general steps:

[0062] - Step 1 . Gene family selection:

[0063] In order to maximize the functional coverage, a list of marker gene families needs to be compiled. That is, there is a need to identify genes from a database within or associated to a specified group of molecular interactions and reaction networks.

[0064] For example, to design and manufacture a kit targeting the precise identification of all microorganisms present in a metagenomics sample, we used a set of universal, single copy phylogenetic markers genes. Converserly, to unveil all the biodiversity of metabolic genes involved in the degradation of soil pollutants, we would target maker enzymes and transporters related to the biodegradation of drugs (including antibiotics), dioxins, hydrocarbures, benzones, and other pollutants such as heavy metal resistance genes.

[0065] Thus, for this first step, we shall depart from the already known sequences associated with each of the functional descriptions considered as a functional marker. This is typically referred to as a functional orthologous group (OG).

[0066] - Step 2 Data mining to maximize phylogenetic coverage:

[0067] In this step, we shall identify from one or more databases nucleotide sequences having a percentage of identity with the genes and OGs identified in step 1 of at least 60%, 70%, 80%, 90%, 95%, 99% or more. For our specific examples, for each of the selected OGs, we will pull all available sequences from (i) global genomic repositories and (ii) public metagenomic datasets. By mining both genomic and metagenomic variability, we aim at ensuring that subsequent enrichment experiments are not biased towards data from cultivated taxa or specific lineages. The resulting corpus of genomic and metagenomics sequences matching each of the targeted OG represents all known gene sequence variability across all cultivated and uncultivated microbial organisms. Due to the volume and extent of such corpus of information, it is currently unfeasible to target all such variability at once in a single sequencing experiment

[0068] Step 3. Design of capture-kits:

[0069] In this step, we analyze the data corpus produced in Step 2 in order to reduce its dimensionality to make it accessible for a single targeted sequencing experiment, while maintaining the same phylogenetic coverage and sensibility and improving its precision. This involved a series of in-vitro computational analysis and actions as detailed in the following pseudocode:

[0070] Loop 1 : For each targeted OG (tOG) included in a kit:

[0071] A. All nucleotide sequences comprising the dataset obtained in step 2 for the tOG are decomposed into non-overlapping chunks of 80 pair bases, herein referred to as probes.

[0072] B. All nucleotide sequences comprising the dataset obtained in step 2 are aligned into a multiple sequence alignment, where the conservation level of each column is calculated.

[0073] C. Optionally the multiple sequence alignment produced is used to generate a phylogenetic tree of all non-redundant sequences.

[0074] D. All the generated probes in substep A are compared in an all-against-al I fashion to calculate their pairwise sequence similarity, computed as the % of nucleotides identities observed between two probes.

[0075] Loop 2: For each of the possible nucleotide sequence identity (NSI) in 100%, 90%, 80%, 70%:

[0076] E. The all-against-all probe comparison matrix is then used to generate clusters of probes at the NSI threshold. Here, we intend to produce a reduced set of probes compatible with the current sensitivity of capture protocols, which varies from 100% to 30% NSI. This is possible because, for in-solution hybridization capture kits, capture probes can be used in the same way as degenerate primers can, albeit providing capture ranges of up to 30% divergence from the original template

[0077] F. All probe clusters are mapped back to the same original genomic and metagenomic databases used in step 2 using BLASTN searches.

[0078] G. The significant hits produced in substep F are analyzed to compute the ratio of positive matches (i.e. hits against a gene belonging to the expected OG) over negative matches (i.e. hits against a gene belonging to an unwanted OG). Notice that, although all probes were generated from targeted OG genes, many might belong to promiscuous gene regions that exist in many other OGs, potentially dismissing the precision of the targeted OG detections.

[0079] H. All probe clusters with a ratio of positive matches lower than 1.0 (i.e. probes producing at least one mismatch) are discarded from the set, therefore keeping only highly specific probes of the tOG and reducing potential off-target hybridization in the kit.

[0080] I. All the remaining selected probe clusters are then mapped to the MSA generated in substep B), recording the exact position of matching residues of the original probe on the MSA.

[0081] J. Probes matching columns of the alignment produced in B with a residue conservation score below 0.75 are discarded, keeping only specific probes that are also largely universal for the tOGs (i.e. discarding marginal probes).

[0082] K. The final set of selected probes is then mapped back to the original sequences selected in Step 2 using BLASTN, calculating: number of original sequences recovered by the probes (RecovT), the average length of the original sequence covered by the probe-selection (RecovLen) and the phylogenetic coverage based on the the generated in substep C (PhyloC)

[0083] The total number of probes selected (Nprobes), together with their RecovT, RecovL and PhyloC values obtained for each clustering experiment under each NSI are automatically evaluated to select the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing NProbes. Finally, the best probe set for each tOG is combined into a single dataset, which is considered the final design of the panel.

[0084] Therefore, in a first aspect of the invention, we herein provide an in vitro or computer implemented method for providing a collection of highly optimized sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions, providing full phylogenetic coverage and higher sensitivity and precision than current methods based on shotgun and amplicon based sequencing, said method comprising the steps of:

[0085] - Step 1 . Gene family selection:

[0086] Compiling a list of marker gene families or functional orthologous groups (OG) by, preferably, identifying genes from a database within or associated to a specified group of molecular interactions and reaction networks.

[0087] Thus, for this first step, we shall depart from the already known sequences associated with each of the functional descriptions considered as a functional marker.

[0088] - Step 2 Data mining to maximize phylogenetic coverage:

[0089] Identifying and providing from one or more databases, nucleotide sequences having a percentage of identity with the genes and OGs identified in step 1 of at least 60%, 70%, 80%, 90%, 95%, 99% or more. Preferably, for each of the selected OGs of step 1 , retrieving all available sequences from (i) global genomic repositories and (ii) public metagenomic datasets. By mining both genomic and metagenomic variability, we aim at ensuring that subsequent enrichment experiments are not biased towards data from cultivated taxa or specific lineages. The resulting corpus of genomic and metagenomics sequences matching each of the targeted OG represents all known gene sequence variability across all cultivated and uncultivated microbial organisms. As already indicated due to the volume and extent of such corpus of information, it is currently unfeasible to target all such variability at once in a single sequencing experiment

[0090] Step 3. Design of capture-kits:

[0091] Analyzing the data corpus produced in Step 2 in order to reduce its dimensionality to make it accessible for a single targeted sequencing experiment, while maintaining the same phylogenetic coverage and sensibility and improving its precision. Preferably, by carrying out the following steps:

[0092] Substep 1 : For each targeted OG (tOG) included in a kit:

[0093] A. Decomposing into non-overlapping chunks of from 10 - 200 pair bases, preferably from 50 to 100 pair bases, more preferably about 80 pair bases, all or part of the nucleotide sequences comprising the dataset obtained in step 2 for each of the tOG, herein referred to as probes. B. Aligning the nucleotide sequences comprising the dataset obtained in step 2 into a multiple sequence alignment (MSA), where the conservation level of each column is calculated.

[0094] C. The multiple sequence alignment produced is used to generate a phylogenetic tree of all or part of the non-redundant sequences.

[0095] D. All the generated probes in substep A are compared in an all-against-all fashion to calculate their pairwise sequence similarity, computed as the % of nucleotides identities observed between two probes and provide an all-against- all probe comparison matrix.

[0096] Substep 2:

[0097] E. Using the all-against-all probe comparison matrix of step D to generate clusters of probes at the NSI threshold. Here, we intend to produce a reduced set of probes compatible with the current sensitivity of capture protocols, which varies from 100% to 30% NSI. This is possible because, for in-solution hybridization capture kits, capture probes can be used in the same way as degenerate primers can, albeit providing capture ranges of up to 30% divergence from the original template

[0098] F. Mapping all or part of the probe clusters of step E back to the same original genomic and metagenomic databases used in step 2 using BLASTN searches.

[0099] G. Analyzing to compute the ratio of positive matches (i.e. hits against a gene belonging to the expected OG) over negative matches (i.e. hits against a gene belonging to an unwanted OG) of the hits produced in step F. Notice that, although all probes were generated from targeted OG genes, many might belong to promiscuous gene regions that exist in many other OGs, potentially dismissing the precision of the targeted OG detections.

[0100] H. Discarding all probe clusters with a ratio of positive matches lower than 1 .0 (i.e. probes producing at least one mismatch) from the set. This step thus provides keeping only highly specific probes of the tOG and reducing potential off-target hybridization in the kit.

[0101] I. Mapping all or part of the remaining selected probe clusters to the MSA generated in substep B), preferably by recording the exact position of matching residues of the original probe on the MSA. J. Discarding probes matching columns of the alignment produced in B with a residue conservation score, preferably of below 0.75. This step provides keeping only specific probes that are also largely universal for the tOGs (i.e. discarding marginal probes).

[0102] K. Mapping back the final set of selected probes to the original sequences selected in Step 2, preferably by using BLASTN, and preferably calculating: number of original sequences recovered by the probes (RecovT), the average length of the original sequence covered by the probe-selection (RecovLen) and the phylogenetic coveraged based on the the generated in substep C (PhyloC)

[0103] In a preferred embodiment, the total number of probes selected (Nprobes), together with their RecovT, RecovL and PhyloC values obtained for each clustering experiment under each NS I are automatically evaluated to select the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing NProbes. Finally, the best probe set for each tOG is combined into a single dataset, thus providing the design of the panel.

[0104] In one embodiment of the first aspect of the invention, the databases of step 2 (Data mining) comprise genomic and metagenomic databases.

[0105] A second aspect of the invention refers to a set of oligonucleotides or probes identified or produced in accordance with the first aspect of the invention, or in accordance with any of the preferred embodiments of the first aspect of the invention. Preferably, the oligonucleotides or probes are DNA or RNA oligonucleotides or probes.

[0106] A third aspect of the invention refers to the in vitro use of a set of oligonucleotides or probes identified or produced in accordance with the first aspect of the invention, or in accordance with any of the preferred embodiments of the first aspect of the invention, in capture and target enrichment assays to identify known or unknown organisms. Kits comprising the above mentioned set of oligonucleotides or probes identified or produced in accordance with the first aspect of the invention also form part of the present invention.

[0107] The combined technical effect of this workflow arises from the integration of all the above elements — metagenomic data inclusion, multi-threshold clustering, off-target filtering, conservation-based exclusion, and quantitative optimization, resulting in a unique capacity to achieve full phylogenetic coverage, strain-level taxonomic resolution, and efficient enrichment of low-abundance taxa in a single sequencing experiment. These combined effects are not achievable by any existing profiling or probe-design system, and represent the main inventive contribution and value of the present patent.

[0108] The panels and methods described herein are designed for use in the capture and target enrichment of environmental or biological DNA samples prior to shotgun sequencing, with the objective of obtaining a comprehensive taxonomic and functional profile of microbial communities with full phylogenetic coverage.

[0109] Specifically, the use of the panel according to step e) of the method enables direct enrichment assays on complex metagenomic samples, including soil, water, sediment, human gut, plant-associated, and extreme environments, before sequencing. The method provides simultaneous detection of high- and low-abundance taxa with strainlevel resolution, and allows for the discovery and characterization of previously undescribed or unknown microbial species and strains not present in existing reference databases. The combined use of the produced panels with standard taxonomic profiling methods (e.g., Kraken2, MetaPhlAn, or mOTUs) significantly improves the number of microbial species and strains identified compared to untargeted shotgun sequencing methods, while maintaining compatibility with established bioinformatics pipelines. The combined use of the produced panels with standard taxonomic profiling does not significantly modify relative abundance estimations of the identified microorganisms, thereby enabling quantitative metagenomic studies without introducing capture-related bias.

[0110] In preferred embodiments, the panels can be used to selectively enrich a defined pool of universally conserved marker genes (e.g., ribosomal proteins, RNA polymerase subunits, gyrases, or elongation factors) or specific functional pathways (e.g., nitrogen fixation, methane metabolism, antibiotic resistance). This enables fine-resolution profiling of microbial communities and functional gene distributions with a depth unattainable by traditional metagenomic sequencing approaches.

[0111] Furthermore, the method is compatible with low-input DNA samples and degraded environmental DNA (eDNA), which broadens its applicability to clinical or host- contaminated samples.

[0112] The following examples serve to illustrate the present invention but do not limit the same. EXAMPLES

[0113] Opti mized target capture sequencing of ten maker genes provide full phylogenetic coverage and strain resolution in a single experiment:

[0114] Using the methodology described in previous steps, we designed a custom set of DNA oligonucleotide probes targeting ten universally-conserved marker genes (MGs) suitable for high resolution phylogenetic analysis and taxonomic profiling of metagenomics samples and commonly employed by bioinformatic software (mOTUs, https: / / www.nature.com / articles / s41467-019-08844-4). Each of the ten genes selected represent a targeted orthologous groups (tOG) in substeps A-K. Probes were generated from the coding nucleotide sequences of the ten marker genes extracted from reference genomic and metagenomic databases, comprising a total of 7,726 prokaryotic genomes and metagenomic assemble genomes and providing coverage for nearly full bacterial and archeal biodiversity.

[0115] From the total number of probes obtained in step A (~3 million), the optimization substeps generated ~500,000 total selected probes distributed along the ten tOGs. Next, we manufactured a custom hybridization capture kit using the 500k selected probes and performed enrichment experiments on various environmental samples. In particular, we analyzed the sensitivity, phylogenetic coverage and precision of the generated kit on 57 metagenomic samples from human gut (32 samples), soil (16 samples), and microbial mats (9 samples), for which the standard shotgun metagenomics and 16 amplicon sequencing approaches were also performed for comparison reasons.

[0116] Capture experiments were successful for all samples, achieving an on-target enrichment rate (i.e. total reads matching probes) from 70% (human gut) to 20% (other natural environments).

[0117] High sensitivity and phylogenetic coverage: Accurate detection of low abundance species

[0118] To test the sensitivity of produced capture kit design, we estimated the taxonomic profile of all samples using the mOTUs tool, which is considered an standard method in the field and is optimized to obtain the presence and relative abundance of microorganisms in a metagenomic samples using the same ten marker genes used in our capture panel. On average, our target capture resulted in the detection of up to 14-fold more mOTUs species per sequencing depth unit than regular shotgun across all samples (Figure 1 B). This increase was consistent regardless of the stringency level used for species calls in the mOTUs tool. Most species detected in the shotgun data were also detected in the capture data (% average match), and their relative abundances were strongly correlated, demonstrating that the capture method increased detection without biasing abundance estimates (Figure 1 ). In particular, capture detections were significantly enriched in the low abundance range (Figure 1 histograms), highlighting a high level of sensitivity currently unreachable by either shotgun or amplicon methods. The number of mOTUs-based species detections derived from target capture data increased not only compared to mOTUs shotgun results, but also compared to the number of species detected by other standard tools such as MetaPhlAn4 and and Kraken2.

[0119] High precision: Strain resolution level in the identification of microorganisms in the samples

[0120] Levering on the high sequencing depth obtained for the targeted makers genes, we evaluated whether fine-grained taxonomic profiling based on capture-based data could achieve strain level resolution in the identification of microorganisms, which remains largely unfeasible with current shotgun and amplicon based approaches. For this, we used the bioinformatic software metaSNV to call for unique single nucleotide polymorphisms (SNP) in the ten marker genes that identified different strains in a sample.

[0121] As shown in table 1 below, the number of unique strains identified based on SNP detections under strict thresholds (i.e. SNPs supported by at least five reads and across various samples) were ten times larger using the capture-based targeted sequencing approach described in this document. Table 1

[0122] DNA capture and sequencing

[0123] DNA from the different experiments was mechanically fragmented with Covaris (Covaris, Woburn, Massachusetts, USA) to an average fragment size of 300pb. Library preparation of the samples was done according to NEBNext Ultra II DNA Library Prep Kit with manufacturer's instructions. The samples were equimollary pooled, and captured using the Nimblegen baits, following the SeqCap EZ HyperCap Workflow v2.3. Illumina sequencing was performed afterwards, obtaining paired-end reads of 2x150 nucleotides. Shotgun data was already available for human gut samples, for extreme environment experiments and for soil samples.

[0124] Reads preprocessing

[0125] Obtained Illumina reads were inspected with FastQC (vO.11.9) and MultiQC (v1 .10.1 ). Duplicates were removed using BBmap’s (v38.91 -1 ) clumpify.sh (dedupe subs=0). Next, deduplicated reads were trimmed using fastp (v0.20.1 ; - disable_quality_filtering — disable_length_filtering --correction --trim_poly_g --overrepresentation_analysis) and Trimmomatic (0.38;

[0126] ILLUMINACLIP:$EBROOTTRIMMOMATIC / adapters / TruSeq3-PE.fa:2:30:10 MAXINFO:40:0.2 MINLEN:50). From the sets of clean reads we created standardized sets of read pairs of length 100, using Trimmomatic (v0.38; CROP:100 MINLEN:100). These were randomly subsampled to 1 M read pairs for all samples, and also to the minimum number of reads available for each capture-shotgun sample pair.

[0127] Taxonomic classification of reads. We ran mOTUs version 3.0.1 with both -g1 and -g3, with the sets of standardized reads, with both subsampling to 1 M read pairs and to the minimum of capture / shotgun sample pairs. For the latter we also ran mOTUs with - g5, -g7, -g9 and -g10. We also ran motus for all the available read pairs, both capture and shotgun data, with both -g 1 and -g 3, producing both absolute abundances from raw counts and relative abundances from scaled counts.

[0128] Saturation curves were built with mOTUs results (default parameters) for the samples with the largest number of standardized 2x100 read pairs for each environment (human gut: 16s001897, soil: NoOTCNoCrust3, extreme environments: LLA9C1 ), to compare capture and shotgun samples. Subsampling was performed in steps of 1 M read pairs, reducing the number of replicates with the number of reads to be subsampled. For the human gut data, subsamples ranged 1 M (5 replicates) to 7M read pairs (1 replicate). For soil, subsamples ranged from 1 M (5 replicates) to 18M read pairs (1 replicate). For extreme environments, subsamples ranged 1 M (5 replicates) to 13M read pairs (1 replicate).

[0129] We also used Kraken2. Some technical comments: most sequences are assigned with very low confidence thresholds (0.0, 0.02). No clear differences observed (which is somewhat surprising) analyzing the same number of reads from capture and shotgun data. In both cases, the kraken2 confidence parameter has dramatic impact in the number of species detected and % of reads mapped to all the taxonomic ranges, which are drastically decreased above confidence 0.2 (~1 / 5 of the k-mers of the read identified), specially for ext_env and soil. For human_gut, the number of species detected is much more affected than the % of reads mapped, by the confidence threshold. Overall, the diversity detected using Kraken2 is comparable, capture vs shotgun, but capture data is enriched in genes (COG MGs) which can be directly compared, in contrast with shotgun which will show a mixture of sequences from diverse genomic loci which will not directly compare.

[0130] Also, as it happens with motus, the kraken2 database seems biased towards human_gut diversity. We also ran MetaPhlAn version 4 with default parameters and unclassified fraction estimation.

[0131] 16S sequencing and analysis

[0132] We obtained sequencing reads for 50 samples out of the 57 used for capture and shotgun data. PCR primers were removed with Cutadapt. Following dada2’s pipeline, reads were trimmed, denoised, merged and subjected to chimera detection. Sample inference and ASV detection was carried out with dada2 default procedures. Taxonomic classification of ASVs was carried out with DECIPHER, both with the SILVA_SSU_r138_2019 and the GTDB_r207-mod_April2022 databases. We also used a GTDB custom database to classify the ASVs, as well as clustered ASVs to 97% identity, resembling a rather standard threshold for OTUs delineation.

[0133] Detection of on-target reads Clean reads were mapped to the original marker genes using Diamond (v2.0.11 .149; blastx mode --ultra-sensitive --iterate -e 0.001 --matrix BLOSUM45 --top 3), and . Each read pair was kept, based on mOTUs BWA and Diamond hits, if at least one read of the pair maps to a marker gene, unless the other read maps to a different marker gene. The remaining read pairs were mapped to eggNOG 5 with eggNOG-mapper 2.1.6 (-m diamond --sensmode ultra-sensitive -dmnd terate yes -matrix BLOSUM45). If at least one read of each pair mapped to an OG different from the marker genes, the read pair was discarded. The remaining read pairs were classified as on-target.

[0134] CLAUSES

[0135] 1 . An in vitro method for providing a collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions with full phylogenetic coverage and high sensitivity, said method comprising the steps of: a. Step 1 . Gene family selection: Compiling a list of marker gene families or functional orthologous groups (OG) by, preferably, identifying genes from a database within or associated to a specified group of molecular functions; b. Step 2 Data mining: Identifying and providing from one or more databases, a data corpus or dataset of nucleotide sequences having a percentage of identity with the genes and OGs identified in step 1 of at least 60%, 70%, 80%, 90%, 95%, 99% or more; c. Step 3. Design of capture-kits: Analyzing the data corpus or dataset provided in Step 2 in order to reduce its dimensionality to make it accessible for a single targeted sequencing experiment, preferably while maintaining the same phylogenetic coverage and sensibility; d. Providing the design of the panel or collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families or functions selected in Step 1 ; e. Optionally, manufacturing or producing the panel of step d.

[0136] 2. The method according to clause 1 , wherein in Step 2 the data corpus of nucleotide sequences is provided by retrieving for each of the selected OGs of step 1 , the sequences from (i) global genomic repositories and (ii) public metagenomic datasets.

[0137] 3. The method according to clause 2, wherein Step 3 is carried out by a method comprising the following steps i) to iv) for each targeted OG (tOG) included in a kit: i. Decomposing into non-overlapping chunks of from 10 - 200 pair bases, preferably from 50 to 100 pair bases, more preferably about 80 pair bases, all or part of the nucleotide sequences comprising the dataset obtained in Step 2 for each of the tOG, herein referred to as probes. ii. Aligning the nucleotide sequences comprising the dataset obtained in step 2 into a multiple sequence alignment (MSA), where the conservation level of each column is calculated. iii. Optionally, the multiple sequence alignment produced is used to generate a phylogenetic tree of all or part of the non-redundant sequences. iv. All or part of the generated probes in step i) are compared in an all-against-all fashion to calculate their pairwise sequence similarity, computed as the % of nucleotides identities observed between two probes and provide an all-against-all probe comparison matrix. wherein the method further comprises the following steps: v. Using the all-against-all probe comparison matrix of step iv to generate clusters of probes at the NSI threshold, wherein the NSI threshold is selected from one or more values in the range of from 30% to 100%. vi. Mapping all or part of the probe clusters of step v back to the same original genomic and metagenomic databases used in Step 2 using BLASTN searches. vii. Analyzing to compute the ratio of positive matches (i.e. hits against a gene belonging to the expected OG) over negative matches (i.e. hits against a gene belonging to an unwanted OG) of the hits produced in step vi. viii. Discarding all probe clusters with a ratio of positive matches lower than 1.0 (i.e. probes producing at least one mismatch) from the set. ix. Mapping all or part of the remaining selected probe clusters to the MSA generated in step ii), preferably by recording the exact position of matching residues of the original probe on the MSA. x. Discarding probes matching columns of the alignment produced in ii with a residue conservation score, preferably of below 0.75. xi. Mapping back the final set of selected probes to the original sequences selected in Step 2, preferably by using BLASTN, and preferably calculating: number of original sequences recovered by the probes (RecovT), the average length of the original sequence covered by the probe-selection (RecovLen) and the phylogenetic coveraged based on the the generated in substep C (PhyloC) The method according to clause 3, wherein for each tOG the total number of probes selected (Nprobes), together with their RecovT, RecovL and PhyloC values obtained for each clustering experiment under each NSI are automatically evaluated to select the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing NProbes. The method according to clause 4, wherein the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing Nprobes, for each tOG, is combined into a single dataset, and providing the design of the panel. The method according to clause 5, wherein the method further comprises the manufacturing or production of the panel of probes. A panel obtained or manufactured or produced by the method of clause 6, wherein the probes are polynucleotides, preferably DNA or RNA, probes. Use of the panel according to clause 6, for the capture and target enrichment assays directly on environmental DNA samples, prior to shotgun sequencing. A kit comprising the panel obtained or manufactured or produced by the method of clause 6

Claims

CLAIMS1. A computer implemented method for providing a collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families, said method comprising the steps of: a. Step 1 . Gene family selection: Compiling a list of functional orthologous groups (OG) by, preferably, identifying genes from a database within or associated to a specified group of molecular functions; b. Step 2 Data mining: Identifying and providing from one or more databases, a data corpus or dataset of nucleotide sequences having a percentage of identity with the genes and OGs identified in step 1 of at least 60%, 70%, 80%, 90%, 95%, 99% or more; c. Step 3. Design of capture-kits: Analyzing the data corpus or dataset provided in Step 2 in order to reduce its dimensionality to make it accessible for a single targeted sequencing experiment, preferably while maintaining the same phylogenetic coverage and sensibility; d. Providing the design of the panel or collection of sequencing probes that enable targeting, preferably in a single sequencing experiment, one or various gene families selected in Step 1 and 2; and wherein Step 3 is carried out by a method comprising the following steps i) to iv) for each targeted OG (tOG): i. Decomposing into non-overlapping chunks of from 10 - 200 pair bases, preferably from 50 to 100 pair bases, more preferably about 80 pair bases, all or part of the nucleotide sequences comprising the dataset obtained in Step 2 for each of the tOG, herein referred to as probes; ii. Aligning the nucleotide sequences comprising the dataset obtained in step 2 into a multiple sequence alignment (MSA), where the conservation level of each column is calculated; iii. Optionally the multiple sequence alignment produced is used to generate a phylogenetic tree of all or part of the non-redundant sequences;iv. All or part of the generated probes in step i) are compared in an all-against-all fashion to calculate their pairwise sequence similarity, computed as the % of nucleotides identities observed between two probes and provide an all-against-all probe comparison matrix; wherein step 3 of the method further comprises the following steps: v. Using the all-against-all probe comparison matrix of step iv to generate clusters of probes at the NSI (nucleotide sequence identity) threshold, wherein the NSI threshold is selected from one or more values in the range of from 30% to 100%; vi. Mapping all or part of the probe clusters of step v back to the same original genomic and metagenomic databases used in Step 2 using BLASTN searches. vii. Analyzing to compute the ratio of positive matches, hits against a gene belonging to the expected OG, over negative matches, hits against a gene belonging to an unwanted OG, of the hits produced in step vi. viii. Discarding all probe clusters with a ratio of positive matches lower than 1.0, probes producing at least one mismatch, from the set. ix. Mapping all or part of the remaining selected probe clusters to the MSA generated in step ii), preferably by recording the exact position of matching residues of the original probe on the MSA. x. Discarding probes matching columns of the alignment produced in ii with a residue conservation score below 0.

75. xi. Mapping back the final set of selected probes to the original sequences selected in Step 2, preferably by using BLASTN, and preferably calculating: number of original sequences recovered by the probes (RecovT), the average length of the original sequence covered by the probe-selection (RecovLen) and the phylogenetic coveraged based on the generated in substep C (PhyloC).

2. The method according to claim 1 , wherein in Step 2 the data corpus of nucleotide sequences is provided by retrieving for each of the selected OGs of step 1 , the sequences from (i) global genomic repositories and / or (ii) public metagenomic datasets.

3. The method according to claim 1 or 2, wherein for each tOG the total number of probes selected (Nprobes), together with their RecovT, RecovL and PhyloC values obtained for each clustering experiment under each NSI are automatically evaluated to select the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing NProbes.

4. The method according to claim 3, wherein the set of probes that maximizes RecovT, RecovL and PhyloC, while minimizing Nprobes, for each tOG, is combined into a single dataset, and providing the design of the panel.

5. The method according to claim 4, wherein the method further comprises designing the panel of probes.