Method for predicting and increasing agricultural yield and bacterial compositions for improving the soil

Metagenomic methods and bacterial compositions based on 16S rRNA sequencing predict agricultural output, addressing the challenges of yield prediction and soil fertility optimization, achieving significant variance explanation and yield enhancement.

WO2026115029A1PCT designated stage Publication Date: 2026-06-04SOILYTIX GMBH

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SOILYTIX GMBH
Filing Date
2025-11-27
Publication Date
2026-06-04

Smart Images

  • Figure IMGF000014_0001
    Figure IMGF000014_0001
  • Figure IMGF000014_0002
    Figure IMGF000014_0002
  • Figure IMGF000014_0003
    Figure IMGF000014_0003
Patent Text Reader

Abstract

The present invention relates to a method for predicting agricultural output. In addition, the present invention relates to a composition for increasing agricultural output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Our Ref.: 1173-4 PCT Soilytix GmbH

[0002] METHOD FOR PREDICTING AGRICULTURAL OUTPUT

[0003] The present invention relates to a method for predicting agricultural output. In addition, the present invention relates to a composition for increasing agricultural output.

[0004] BACKGROUND OF THE INVENTION

[0005] It is assumed that feeding the world's population will face problems in the coming decades due to at least two factors: a growing world population on the one hand and the negative effects of climate change on soil fertility on the other. Therefore, it is elementary to optimize yield on agricultural soils through intelligent management without compromising soil fertility. Soil fertility can be viewed as a function of both physical and chemical, as well as biological parameters.

[0006] In the past, the study of bacteria from a specific habitat presupposed the cultivation in a laboratory. It was then possible to analyze the grown bacterial cultures more precisely and identify them. With the development of molecular genetic methods, these laborious procedures are no longer necessary. Today, they are largely replaced by so-called metagenomic methods. With the appearance of high-throughput sequencing technologies, it is possible to determine a significant portion of (micro)biological parameters with a single measurement.

[0007] Therefore, the development of microbiome-based yield prediction would be a promising method to cost-effectively and less laborious optimize yield through targeted field and fruit selection.

[0008] The present inventors have identified microbiome signatures using metagenomic methods which allow the prediction of agricultural output in a cost-effective and less laborious way.

[0009] Specifically, the present inventors have extracted bacterial DNA directly from soil samples taken from different agricultural areas. In these different agricultural areas, the inventors also recorded the maize com yield. The soil samples taken from the different agricultural areas were subsequently sequenced. From the resulting DNA pattern, the present inventors could read which bacteria, particularly amplicon sequence variants (AS Vs) which can be assigned to bacteria, were contained in the soil sample and associated them with the maize corn yield.

[0010] The gene of choice for the metagenomic studies of the present inventors is the so-called 16S rRNA gene, as this gene is ubiquitous in bacteria. It contains some conserved DNA regions that are highly similar in all bacterial taxa, and also contains variable DNA sequences that have been significantly modified in the course of evolution, making it possible to identify the bacteria by sequence diversity. Particularly, the present inventors first amplified a part of the 16S rRNA gene from the DNA isolated from soil samples. Then, the resulting DNA libraries were sequenced by nextgeneration sequencing. The 16S rRNA gene sequences, specifically the amplicon sequence variants (AS Vs) identified therein, thus, obtained from a sample, were then compared with a database to determine the taxonomic affiliation of the bacteria and their abundance in the soil sample.

[0011] The LASSO model created by the present inventors in this process was able to describe 58% (16S V4) of the yield variance in an exemplary maize corn field. Significant correlations between predicted and observed yields were also obtained for external / published datasets, by utilizing 16S rRNA gene abundances.

[0012] The present inventors could, thus, show that the determined microbiome signatures allow prediction of agricultural output.

[0013] SUMMARY OF THE INVENTION

[0014] In a first aspect, the present invention relates to a method for predicting agricultural output comprising the steps of

[0015] (i) determining the relative percentage of at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the genus Sphingomonas, of the family A21b, of the order Rokubacteriales, of the genus Nordella, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, of the order Vicinamibacterales, of the family Vicinamibacteraceae, of the genus RB41, of the family Methyloligellaceae, and of the genus Nocardioides in a soil sample of an agricultural or a prospective agricultural area, and

[0016] (ii) predicting agricultural output on the basis of the relative percentage of the at least one bacterial taxon determined in step (i).

[0017] In a preferred embodiment, the method comprises determining the relative percentage of at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium, of the genus Bacillus, of the genus Pseudonocardia, of the genus Paenisporosarcina, of the genus Ellin6055, of the genus Rummeliibacillus, of the genus Reyranella, of the genus Rhodoplanes, and of the genus Streptomyces in the soil sample of an agricultural or a prospective agricultural area.

[0018] In a second aspect, the present invention relates to a composition comprising at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the family A21b, of the order Rokubacteriales, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, and of the family Methyloligellaceae .

[0019] In a preferred embodiment, the composition comprises at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium and a bacterial taxon of the genus Reyranella.

[0020] In a third aspect, the present invention relates to the use of the composition of the second aspect to increase agricultural output.

[0021] In a fourth aspect, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0022] (i) carrying out the method of the first aspect(, thereby obtaining a (composite) score having a value close to 0 which is indicative for a predicted low agricultural output),

[0023] (ii) applying the composition of the second aspect onto the (prospective) agricultural area, and

[0024] (iii) planting or cultivating agricultural plants onto the agricultural area.

[0025] In a fifth aspect, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0026] (i) applying the composition of the second aspect onto an (a prospective) agricultural area, and

[0027] (ii) planting or cultivating agricultural plants onto the agricultural area.

[0028] In a sixth aspect, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0029] (i) applying the composition of the second aspect onto agricultural plant seeds, thereby coating the agricultural plant seeds with the composition,

[0030] (ii) sowing the coated agricultural plant seeds into the (prospective) agricultural area, and

[0031] (iii) cultivating agricultural plants derived therefrom onto the agricultural area.

[0032] This summary of the invention does not necessarily describe all features of the present invention. Other embodiments will become apparent from a review of the ensuing detailed description.

[0033] DETAILED DESCRIPTION OF THE INVENTION

[0034] Definitions

[0035] Before the present invention is described in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.

[0036] Preferably, the terms used herein are defined as described in “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, Leuenberger, H.G.W, Nagel, B. and Kolbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland).

[0037] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, GenBank Accession Number sequence submissions etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. In the event of a conflict between the definitions or teachings of such incorporated references and definitions or teachings recited in the present specification, the text of the present specification takes precedence.

[0038] The term “comprise” or variations such as “comprises” or “comprising” according to the present invention means the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers. The term “consisting essentially of’ according to the present invention means the inclusion of a stated integer or group of integers, while excluding modifications or other integers which would materially affect or alter the stated integer. The term “consisting of’ or variations such as “consists of’ according to the present invention means the inclusion of a stated integer or group of integers and the exclusion of any other integer or group of integers.

[0039] The terms “a” and “an” and “the” and similar reference used in the context of describing the invention (especially in the context of the claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.

[0040] As used herein, the term “about” indicates a certain variation from the quantitative value it precedes. In particular, the term “about” allows a ±5% variation from the quantitative value it precedes, unless otherwise indicated or inferred. The use of the term “about” also includes the specific quantitative value itself, unless explicitly stated otherwise. For example, the expression “about 80°C” allows a variation of ±4°C, thus referring to range from 76°C to 84°C.

[0041] The present inventors have identified microbiome signatures which allow the prediction of agricultural output. Thus, the present invention relates to a method for predicting agricultural output.

[0042] The term “agriculture”, as used herein, encompasses crop production. The term “agricultural output”, as used herein, refers to an amount, a quantity, or a proportion of crops obtained from an agricultural area. While inputs in agriculture are seeds, fertilizers, machinery, labor, etc., the outputs of the farming activity are the harvested crops. Thus, agricultural output is a direct measure of the total amount, quantity, or proportion of crops produced.

[0043] The term “low agricultural output”, as used herein, refers to an output which is below the usually expected output and / or average of the last five cultivation periods in relation to the cultivated agricultural land under usual climate conditions in the respective geographical area. In one preferred embodiment, the agricultural output of maize or green fodder is predicted.

[0044] In this regard, a predicted low crop yield of maize is preferably indicative for a crop yield size of < 20 tons / hectare.

[0045] In this regard, a predicted low green fodder yield is preferably indicative for a low green fodder yield size of the respective crop.

[0046] The term “high agricultural output”, as used herein, refers to an output which is above the usually expected output and / or average of the last five cultivation periods in relation to the cultivated agricultural land under usual climate conditions in the respective geographical. In one preferred embodiment, the agricultural output of maize or green fodder is predicted.

[0047] In this regard, a predicted high crop yield of maize is preferably indicative for a crop yield size of > 35 tons / hectare.

[0048] In this regard, a predicted high green fodder yield is preferably indicative for a high green fodder yield size of the respective crop.

[0049] The term “crop yield”, as used herein, refers to the harvested production per unit of harvested area for crop products.

[0050] Crop production depends on the availability of arable land and is affected in particular by yields, macroeconomic uncertainty, as well as consumption patterns; it also has a great incidence on agricultural commodities' prices. The importance of crop production is related to harvested areas, returns per hectare (yields) and quantities produced. In most of the cases yield data are not recorded, but are obtained by dividing the production data by the data on area harvested. The actual yield that is captured on farm depends on several factors such as the crop's genetic potential, the amount of sunlight, water and nutrients absorbed by the crop, the presence of weeds and pests. This indicator is presented for wheat, maize, rice and soybean. Crop production is measured in tonnes per hectare, in thousand hectares and thousand tonnes.

[0051] The term “agricultural productivity”, as used herein, is the ratio of agricultural inputs to outputs. The greater the agricultural output (for a given input), the higher the agricultural productivity of a farm or farm land. In simple terms, agricultural productivity is: output input = productivity.

[0052] While individual products are usually measured by weight, which is known as crop yield, varying products make measuring overall agricultural output difficult. Therefore, agricultural productivity is usually measured as the market value of the final output. This productivity can be compared to many different types of inputs such as labour or land. Such comparisons are called partial measures of productivity.

[0053] Agricultural productivity may also be measured by what is termed total factor productivity (TFP). This method of calculating agricultural productivity compares an index of agricultural inputs to an index of outputs. This measure of agricultural productivity was established to remedy the shortcomings of the partial measures of productivity; notably that it is often hard to identify the factors cause them to change. Changes in TFP are usually attributed to technological improvements.

[0054] Agricultural productivity is an important component of food security. Increasing agricultural productivity through sustainable practices can be an important way to decrease the amount of land needed for farming and slow environmental degradation as well as climate change.

[0055] The term “agricultural area”, as used herein, refers to land devoted to / used for agriculture. It is land used for the production of crops. It is generally synonymous with both farmland or cropland. It is often expressed in hectares (ha). The term “agricultural area”, as used herein, also refers to an agricultural region unit of measurement used in statistics and management, especially for production indicators such as yields. The term “agricultural area” includes, but is not limited to, cropland, specifically for the production of crops, and grass land, specifically for the production of green fodder.

[0056] The term “prospective agricultural area”, as used herein, refers to land that may be devoted to / used for agriculture, e.g. if diverse preconditions such as microbiome status / composition are met.

[0057] The term “agricultural plant”, as used herein, refers to any plant, or part thereof, grown, maintained, or otherwise produced for commercial purposes, including growing, maintaining or otherwise producing plants for sale or trade, for research or experimental purposes, or for use in part or their entirety in another location. In a preferred embodiment, the agricultural plant is maize or grass (e.g. used as green fodder).

[0058] The term “plant seed”, as used herein, refers to an embryonic plant enclosed in a protective outer covering. The formation of the seed is part of the process of reproduction in seed plants, the spermatophytes. The term “coating”, as used herein, refers to a covering that is applied to the plant or plant seed, in particular to the surface of the plant or plant seed, to be coated. The coating itself may be an all-over coating, completely covering the plant or plant seed, or it may only cover parts of the plant or plant seed.

[0059] The term “microbiome”, as used herein, refers to a community of microorganisms that can usually be found living together in a given habitat, e.g. in an agricultural area.

[0060] The term “soil sample”, as used herein, refers to a sample collected in a representative location of an (a prospective) agricultural area. A soil sample may be taken using ASTM El 727, “Standard Practice for Field Collection of Soil Samples for Lead Determination by Atomic Spectrometry Techniques”, or equivalent method.

[0061] The term “soil sampling”, as used herein, refers to a process of extracting a small volume of soil for subsequent analysis at a lab. Soil testing is an essential component of soil resource management. Each sample collected must be a true representative of the area being sampled. Utility of the results obtained from the laboratory analysis depends on the sampling precision. Hence, collection of large number of samples is advisable so that sample of desired size can be obtained by sub-sampling. In general, sampling is done at the rate of one sample for every two- hectare area. However, at-least one sample should be collected for a maximum area of five hectares. For soil survey work, samples are collected from a soil profile representative to the soil of the surrounding area.

[0062] In the past, the study of bacteria from a specific habitat presupposed the cultivation in a laboratory. It was then possible to analyze the grown bacterial cultures more precisely and identify them. With the development of molecular genetic methods, these laborious procedures are no longer necessary. Today, they are largely replaced by so-called metagenomic methods. With the appearance of high-throughput sequencing technologies, it is possible to determine a significant portion of (micro)biological or microbiome parameters with a single measurement.

[0063] The term “metagenomics”, as used herein, refers to a technique by which the genetic material (e.g. DNA) is extracted directly from samples taken from the environment (e.g. soil samples). The genetic material is, after multiplication / amplification, sequenced, e.g. in an approach called next generation sequencing or shotgun sequencing.

[0064] Today, bacterial DNA is extracted directly from a sample taken from its normal habitat. This sample is sequenced. From the resulting DNA pattern, researchers can read which bacteria are contained in the sample. The gene of choice for the metagenomic studies described herein is the so-called 16S ribosomal RNA (16S rRNA) gene.

[0065] The term “16S ribosomal RNA (16S rRNA)”, as used herein, is the RNA component of the 30S subunit of prokaryotic ribosomes. The 16S rRNA gene is found in all bacteria. It contains some conserved DNA regions that are the highly similar in all organisms. The gene can be recognized by them. In addition, the 16S rRNA gene also contains variable DNA sequences that have been significantly modified in the course of evolution, making it possible to identify the bacteria. In general, at least a part of the 16S rRNA gene is first amplified from a sample. Then, either the pool of sequences, thus, propagated is cloned and the sequence of individual clones is determined, or state-of-the-art sequencing equipment is used for sequencing. The 16S rRNA sequences, thus, obtained from a sample can then be compared with a database to determine the taxonomic affiliation of the bacteria and their abundance in the sample.

[0066] As mentioned above, the present inventors first amplified a part of the 16S rRNA gene from the DNA isolated from soil samples. Then, the resulting DNA libraries were sequenced by next-generation sequencing. The 16S rRNA gene sequences, specifically the amplicon sequence variants (AS Vs) identified therein, thus, obtained from a sample, were then compared with a database to determine the taxonomic affiliation of the bacteria and their abundance in the soil sample.

[0067] The term “amplicon sequence variants (AS Vs)”, as used herein, refers to highly resolved DNA sequences obtained from microbial communities during amplicon-based studies, such as 16S ribosomal RNA (rRNA) gene sequencing. AS Vs represent unique sequences of a specific amplicon region, typically defined by exact sequence differences rather than operational taxonomic units (OTUs), which rely on clustering sequences at a specified similarity threshold (e.g. 97%).

[0068] The key characteristics of ASVs are:

[0069] ASVs distinguish individual sequence variants down to single nucleotide differences, offering finer resolution than OTUs. They have, thus, a high resolution.

[0070] They are defined as unique, error-corrected, exact sequences, rather than groups of sequences with approximate similarity.

[0071] Advanced algorithms, such as DADA2 or UNOISE3, are used to denoise the sequencing data and correct for sequencing errors, ensuring that ASVs represent true biological sequences rather than artefacts.

[0072] ASVs enable direct comparison across datasets because they do not depend on clustering thresholds that can vary across studies. They have, thus, a standardized comparability.

[0073] The exact sequence-based definition ensures that analyses can be replicated, and ASVs are consistently identifiable in future studies. ASVs are, thus, reproducible.

[0074] In view of the above advantages, the ASVs are used in the present bacterial microbiome study to investigate microbial diversity in soil samples. As ASVs can be classified to / assigned to a bacterial taxon, they are particularly used herein to identify specific taxa at high resolution in the soil samples. Specifically, agricultural output can be predicted by determining the relative percentage of at least one amplicon sequence variant (ASV) selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 25 classified to / assigned to a bacterial taxon in a soil sample of an agricultural or a prospective agricultural area.

[0075] In any case, AS Vs offer advantages over traditional OTU-based approaches in both accuracy and interpretability.

[0076] The term “amplifying”, as used herein, refers to any means by which at least a part of a nucleic acid molecule described herein is reproduced, typically in a template-dependent manner, including without limitation, a broad range of techniques for amplifying nucleic acid sequences, either linearly or exponentially. Any of several methods can be used to amplify the nucleic acid molecule. Any in vitro means for multiplying the copies of a target sequence of nucleic acid can be utilized. These include linear, exponential, or other amplification methods.

[0077] Examples of amplification techniques that can be used include, but are not limited to, PCR, quantitative PCR, quantitative fluorescent PCR (QF-PCR), multiplex fluorescent PCR (MF -PCR), real time PCR (RT-PCR), single cell PCR, restriction fragment length polymorphism PCR (PCR- RFLP), hat start PCR, nested PCR, in situ polony PCR, in situ rolling circle amplification (RCA), bridge PCR, picotiter PCR, and emulsion PCR. Other suitable amplification methods include the ligase chain reaction (LCR), transcription amplification, self-sustained sequence replication, selective amplification of target polynucleotide sequences, consensus sequence primed polymerase chain reaction (CP -PCR), arbitrarily primed polymerase chain reaction (AP-PCR), degenerate oligonucleotide-primed PCR (DOP-PCR), and nucleic acid-based sequence amplification (NABSA).

[0078] In various embodiments, nucleotide sequences are sequenced. The term “sequencing”, as used herein, includes any method of determining the sequence of a nucleic acid molecule. Such methods include Maxam-Gilbert sequencing, Chain-termination methods, Shot gun sequencing, PCR sequencing, Bridge PCR, massively parallel signature sequencing (MPSS), Polony sequencing, pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, Single molecule real time (SMRT) sequencing, Nanopore DNA sequencing, sequencing by hybridization, sequencing with mass spectrometry, microfluidic Sanger sequencing, microscopybased techniques, RNAP sequencing, high- throughput sequencing (HTS).

[0079] The term “next generation sequencing (NGS)” as used herein, refers to a new method for sequencing nucleotide sequences at high speed and at low cost. Next-generation sequencing (NGS) is, thus, a high-throughput methodology that enables rapid sequencing of the base pairs in DNA or RNA samples. Supporting a broad range of applications, including gene expression profiling, chromosome counting, detection of epigenetic changes, and molecular analysis, NGS is driving discovery and enabling the future of personalized medicine. NGS is also known as second generation sequencing (SGS) or massively parallel sequencing (MPS).

[0080] The term “sequence identity”, as used herein, refers to a measurement which allows to indicate the similarity of nucleotide and amino acid sequences. The percentage of sequence identity can be determined via sequence alignments. Such alignments can be carried out with several art-known algorithms, preferably with the mathematical algorithm of Karlin and Altschul (Karlin & Altschul (1993) Proc. Natl. Acad. Sci. USA 90: 5873-5877), with hmmalign (HMMER package) or with the CLUSTAL algorithm (Thompson, J. D., Higgins, D. G. & Gibson, T. J. (1994) Nucleic Acids Res. 22, 4673-80) or the CLUSTALW2 algorithm (Larkin MA, Blackshields G, Brown NP, Chenna R, McGettigan PA, McWilliam H, Valentin F, Wallace IM, Wilm A, Lopez R, Thompson JD, Gibson TJ, Higgins DG. (2007). Clustal W and Clustal X version 2.0. Bioinformatics, 23, 2947-2948).

[0081] The grade of sequence identity (sequence matching) may be calculated using e.g. BLAST, BLAT or BlastZ (or BlastX). A similar algorithm is incorporated into the BLASTN and BLASTP programs of Altschul et al. (1990) J. Mol. Biol. 215: 403-410. BLAST protein searches are performed with the BLASTP program available e.g. on the web site: http: / / blast.ncbi. nlm.nih.gov / Blast.cgi?PROGRAM=blastp&BLAST_PROGRAMS=blastp&PA GE_TYPE=BlastSearch&SHOW_DEFAULTS=on&LINK_LOC=blasthome

[0082] Preferred algorithm parameters used are the default parameters as they are set on the indicated web site:

[0083] Expect threshold = 10, word size = 3, max matches in a query range = 0, matrix = BLOSUM62, gap costs = Existence: 11 Extension: 1, compositional adjustments = conditional compositional score matrix adjustment together with the database of non-redundant protein sequences (nr).

[0084] To obtain gapped alignments for comparative purposes, Gapped BLAST is utilized as described in Altschul et al. (1997) Nucleic Acids Res. 25: 3389-3402. When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs are used. Sequence matching analysis may be supplemented by established homology mapping techniques like Shuffle-LAGAN (Brudno M., Bioinformatics 2003b, 19 Suppl 1 :154-I62) or Markov random fields.

[0085] The term “relative percentage”, as used herein, refers to dividing the read counts of a given gene sequence by the summed read counts of all gene sequences and multiplying it by 100.

[0086] Embodiments of the invention

[0087] The present invention will now be further described. In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous, unless clearly indicated to the contrary.

[0088] The present inventors have identified microbiome signatures using metagenomic methods which allow the prediction of agricultural output.

[0089] Specifically, the present inventors have extracted bacterial DNA directly from soil samples taken from different agricultural areas. In these different agricultural areas, the inventors also recorded the maize com yield. The soil samples taken from the different agricultural areas were subsequently sequenced. From the resulting DNA pattern, the present inventors could read which bacteria, particularly amplicon sequence variants (AS Vs) which can be assigned to bacteria, were contained in the soil sample and associated them with the maize corn yield.

[0090] The gene of choice for the metagenomic studies of the present inventors is the so-called 16S rRNA gene, as this gene is ubiquitous in bacteria. It contains some conserved DNA regions that are highly similar in all bacterial taxa, and also contains variable DNA sequences that have been significantly modified in the course of evolution, making it possible to identify the bacteria by sequence diversity.

[0091] Particularly, the present inventors first amplified a part of the 16S rRNA gene from the DNA isolated from soil samples. Then, the resulting DNA libraries were sequenced by nextgeneration sequencing. The 16S rRNA gene sequences, specifically the amplicon sequence variants (AS Vs) identified therein, thus, obtained from a sample were then compared with a database to determine the taxonomic affiliation of the bacteria and their abundance in the soil sample.

[0092] The LASSO model created by the present inventors in this process was able to describe 58% (16S V4) of the yield variance in an exemplary maize corn field. Significant correlations between predicted and observed yields were also obtained for external / published datasets, by utilizing 16S rRNA gene abundances.

[0093] The present inventors could, therefore, show that the determined microbiome signatures allow prediction of agricultural output.

[0094] Thus, in a first aspect, the present invention relates to a method for predicting agricultural output comprising the steps of:

[0095] (i) determining the relative percentage of at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the genus Sphingomonas, of the family A21b, of the order Rokubacteriales, of the genus Nordella, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, of the order Vicinamibacterales, of the family Vicinamibacteraceae, of the genus RB41, of the family Methyloligellaceae, and of the genus Nocardioides in a soil sample of an agricultural or a prospective agricultural area, and

[0096] (ii) predicting agricultural output on the basis of the relative percentage of the at least one bacterial taxon determined in step (i).

[0097] For example, the relative percentage(s) of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 bacterial taxon / taxa selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the genus Sphingomonas, of the family A21b, of the order Rokubacteriales, of the genus Nordella, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, of the order Vicinamibacterales, of the family Vicinamibacteraceae, of the genus RB41, of the family Methyloligellaceae, and of the genus Nocardioides is (are) determined.

[0098] In one preferred embodiment, the method comprises determining the relative percentage of at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium, of the genus Bacillus, of the genus Pseudonocardia, of the genus Paenisporosarcina, of the genus Ellin6055, of the genus Rummeliibacillus, of the genus Reyranella, of the genus Rhodoplanes, and of the genus Streptomyces in the soil sample of an agricultural or a prospective agricultural area.

[0099] For example, the relative percentage(s) of at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 further bacterial taxon / taxa selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium, of the genus Bacillus, of the genus Pseudonocardia, of the genus Paenisporosarcina, of the genus Ellin6055, of the genus Rummeliibacillus, of the genus Reyranella, of the genus Rhodoplanes, and of the genus Streptomyces is (are) determined.

[0100] In one more preferred embodiment, the relative percentage of the at least one (further) bacterial taxon is determined by

[0101] (i) isolating microbial DNA from the soil sample of an agricultural or a prospective agricultural area,

[0102] (ii) amplifying the bacterial 16S rRNA genes comprised in the microbial DNA,

[0103] (iii) sequencing the bacterial 16S rRNA genes comprised in the microbial DNA,

[0104] (iv) assigning the sequenced bacterial 16S rRNA genes to the at least one bacterial taxon, and

[0105] (v) dividing (normalizing) the read counts of the bacterial 16S rRNA genes of the at least one bacterial taxon by the summed read counts of all 16S rRNA genes, thereby determining the relative percentage of the at least one bacterial taxon.

[0106] Specifically, the amplification is carried out using a polymerase chain reaction (PCR), and / or the sequencing is carried out using next generation sequencing, Maxam-Gilbert sequencing, Chain-termination methods, Shot gun sequencing, PCR sequencing, Bridge PCR, massively parallel signature sequencing (MPSS), Polony sequencing, pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, Single molecule real time (SMRT) sequencing, Nanopore DNA sequencing, sequencing by hybridization, sequencing with mass spectrometry, microfluidic Sanger sequencing, microscopy-based techniques, RNAP sequencing, or high- throughput sequencing (HTS).

[0107] More specifically, the PCR is selected from the group consisting of conventional PCR or real-time PCR (quantitative PCR or qPCR), and any derivative of both, such as TaqMan qPCR, multiplex PCR, nested PCR, high fidelity PCR, fast PCR, hot start PCR, and GC-rich PCR, and / or the sequencing is next generation sequencing.

[0108] In one even more preferred embodiment the prediction in step (ii) is carried out by inputting the relative percentages of at least two bacterial taxa in a mathematical function where p0denotes the intercept, n denotes the number of bacterial taxa used in the equation, to Pndenote the slope coefficients of the bacterial taxa, and x to xndenote the relative percentages of the bacterial taxa, which sums the weighted relative percentages of each bacterial taxon, returning a composite score reflective of the predicted agricultural output, or inputting the relative percentage of one bacterial taxon in a mathematical function where p0is the intercept, the slope coefficient, and x refers to the relative percentage of one bacterial taxon, returning a score reflective of the predicted agricultural output.

[0109] Specifically, the (composite) score (y) has a value between 0 and 1, e.g. 0, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, or 1. Thus, the (composite)s score (y) can be any n-digit number between 0 and 1, so any value between 0 and 1.

[0110] More specifically, a value close to 0 (e.g. 0, 0.1, 0.2, or 0.3) indicates a predicted low agricultural output, or a value close to 1 (e.g. 0.7, 0.8, 0.9, or 1) indicates a predicted high agricultural output.

[0111] In one still even more preferred embodiment, the relative percentage(s) is (are) obtained from

[0112] (i) a bacterial taxon of the class KD4-96,

[0113] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the order Rokubacteriales,

[0114] (iii) a bacterial taxon of the class KD4-96. a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the family Methyloligellaceae, or

[0115] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the genus Sphingomonas, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Nordella, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the order Vicinamibacterales, a bacterial taxon of the family Vicinamibacteraceae, a bacterial taxon of the genus RB41, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Nocardioides.

[0116] In one most preferred embodiment, the relative percentages are obtained from

[0117] (i) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Reyranella,

[0118] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Hyphomicrobium,

[0119] (iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the genus Reyranella,

[0120] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Reyranella, or

[0121] (v) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the genus Sphingomonas, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Nordella, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the order Vicinamibacterales, a bacterial taxon of the family Vicinamibacteraceae, a bacterial taxon of the genus RB41, a bacterial taxon of the family Methyloligellaceae, a bacterial taxon of the genus Nocardioides, a bacterial taxon of the genus Hyphomicrobium, a bacterial taxon of the genus Bacillus, a bacterial taxon of the genus Pseudonocardia, a bacterial taxon of the genus Paenisporosarcina, a bacterial taxon of the genus Ellin6055, a bacterial taxon of the genus Rummeliibacillus, a bacterial taxon of the genus Reyranella, a bacterial taxon of the genus Rhodoplanes, and a bacterial taxon of the genus Streptomyces. As mentioned above, bacterial DNA comprises variable DNA sequences, giving rise to amplicon sequence variants (AS Vs), that have been significantly modified in the course of evolution, making it possible to identify the bacteria, e.g. in a soil sample of an agricultural or a prospective agricultural area.

[0122] Thus, in one particularly preferred embodiment, the present invention relates to a method for predicting agricultural output comprising the steps of

[0123] (i) determining the relative percentage of at least one amplicon sequence variant (ASV) having a sequence selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 25 (classified to / assigned to a bacterial taxon), a fragment thereof, and a sequence having at least 90%, preferably at least 95%, more preferably at least 97%, and even more preferably at least 99%, e.g. at least 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%, sequence identity thereto in a soil sample of an agricultural or a prospective agricultural area, and

[0124] (ii) predicting agricultural output on the basis of the relative percentage of the at least one ASV determined in step (i).

[0125] For example, the relative percentage(s) of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amplicon sequence variants (ASVs) having a sequence selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 25, a fragment thereof, and a sequence having at least 90%, preferably at least 95%, more preferably at least 97%, and even more preferably at least 99%, e.g. at least 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%, sequence identity thereto, are determined.

[0126] Especially, the relative percentage of

[0127] (i) at least one amplicon sequence variant (ASV) having a sequence selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 25,

[0128] (ii) a sequence that is a fragment of the ASV according to (i), preferably, a sequence that is a fragment which is between 1 and 5, preferably between 1 and 3, and more preferably between 1 and 2, e.g. 1, 2, 3, 4, or 5, nucleotides shorter than the ASV according to (i), or

[0129] (iii) a sequence that has at least 90%, preferably at least 95%, more preferably at least 97%, and even more preferably at least 99%, e.g. at least 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%, sequence identity to the sequence according to (i) or sequence fragment according to (ii) is determined in a soil sample of an agricultural or a prospective agricultural area.

[0130] In one particularly more preferred embodiment, the relative percentage of the at least one amplicon sequence variant (ASV) (classified to / assigned to a bacterial taxon) is determined by (i) isolating microbial DNA from the soil sample of an agricultural or a prospective agricultural area, (ii) amplifying the bacterial 16S rRNA genes comprised in the microbial DNA,

[0131] (iii) sequencing the bacterial 16S rRNA genes comprised in the microbial DNA,

[0132] (iv) identifying at least one ASV from the sequenced bacterial 16S rRNA genes, and

[0133] (v) dividing (normalizing) the read counts of the bacterial 16S rRNA genes of the at least one ASV by the summed read counts of all 16S rRNA genes, thereby determining the relative percentage of the at least one ASV.

[0134] Specifically, the amplification is carried out using a polymerase chain reaction (PCR), and / or the sequencing is carried out using next generation sequencing, Maxam-Gilbert sequencing, Chain-termination methods, Shot gun sequencing, PCR sequencing, Bridge PCR, massively parallel signature sequencing (MPSS), Polony sequencing, pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, Single molecule real time (SMRT) sequencing, Nanopore DNA sequencing, sequencing by hybridization, sequencing with mass spectrometry, microfluidic Sanger sequencing, microscopy-based techniques, RNAP sequencing, or high- throughput sequencing (HTS).

[0135] More specifically, the PCR is selected from the group consisting of conventional PCR or real-time PCR (quantitative PCR or qPCR), and any derivative of both, such as TaqMan qPCR, multiplex PCR, nested PCR, high fidelity PCR, fast PCR, hot start PCR, and GC-rich PCR, and / or the sequencing is next generation sequencing.

[0136] In one particularly even more preferred embodiment, the prediction in step (ii) is carried out by inputting the relative percentages of at least two AS Vs in a mathematical function where p0denotes the intercept, n denotes the number of AS Vs used in the equation, to pndenote the slope coefficients of the AS Vs, and x to xndenote the relative percentages of the AS Vs, which sums the weighted relative percentages of each ASV, returning a composite score reflective of the predicted agricultural output, or inputting the relative percentage of one ASV in a mathematical function where0is the intercept, ?xthe slope coefficient, and refers to the relative percentage of one ASV, returning a score reflective of the predicted agricultural output.

[0137] Specifically, the (composite) score (y) has a value between 0 and 1, e.g. 0, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29, 0.3, 0.31, 0.32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4, 0.41, 0.42, 0.43, 0.44, 0.45, 0.46, 0.47, 0.48, 0.49, 0.5, 0.51, 0.52, 0.53, 0.54, 0.55, 0.56, 0.57, 0.58, 0.59, 0.6, 0.61 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, or 1. Thus, the (composite)s score (y) can be any n-digit number between 0 and 1, so any value between 0 and 1.

[0138] More specifically, a value close to 0 (e.g. 0, 0.1, 0.2, or 0.3) indicates a predicted low agricultural output, or a value close to 1 (e.g. 0.7, 0.8, 0.9, or 1) indicates a predicted high agricultural output.

[0139] As mentioned above, the method for predicting agricultural output comprises step (i) of determining the relative percentage of at least one amplicon sequence variant (ASV) having a sequence selected from the group consisting of SEQ ID NO: 1 to SEQ ID NO: 25 (classified to / assigned to a bacterial taxon).

[0140] For example, the relative percentage of at least one amplicon sequence variant (ASV) having a sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 22, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 2, SEQ ID NO: 16, SEQ ID NO: 23, SEQ ID NO: 21, SEQ ID NO: 12, SEQ ID NO: 18, SEQ ID NO: 24, SEQ ID NO: 11, SEQ ID NO: 17, SEQ ID NO: 5, and SEQ ID NO: 19 is determined in step (i) of the above method. This determination is particularly supplemented by the determination of the relative percentage of at least one further amplicon sequence variant (ASV) having a sequence selected from the group consisting of SEQ ID NO: 6, SEQ ID NO: 3, SEQ ID NO: 9, SEQ ID NO: 4, SEQ ID NO: 10, SEQ ID NO: 7, SEQ ID NO: 20, SEQ ID NO: 25, and SEQ ID NO: 8 in step (i) of the above method.

[0141] In one particularly still even more preferred embodiment, the relative percentage(s) of

[0142] (i) the ASV having a sequence according to SEQ ID NO: 1,

[0143] (ii) the ASV having a sequence according to SEQ ID NO: 1 and the ASV having a sequence according to SEQ ID NO: 2,

[0144] (iii) the ASV having a sequence according to SEQ ID NO: 1, the ASV having a sequence according to SEQ ID NO: 2, and the ASV having a sequence according to SEQ ID NO: 5, or (iv) the ASV having a sequence according to SEQ ID NO: 1, the ASV having a sequence according to SEQ ID NO: 22, the ASV having a sequence according to SEQ ID NO: 13, the ASV having a sequence according to SEQ ID NO: 14, the ASV having a sequence according to SEQ ID NO: 15, the ASV having a sequence according to SEQ ID NO: 2, the ASV having a sequence according to SEQ ID NO: 16, the ASV having a sequence according to SEQ ID NO: 23, the ASV having a sequence according to SEQ ID NO: 21, the ASV having a sequence according to SEQ ID NO: 12, the ASV having a sequence according to SEQ ID NO: 18, the ASV having a sequence according to SEQ ID NO: 24, the ASV having a sequence according to SEQ ID NO: 11, the ASV having a sequence according to SEQ ID NO: 17, the ASV having a sequence according to SEQ ID NO: 5, and the ASV having a sequence according to SEQ ID NO: 19 is (are) determined in a soil sample of an agricultural or a prospective agricultural area. In one particularly most preferred embodiment, the relative percentages of

[0145] (i) the ASV having a sequence according to SEQ ID NO: 1 and the ASV having a sequence according to SEQ ID NO: 20,

[0146] (ii) the ASV having a sequence according to SEQ ID NO: 1 and the ASV having a sequence according to SEQ ID NO: 6,

[0147] (iii) the ASV having a sequence according to SEQ ID NO: 1, the ASV having a sequence according to SEQ ID NO: 2, and the ASV having a sequence according to SEQ ID NO: 20,

[0148] (iv) the ASV having a sequence according to SEQ ID NO: 1, the ASV having a sequence according to SEQ ID NO: 2, the ASV having a sequence according to SEQ ID NO: 5, and the ASV having a sequence according to SEQ ID NO: 20, or

[0149] (iv) the ASV having a sequence according to SEQ ID NO: 1, the ASV having a sequence according to SEQ ID NO: 22, the ASV having a sequence according to SEQ ID NO: 13, the ASV having a sequence according to SEQ ID NO: 14, the ASV having a sequence according to SEQ ID NO: 15, the ASV having a sequence according to SEQ ID NO: 2, the ASV having a sequence according to SEQ ID NO: 16, the ASV having a sequence according to SEQ ID NO: 23, the ASV having a sequence according to SEQ ID NO: 21, the ASV having a sequence according to SEQ ID NO: 12, the ASV having a sequence according to SEQ ID NO: 18, the ASV having a sequence according to SEQ ID NO: 24, the ASV having a sequence according to SEQ ID NO: 11, the ASV having a sequence according to SEQ ID NO: 17, the ASV having a sequence according to SEQ ID NO: 5, the ASV having a sequence according to SEQ ID NO: 19, the ASV having a sequence according to SEQ ID NO: 6, the ASV having a sequence according to SEQ ID NO: 3, the ASV having a sequence according to SEQ ID NO: 9, the ASV having a sequence according to SEQ ID NO: 4, the ASV having a sequence according to SEQ ID NO: 10, the ASV having a sequence according to SEQ ID NO: 7, the ASV having a sequence according to SEQ ID NO: 20, the ASV having a sequence according to SEQ ID NO: 25, and the ASV having a sequence according to SEQ ID NO: 8 are determined in a soil sample of an agricultural or a prospective agricultural area.

[0150] In the above examples, the relative percentage of the ASV having a sequence according to SEQ ID NO: 14 or the relative percentage of the ASV having a sequence according to SEQ ID NO: 15 may not be determined as both AS Vs are assigned to the same bacterial taxon. In the above examples, the relative percentage of the ASV having a sequence according to SEQ ID NO: 12 or the relative percentage of the ASV having a sequence according to SEQ ID NO: 18 may not be determined as both AS Vs are assigned to the same bacterial taxon. In the above examples, the relative percentage of the ASV having a sequence according to SEQ ID NO: 14 or the relative percentage of the ASV having a sequence according to SEQ ID NO: 15 and the relative percentage of the ASV having a sequence according to SEQ ID NO: 12 or the relative percentage of the ASV having a sequence according to SEQ ID NO: 18 may not be determined as both ASVs are assigned to the same bacterial taxon.

[0151] Preferably, the agricultural output includes / represents / is crop yield of maize or green fodder yield.

[0152] Thus, the present invention preferably relates to a method of predicting crop yield of maize.

[0153] In this regard, a predicted low crop yield of maize (as mentioned above) is preferably indicative for a crop yield size of < 20 tons / hectare, e.g. 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or less tons / hectare, or a predicted high crop yield of maize (as mentioned above) is preferably indicative for a crop yield size of > 35 tons / hectare, e.g. 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more tons / hectare.

[0154] Alternatively, the present invention relates to a method of predicting green fodder yield. In this regard, a predicted low green fodder yield (as mentioned above) is indicative for a low green fodder yield size of the respective crop, or a predicted high green fodder yield (as mentioned above) is preferably indicative for a high green fodder yield size of the respective crop.

[0155] Knowing the agricultural output and the agricultural input further allows the determination of the agricultural productivity. In a second aspect, the present invention relates to a (an agricultural) (therapeutic) composition comprising at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the family A21b, of the order Rokubacteriales, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, and of the family Methyloligellaceae .

[0156] The composition may also be designated as (agricultural) cocktail.

[0157] For example, the composition may comprise at least 1, 2, 3, 4, 5, 6, 7, or 8 bacterial taxon / taxa selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the family A21b, of the order Rokubacteriales, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, and of the family Methyloligellaceae .

[0158] Preferably, the composition comprises

[0159] (i) a bacterial taxon of the class KD4-96,

[0160] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the order Rokubacteriales,

[0161] (iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the family Methyloligellaceae, or

[0162] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, and a bacterial taxon of the family Methyloligellaceae .

[0163] More preferably, the composition comprises at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium and a bacterial taxon of the genus Reyranella.

[0164] For example, the composition may comprise at least 1 or 2 further bacterial taxon / taxa selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium and a bacterial taxon of the genus Reyranella.

[0165] Even more preferably, the composition comprises

[0166] (i) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Reyranella,

[0167] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Hyphomicrobium,

[0168] (iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the genus Reyranella,

[0169] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Reyranella, or

[0170] (v) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the family Methyloligellaceae, a bacterial taxon of the genus Hyphomicrobiiim. and a bacterial taxon of the genus Reyranella.

[0171] The present inventors found that the above bacterial taxa positively correlate with the agricultural output (see experimental section and FIGURES 3 and 4). Specifically, the agricultural output is increased when the abundance of the above bacterial taxa is increased (see experimental section).

[0172] The composition preferably further comprises phosphates, manganese, iron and / or other trace elements.

[0173] The composition may be in dry (e.g. freeze-dried or powdery) or liquid form. In case the composition is in liquid form, the bacteria are part of / comprised in a solution which stabilizes the bacteria. The solution may comprise a solvent. Solvents suitable for use in the composition include without limitation water, a citrate solution, or combinations thereof. The solvent may be present in an amount sufficient to meet some user and / or process needs. For example, the solvent may be present in an amount of from about 1 wt.% to about 99 wt.%, alternatively about 10 wt.% to about 50 wt.%, or alternatively from about 99 wt.% to about 1 wt.%. Specifically, the composition may be a solution, a suspension, a dispersion, or a powder.

[0174] The composition may also comprise one or more agriculturally acceptable excipients. Said agriculturally acceptable excipients are intended to enhance the agricultural output, e.g. the yield of crops. The agriculturally acceptable excipients comprise, but are not limited to, the use of wheat flour, corn starch, gelatine, potato starch, silicon dioxide, citric acid, bicarbonate, polysorbates like Tweens, lactose, soy lecithin, casein, carboxymethyl cellulose or cellulose gum, sucrose esters, mannitol, sorbitans, Pluronic F68, alginate, xanthan gum, PEG (polyethylene glycol), corn syrup, egg, milk, glycerol, fructose, pectins, mineral oil, ester gum, and / or long-chain triglycerides.

[0175] The above composition allows to increase agricultural output. Specifically, the above composition allows to increase agricultural output by at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to a situation where no composition is used. Especially, the crop yield of maize or green fodder yield is increased.

[0176] In a third aspect, the present invention relates to the use of the (agricultural) (therapeutic) composition according to the second aspect to increase agricultural output.

[0177] Specifically, the use of the (agricultural) composition increases the agricultural output by at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to a situation where no (agricultural) composition is used. Especially, the crop yield of maize or green fodder yield is increased. In case of maize, the crop yield of maize may be increased to > 35 tons / hectare, e.g. 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more tons / hectare. In a fourth aspect, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0178] (i) carrying out the method of the first aspect,

[0179] (ii) applying the composition of the second aspect onto the (prospective) agricultural area, and

[0180] (iii) planting or cultivating agricultural plants onto the agricultural area.

[0181] In one preferred embodiment - by carrying out the method according to the first aspect - a composite (score) having a value close to 0 is obtained which is indicative for a predicted low agricultural output. In this case, the composition according to the second aspect is applied onto the (prospective) agricultural area. Subsequently, agricultural plants are planted or cultivated onto the agricultural area. In this way, the agricultural output can be increased.

[0182] Thus, in one particular embodiment, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0183] (i) carrying out the method according to the first aspect, thereby obtaining a (composite) score having a value close to 0 which is indicative for a predicted low agricultural output,

[0184] (ii) applying the composition according to the second aspect onto the (prospective) agricultural area, and

[0185] (iii) planting or cultivating agricultural plants onto the agricultural area.

[0186] In this way, the agricultural area with low agricultural output can effectively be treated with the composition in order to increase said output.

[0187] The composition may be in liquid or solid form. In embodiments in which the composition is in liquid form, it may be sprayed or poured onto the agricultural area. In aspects in which the composition is in solid (e.g. freeze-dried or powdery) form, it may be spread onto the surface of the agricultural area or it may be mixed into the soil which is part of the agricultural area.

[0188] Specifically, the use of the composition increases the agricultural output by at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to a situation where no (agricultural) composition is used.

[0189] In one more preferred embodiment, the agricultural plants are maize or grass / green fodder. In this case, the crop yield of maize or grass / green fodder yield is increased by the application of the composition.

[0190] In a fifth aspect, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0191] (i) applying the composition of the second aspect onto an (a prospective) agricultural area, and

[0192] (ii) planting or cultivating agricultural plants onto the agricultural area. The composition may be in liquid or solid form. In embodiments in which the composition is in liquid form, it may be sprayed or poured onto the agricultural area. In aspects in which the composition is in solid (e.g. freeze-dried or powdery) form, it may be spread onto the surface of the agricultural area or it may be mixed into the soil which is part of the agricultural area. Particularly, the application is carried out by spraying the composition onto the (prospective) agricultural area or by (finely) distributing the composition onto the (prospective) agricultural area.

[0193] Specifically, the use of the composition increases the agricultural output by at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to a situation where no (agricultural) composition is used.

[0194] In one preferred embodiment, the agricultural plants are maize or grass / green fodder. In this case, the crop yield of maize or grass / green fodder yield is increased by the application of the composition.

[0195] In a sixth aspect, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0196] (i) applying the composition of the second aspect onto agricultural plant seeds, thereby coating the agricultural plant seeds with the composition,

[0197] (ii) sowing the coated agricultural plant seeds into the (prospective) agricultural area, and

[0198] (iii) cultivating agricultural plants derived therefrom onto the agricultural area.

[0199] Alternatively, the composition may be contacted with a part of a plant that is above the ground, for example, the leaves, flowers, fruit, and / or stem.

[0200] In this case, the present invention relates to a method of increasing agricultural output comprising the steps of:

[0201] (i) applying the composition according to the second aspect onto a part of a (an agricultural) plant that is above the ground, thereby coating the part of the (agricultural) plant that is above the ground with the composition, and

[0202] (ii) cultivating the (agricultural) plant onto / in the agricultural area.

[0203] The composition may be in liquid or solid form. In embodiments in which the composition is in liquid form, it may be sprayed or poured onto the seeds as well as plants or parts thereof. In aspects in which the composition is in solid (e.g. freeze-dried or powdery) form, it may be spread onto the surface of the seeds as well as plants or parts thereof.

[0204] Particularly, the application of the composition onto the agricultural plant seeds may take place / be carried out via dip coating or spray coating. Alternatively, the application of the composition onto the part of the (agricultural) plant may take place / be carried out by spray coating.

[0205] Preferably, the coating covers at least 1 %, more preferably at least 50%, even more preferably at least 80%, and most preferably at least 90% or even 100%, of the plant or plant seed, e.g. at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26,

[0206] 27, 28 ,29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52,

[0207] 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78,

[0208] 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99%, or 100% of the plant or plant seed.

[0209] Specifically, the use of the composition increases the agricultural output by at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to a situation where no (agricultural) composition is used.

[0210] In one preferred embodiment, the agricultural seeds are maize or grass / green fodder seeds or the agricultural plants are maize or grass / green fodder plants. In this case, the crop yield of maize or grass / green fodder yield is increased by the application of the composition.

[0211] The present invention is summarized as follows:

[0212] 1. A method for predicting agricultural output comprising the steps of:

[0213] (i) determining the relative percentage of at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the genus Sphingomonas, of the family A21b, c the order Rokubacteriales, of the genus Nordella, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, of the order Vicinamibacterales, of the family Vicinamibacteraceae, of the genus RB41, of the family Methyloligellaceae, and of the genus Nocardioides in a soil sample of an agricultural or a prospective agricultural area, and

[0214] (ii) predicting agricultural output on the basis of the relative percentage of the at least one bacterial taxon determined in step (i).

[0215] 2. The method of item 1, wherein the method comprises determining the relative percentage of at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium, of the genus Bacillus, of the genus Pseudonocardia, of the genus Paenisporosarcina, of the genus Ellin6055, of the genus Rummeliibacillus, of the genus Reyranella, of the genus Rhodoplanes, and of the genus Streptomyces in the soil sample of an agricultural or a prospective agricultural area.

[0216] 3. The method of items 1 or 2, wherein the relative percentage of the at least one bacterial taxon is determined by

[0217] (i) isolating microbial DNA from the soil sample of an agricultural or a prospective agricultural area,

[0218] (ii) amplifying the bacterial 16S rRNA genes comprised in the microbial DNA,

[0219] (iii) sequencing the bacterial 16S rRNA genes comprised in the microbial DNA, (iv) assigning the sequenced bacterial 16S rRNA genes to the at least one bacterial taxon, and

[0220] (v) dividing (normalizing) the read counts of the bacterial 16S rRNA genes of the at least one bacterial taxon by the summed read counts of all 16S rRNA genes, thereby determining the relative percentage of the at least one bacterial taxon. The method of item 3, wherein the amplification is carried out using a polymerase chain reaction (PCR). The method of any one of items 3 or 4, wherein the sequencing is next generation sequencing. The method of any one of items 1 to 5, wherein the prediction in step (ii) is carried out by inputting the relative percentages of at least two bacterial taxa in a mathematical function where p0denotes the intercept, n denotes the number of bacterial taxa used in the equation, to pndenote the slope coefficients of the bacterial taxa, and x to xndenote the relative percentages of the bacterial taxa, which sums the weighted relative percentages of each bacterial taxon, returning a composite score reflective of the predicted agricultural output, or inputting the relative percentage of one bacterial taxon in a mathematical function where p0is the intercept, the slope coefficient, and x refers to the relative percentage of one bacterial taxon, returning a score reflective of the predicted agricultural output. The method of item 6, wherein the (composite) score has a value between 0 and 1. The method of item 7, wherein a value close to 0 indicates a predicted low agricultural output, or a value close to 1 indicates a predicted high agricultural output. The method of any one of items 1 to 8, wherein the relative percentage(s) is (are) obtained from

[0221] (i) a bacterial taxon of the class KD4-96,

[0222] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the order Rokubacteriales,

[0223] (iii) a bacterial taxon of the class KD4-96. a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the family Methyloligellaceae, or (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the genus Sphingomonas, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Nordella, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the order Vicinamibacterales, a bacterial taxon of the family Vicinamibacteraceae, a bacterial taxon of the genus RB41, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Nocardioides. The method of any one of items 2 to 9, wherein the relative percentages are obtained from

[0224] (i) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Reyranella,

[0225] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Hyphomicrobium,

[0226] (iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the genus Reyranella,

[0227] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Reyranella, or

[0228] (v) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the genus Sphingomonas, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Nordella, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the order Vicinamibacterales, a bacterial taxon of the family Vicinamibacteraceae, a bacterial taxon of the genus RB41, a bacterial taxon of the family Methyloligellaceae, a bacterial taxon of the genus Nocardioides, a bacterial taxon of the genus Hyphomicrobium, a bacterial taxon of the genus Bacillus, a bacterial taxon of the genus Pseudonocardia, a bacterial taxon of the genus Paenisporosarcina, a bacterial taxon of the genus Ellin6055, a bacterial taxon of the genus Rummeliibacillus, a bacterial taxon of the genus Reyranella, a bacterial taxon of the genus Rhodoplanes, and a bacterial taxon of the genus Streptomyces. The method of any one of items 1 to 10, wherein the agricultural output includes / represents / is crop yield of maize or green fodder yield. A composition comprising at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the family A21b, of the order Rokubacteriales, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gate Hales, and of the family Methyloligellaceae .

[0229] 13. The composition of item 12, wherein the composition comprises

[0230] (i) a bacterial taxon of the class KD4-96,

[0231] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the order Rokubacteriales,

[0232] (iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the family Methyloligellaceae, or

[0233] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Altererythrobacter , a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, and a bacterial taxon of the family Methyloligellaceae.

[0234] 14. The composition of items 12 or 13, wherein the composition comprises at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium and a bacterial taxon of the genus Reyranella.

[0235] 15. The composition of item 14, wherein the composition comprises

[0236] (i) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Reyranella,

[0237] (ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Hyphomicrobium,

[0238] (iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the genus Reyranella,

[0239] (iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Reyranella, or

[0240] (v) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Altererythrobacter , a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the family Methyloligellaceae, a bacterial taxon of the genus Hyphomicrobium, and a bacterial taxon of the genus Reyranella.

[0241] 16. The composition of any one of items 12 to 15, wherein the composition is in dry or liquid form.

[0242] 17. Use of the composition of any one of items 12 to 16 to increase agricultural output.

[0243] 18. The use of item 17, wherein crop yield of maize or green fodder yield is increased. 19. A method of increasing agricultural output comprising the steps of:

[0244] (i) carrying out the method of any one of items 1 to 11(, thereby obtaining a (composite) score having a value close to 0 which is indicative for a predicted low agricultural output),

[0245] (ii) applying the composition of any one of items 12 to 16 onto the (prospective) agricultural area, and

[0246] (iii) planting or cultivating agricultural plants onto the agricultural area.

[0247] 20. The method of item 19, wherein the agricultural plants are maize or grass / green fodder.

[0248] 21. A method of increasing agricultural output comprising the steps of:

[0249] (i) applying the composition of any one of items 12 to 16 onto an (a prospective) agricultural area, and

[0250] (ii) planting or cultivating agricultural plants onto the agricultural area.

[0251] 22. The method of item 21, wherein the agricultural plants are maize or grass / green fodder.

[0252] 23. The method of items 21 or 22, wherein the composition is in liquid or solid, preferably powdery, form.

[0253] 24. The method of any one of items 21 to 23, wherein the application is carried out by spraying the composition onto the (prospective) agricultural area or by (finely) distributing the composition onto the (prospective) agricultural area.

[0254] 25. A method of increasing agricultural output comprising the steps of:

[0255] (i) applying the composition of any one of items 12 to 16 onto agricultural plant seeds, thereby coating the agricultural plant seeds with the composition,

[0256] (ii) sowing the coated agricultural plant seeds into the (prospective) agricultural area, and

[0257] (iii) cultivating agricultural plants derived therefrom onto the agricultural area.

[0258] 26. The method of item 25, wherein the agricultural plants are maize or grass / green fodder.

[0259] 27. The method of items 25 or 26, wherein the composition is in liquid form.

[0260] 28. The method of items 25 to 27, wherein the application is carried out by dip coating or spray coating.

[0261] Various modifications and variations of the invention will be apparent to those skilled in the art without departing from the scope of invention. Although the invention has been described in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention which are obvious to those skilled in the art in the relevant fields are intended to be covered by the present invention. BRIEF DESCRIPTION OF THE FIGURES

[0262] The following Figures are merely illustrative of the present invention and should not be construed to limit the scope of the invention as indicated by the appended claims in any way.

[0263] FIGURE 1: Schematic representation of the process used to develop a maize crop yield prediction model.

[0264] FIGURE 2: Left: The 290 ASVs obtained from a core microbiome analysis were input into a LASSO regression machine learning algorithm with Monte Carlo cross-validation. The optimal lambda parameter resulted in an R2 of 0.578 in cross-validation. The identity function y=x is plotted. Right: To show the variable importance of the final model, the absolute values of the non-zero coefficients for the final model were plotted and their signs indicated by color. Note that the features were z-scored prior to model building and thus the regression coefficients / variable importances are independent of ASV abundance.

[0265] FIGURE 3: For each of the 8 test datasets (see TABLE 1), we evaluated the cross-study predictive performance of the locally trained model. Therefore, yield scores were calculated and compared to the studies’ respective standardized productivity metric. Note that we used Spearman’s rank correlation coefficient, as the relationship between predicted and observed values showed in some cases a non-linear pattern.

[0266] FIGURE 4: For each of the 8 test datasets (see TABLE 1), the 25 ASVs from the locally trained Grabko model (if detected in the respective study) were subjected to a correlation analysis against the respective yield metric. Statistical significance of the Spearman‘s rank correlation coefficients, after adjusting for multiple testing via false discovery rate is indicated by color (alpha level = 0.05). Features were grouped according to their influence (positive or negative coefficient) in the locally trained Grabko model. Note that in some cases no analysis was possible due to absence of the respective AVS from the dataset.

[0267] FIGURE 5: For each of the 25 ASVs from the locally trained Grabko model, the lowest taxon that could be taxonomically assigned was subjected to a correlation analysis against the respective yield metric for each of the 8 test datasets (if present) (see Table 1). Statistical significance of the Spearman‘s rank correlation coefficients, after adjusting for multiple testing via false discovery rate is indicated by color (alpha level = 0.05). Features were grouped according to their influence (positive or negative coefficient) in the locally trained Grabko model. Note that in some cases no analysis was possible due to absence of the respective taxon from the dataset. EXAMPLES

[0268] The examples given below are for illustrative purposes only and do not limit the invention described above in any way.

[0269] MATERIAL & METHODS

[0270] Yield and yield-related variable measurement

[0271] A schematic representation of the process used to develop an agricultural output prediction model is shown in FIGURE 1.

[0272] The cooperating agricultural company in Grabko, Brandenburg, provided yield mapping data (local yield measured via volume flow + GPS coordinates). These yield data, measured in tons per hectare, were transformed into a maize yield score ranging from 0 to 1 using Min-Max normalization.

[0273] For validation of the model and the thereby found bacterial taxa, published datasets were aggregated, which measured the prokaryotic soil microbiome (16S) and, for which yield, or yield- related measurements were either directly available from the study or which could be retrieved through the geographical coordinates, by subsequent remote sensing data retrieval using the MODIS satellite to obtain Normalized Difference Vegetation Index (NDVI) data. TABLE 1 summarizes the datasets used for this meta-analysis. TABLE 1: Overview of datasets used for model training and testing

[0274] *BNPP = belowground net primary production; **NDVI = normalized difference vegetation index

[0275] DNA isolation from soil samples and preparation of sequencing libraries

[0276] For DNA isolation, 250 mg of soil was used per sample. DNA isolation was carried out using the DNeasy PowerLyzer PowerSoil Kit (Qiagen) following an internal standard protocol. The 80 samples were processed in batches of up to 12 samples per batch, following a randomization approach. The sequencing libraries for 16S V4 (bacteria) were prepared according to a standard protocol. Sequencing of the samples was performed using an Illumina iSeq 100 sequencer.

[0277] Bioinformatic analysis of sequencing data

[0278] The processing of the high-throughput sequencing data was performed within RStudio [1,2]. Essentially, the software packages dada2 [3] and microeco [4] were used for processing of the sequencing data. For taxonomic classification the SILVA database was used (silva_nr99_vl38.1) [5], Machine learning was done using the caret package [6],

[0279] RESULTS

[0280] Initially, a core microbiome analysis was performed, in which amplicon sequence variants (ASVs) were retained harboring both a prevalence of at least 50 % (i.e. were present in at least 40 of the 80 samples analyzed) and possessing a mean abundance of at least 0.05 % over all samples. This gave rise to a core microbiome of 290 ASVs, which were used as potential features in a supervised machine learning algorithm using least absolute shrinkage and selection operator (LASSO) regression, to predict the maize yield on the field in Grabko. The resulting model was able to predict ~ 58 % of variation in the maize yield in cross-validation and utilized 25 ASVs as features. The 25 ASVs are summarized in TABLE 2.

[0281] TABLE 2: ASV Summary

[0282] FIGURE 2 summarizes the model performance, as well as the importances and contributions (positive or negative) to the model prediction.

[0283] This model was validated using data from previously published studies, where global soil microbiome data, along with various productivity metrics, such as yield, net primary production or soil health were available (see TABLE 1). Predictions using our locally trained model generally agreed well with the test dataset’s respective productivity metric. In 5 out of the 8 test datasets, highly significant spearman’s rank correlation coefficients were observed (see FIGURE 3). Further an in-depth analysis of the individual ASVs was performed. Their association with various productivity metrics was obtained by exploiting publicly available datasets (see TABLE 1). Essentially, from the 11 ASVs that possessed positive regression coefficients in the LASSO model, 10 showed consistent significant positive correlations with the respective productivity metrics (see FIGURE 4) Similar patterns were also observed when the lowest taxa, to which the respective ASVs could be taxonomically assigned, were correlated with the respective productivity metrics.

[0284] The lowest taxon that could be taxonomically assigned for each of the 25 ASVs was further subjected to a correlation analysis against the respective yield metric for each of the 8 test datasets (if present) (see FIGURE 5, TABLE 1). Black color indicates statistically significant Spearman‘s rank correlation coefficients, after adjusting for multiple testing via false discovery rate (alpha level = 0.05). Features were grouped according to their influence (positive or negative coefficient) in the locally trained Grabko model. Note that in some cases no analysis was possible due to absence of the respective taxon from the dataset.

[0285] SUMMARY

[0286] Herein, the discovery of a set of 25 ASVs of prokaryotic origin in soil is described that show consistent responsiveness to a variety of agricultural yield-related variables. Moreover, the lowest taxa, to which the 25 ASVs could be assigned (see TABLE 2), showed similar patterns.

[0287] REFERENCES

[0288] [1] R Core Team (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https: / / www.R-project.org / .

[0289] [2] RStudio Team (2021). RStudio: Integrated Development Environment for R. RStudio, PBC, Boston, MA URL http: / / www.rstudio.com / .

[0290] [3] Callahan BJ, McMurdie PJ, Rosen MJ, Han AW, Johnson AJA, Holmes SP (2016). “DADA2: High-resolution sample inference from Illumina amplicon data.” Nature Methods , *13*, 581- 583. doi: 10.1038 / nmeth.3869 (URL: https: / / doi.org / 10.1038 / nmeth.3869).

[0291] [4] Chi Liu, Yaoming Cui, Xiangzhen Li, Minjie Yao, microeco: an R package for data mining in microbial community ecology, FEMS Microbiology Ecology, Volume 97, Issue 2, February 2021, fiaa255.

[0292] [5] Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, Peplies J, Glbckner FO (2013) The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Opens external link in new windowNucl. Acids Res. 41 (DI): D590-D596.

[0293] [6] Max Kuhn (2022). caret: Classification and Regression Training. R package version 6.0-93. https: / / CRAN.R-project.org / package=caret

[0294] [7] A. Durrer et al., „Organic farming practices change the soil bacteria community, improving soil quality and maize crop yields“, PeerJ, Bd. 9, S. el 1985, Sep. 2021, doi: 10.7717 / peerj.11985.

[0295] [8] W. Hu et al., „Aridity-driven shift in biodiversity-soil multifunctionality relationships^ Nat Commun, Bd. 12, Nr. 1, S. 5350, Sep. 2021, doi: 10.1038 / s41467-021-25641-0.

[0296] [9] M. Labouyrie et al., „Pattems in soil microbial diversity across Europe“, Nat Commun, Bd. 14, Nr. 1, S. 3311, Juni 2023, doi: 10.1038 / s41467-023-37937-4.

[0297]

[0010] A. Fox et al., „Small-scale agricultural grassland management can affect soil fungal community structure as much as continental scale geographic patterns“, FEMS Microbiology Ecology, Bd. 97, Nr. 12, S. fiabl48, Dez. 2021, doi: 10.1093 / femsec / fiabl48.

[0298]

[0011] M. Delgado-Baquerizo et al., „A global atlas of the dominant bacteria found in soil“, Science, Bd. 359, Nr. 6373, S. 320-325, Jan. 2018, doi: 10.1126 / science.aap9516.

[0299]

[0012] R. C. Wilhelm, H. M. van Es, and D. H. Buckley, Predicting measures of soil health using the microbiome and supervised machine learning^ Soil Biology and Biochemistry, Bd. 164, S. 108472, Jan. 2022, doi: 10.1016 / j soilbio.2021.108472.

[0300]

[0013] A. Lanzen et al., „The Community Structures of Prokaryotes and Fungi in Mountain Pasture Soils are Highly Correlated and Primarily Influenced by pH“, Front. Microbiol., Bd. 6, Nov. 2015, doi: 10.3389 / fmicb.2015.01321.

Claims

CLAIMS1. A method for predicting agricultural output comprising the steps of:(i) determining the relative percentage of at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the genus Sphingomonas, of the family A21b, c the order Rokubacteriales, of the genus Nordella, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, of the order Vicinamibacterales, of the family Vicinamibacteraceae, of the genus RB41, of the family Methyloligellaceae, and of the genus Nocardioides in a soil sample of an agricultural or a prospective agricultural area, and(ii) predicting agricultural output on the basis of the relative percentage of the at least one bacterial taxon determined in step (i).

2. The method of claim 1, wherein the method comprises determining the relative percentage of at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium, of the genus Bacillus, of the genus Pseudonocardia, of the genus Paenisporosarcina, of the genus Ellin6055, of the genus Rummeliibacillus, of the genus Reyranella, of the genus Rhodoplanes, and of the genus Streptomyces in the soil sample of an agricultural or a prospective agricultural area.

3. The method of claims 1 or 2, wherein the relative percentage of the at least one bacterial taxon is determined by(i) isolating microbial DNA from the soil sample of an agricultural or a prospective agricultural area,(ii) amplifying the bacterial 16S rRNA genes comprised in the microbial DNA,(iii) sequencing the bacterial 16S rRNA genes comprised in the microbial DNA,(iv) assigning the sequenced bacterial 16S rRNA genes to the at least one bacterial taxon, and(v) dividing (normalizing) the read counts of the bacterial 16S rRNA genes of the at least one bacterial taxon by the summed read counts of all 16S rRNA genes, thereby determining the relative percentage of the at least one bacterial taxon.

4. The method of any one of claims 1 to 3, wherein the prediction in step (ii) is carried out by inputting the relative percentages of at least two bacterial taxa in a mathematical functionwhere p0denotes the intercept, n denotes the number of bacterial taxa used in the equation, to pndenote the slope coefficients of the bacterial taxa, and x to xndenote the relative percentages of the bacterial taxa, which sums the weighted relative percentages of each bacterial taxon, returning a composite score reflective of the predicted agricultural output, or inputting the relative percentage of one bacterial taxon in a mathematical functionwhere p0is the intercept,the slope coefficient, and x refers to the relative percentage of one bacterial taxon, returning a score reflective of the predicted agricultural output.

5. The method of claim 4, wherein the (composite) score has a value between 0 and 1 , wherein preferably a value close to 0 indicates a predicted low agricultural output, or a value close to 1 indicates a predicted high agricultural output.

6. The method of any one of claims 1 to 5, wherein the relative percentage(s) is (are) obtained from(i) a bacterial taxon of the class KD4-96,(ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the order Rokubacteriales,(iii) a bacterial taxon of the class KD4-96. a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the family Methyloligellaceae, or(iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the genus Sphingomonas, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Nordella, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the order Vicinamibacterales, a bacterial taxon of the family Vicinamibacteraceae, a bacterial taxon of the genus RB41, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Nocardioides.

7. The method of any one of claims 2 to 6, wherein the relative percentages are obtained from(i) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Reyranella,(ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Hyphomicrobiiim.(iii) a bacterial taxon of the class KD4-96. a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the genus Reyranella,(iv) a bacterial taxon of the class KD4-96. a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Reyranella, or(v) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the genus Sphingomonas, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Nordella, a bacterial taxon of the genus Alter erythrobacter, a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the order Vicinamibacterales, a bacterial taxon of the family Vicinamibacteraceae, a bacterial taxon of the genus RB41, a bacterial taxon of the family Methyloligellaceae, a bacterial taxon of the genus Nocardioides, a bacterial taxon of the genus Hyphomicrobium, a bacterial taxon of the genus Bacillus, a bacterial taxon of the genus Pseudonocardia, a bacterial taxon of the genus Paenisporosarcina, a bacterial taxon of the genus Ellin6055, a bacterial taxon of the genus Rummeliibacillus, a bacterial taxon of the genus Reyranella, a bacterial taxon of the genus Rhodoplanes, and a bacterial taxon of the genus Streptomyces.

8. A composition comprising at least one bacterial taxon selected from the group consisting of a bacterial taxon of the class KD4-96, of the genus lamia, of the family A21b, of the order Rokubacteriales, of the genus Alter erythrobacter, of the order IMCC26256, of the order Gaiellales, and of the family Methyloligellaceae .

9. The composition of claim 8, wherein the composition comprises(i) a bacterial taxon of the class KD4-96,(ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the order Rokubacteriales,(iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the family Methyloligellaceae, or(iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Altererythrobacter , a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, and a bacterial taxon of the family Methyloligellaceae .

10. The composition of claims 8 or 9, wherein the composition comprises at least one further bacterial taxon selected from the group consisting of a bacterial taxon of the genus Hyphomicrobium and a bacterial taxon of the genus Reyranella.

11. The composition of claim 10, wherein the composition comprises(i) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Reyranella,(ii) a bacterial taxon of the class KD4-96 and a bacterial taxon of the genus Hyphomicrobium,(iii) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, and a bacterial taxon of the genus Reyranella,(iv) a bacterial taxon of the class KD4-96, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the family Methyloligellaceae, and a bacterial taxon of the genus Reyranella, or(v) a bacterial taxon of the class KD4-96, a bacterial taxon of the genus lamia, a bacterial taxon of the family A21b, a bacterial taxon of the order Rokubacteriales, a bacterial taxon of the genus Altererythrobacter , a bacterial taxon of the order IMCC26256, a bacterial taxon of the order Gaiellales, a bacterial taxon of the family Methyloligellaceae, a bacterial taxon of the genus Hyphomicrobium, and a bacterial taxon of the genus Reyranella.

12. Use of the composition of any one of claims 8 to 11 to increase agricultural output.

13. A method of increasing agricultural output comprising the steps of:(i) carrying out the method of any one of claims 1 to 7 (, thereby obtaining a (composite) score having a value close to 0 which is indicative for a predicted low agricultural output),(ii) applying the composition of any one of claims 8 to 11 onto the (prospective) agricultural area, and(iii) planting or cultivating agricultural plants onto the agricultural area.

14. A method of increasing agricultural output comprising the steps of:(i) applying the composition of any one of claims 8 to 11 onto an (a prospective) agricultural area, and (ii) planting or cultivating agricultural plants onto the agricultural area.

15. A method of increasing agricultural output comprising the steps of:(i) applying the composition of any one of claims 8 to 11 onto agricultural plant seeds, thereby coating the agricultural plant seeds with the composition, (ii) sowing the coated agricultural plant seeds into the (prospective) agricultural area, and(iii) cultivating agricultural plants derived therefrom onto the agricultural area.