Decision-making platform to develop customized biofertilizers for high-protein crops

By employing machine learning models to predict the performance of rhizobia strains in specific plant and soil contexts, the method streamlines the development of customized inoculants, addressing the inefficiencies and costs of traditional approaches while enhancing crop yields.

WO2025125613A1PCT designated stage expired Publication Date: 2025-06-19AARHUS UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/086318
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for formulating customized inoculants for legume crops are laborious and costly, requiring extensive field trials or laboratory experiments to identify suitable rhizobia strains adapted to specific plant and soil conditions.

Method used

A computer-implemented method using machine learning models to predict the competitiveness and effectiveness of rhizobia strains for specific plant and soil combinations, allowing for the formulation of customized inoculants without the need for extensive field trials.

Benefits of technology

This approach significantly reduces the time and cost associated with developing customized inoculants, enabling the identification of optimal rhizobia strains for improved nitrogen fixation and crop yields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000024_0001
    Figure IMGF000024_0001
  • Figure IMGF000024_0002
    Figure IMGF000024_0002
  • Figure IMGF000032_0001
    Figure IMGF000032_0001
Patent Text Reader

Abstract

The present disclosure relates to customized inoculant formulations for a plant grown in a soil, as well as methods to provide and manufacture said customized inoculant formulations.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Decision-making platform to develop customized biofertilizers for high-protein crops

[0002] The present disclosure relates to customized inoculant formulations for a plant grown in a soil, as well as methods to provide and manufacture said customized inoculant formulations.

[0003] Background

[0004] Legumes are plants with a high protein content that possess a unique ability to house nitrogen-fixing bacteria called rhizobia. This unique interaction takes place in the root nodules of legume plants, where rhizobia convert atmospheric nitrogen into a usable form for the plants, while the legumes provide the bacteria with a source of carbohydrates and shelter. This natural and symbiotic relationship between legumes and rhizobia is of great significance in sustainable agriculture, offering a cost-effective alternative to chemical nitrogen fertilizers for legume crops.

[0005] In many regions, native rhizobia populations often fail to effectively meet the nitrogen demands of promiscuous cultivars, necessitating a safer approach involving the use of exotic elite rhizobia strains. This is preferred over relying on resident strains, which may have uncertain potential. Additionally, when new plant species are introduced in different locations, co-evolved rhizobia strains are typically lacking in foreign soils. Consequently, successful crop introductions in new regions depend on inoculation with exotic rhizobia. Numerous examples from the Southern hemisphere, like African soils and soybean cultivars, Australian and New Zealand soils with forage and grain legumes, soybean in Argentina, common bean in Brazilian soils, and forage production of Lotus corniculatus and clover in Uruguay, illustrate this phenomenon.

[0006] Currently, commercial inoculants typically consist of rhizobia strains known for their exceptional performance. However, relying on commercial inoculants based on exotic rhizobia strains can lead to problems with promiscuous cultivars. These selected strains, known for their nitrogen-fixing efficiency, may fail to establish successful symbiosis as they are not competitive against inefficient native rhizobia. Consequently, the absence of the inoculated strain in the nodules results in low productivity and inconsistent field performance. To address these inoculation challenges, the key lies in designing tailored inoculants by selecting naturally-evolved and locally-sourced rhizobia with exceptional symbiotic performance (effectiveness) and adaptability to specific agroclimatic conditions in a given region (competitiveness). This approach is crucial for agricultural lands worldwide to tackle the competition problem effectively and increase crop yields without relying on nitrogen fertilizers.

[0007] Indeed, local adaptation plays a more significant role in shaping microbial cooperation than partner choice. The high competitiveness of soil native populations is influenced by various external factors, including the physiological condition and growth stage of the rhizobia culture, the distribution of bacterial cells in the soil, the presence of salt, soil pH at the sampling location, soil nutrient content, soil surface coverage, soil temperatures during the crop season, and the resistance to herbicides used in agricultural practices. The big challenge is to determine which one of these variables is most important to study or control. Thus, to fully understand the effectiveness of elite strains in fixing nitrogen in the presence of native and / or commercial strains, different field scenarios need to be thoroughly investigated, which is laborious and time consuming.

[0008] Thus, at the present moment, the formulation of customized inoculants is but an academic question, as the costs to carry such tests are too prohibitive to be offered as a service that can be afforded by local farmers.

[0009] Summary

[0010] The present disclosure addresses the challenge of formulating customized inoculants which specifically promote the growth of a specific plant in a specific soil.

[0011] The process of identifying rhizobia strains that are best suited for a particular plant and soil combination can be quite laborious. Presently, the only viable approach involves conducting a field trial, which typically lasts 6 months, or controlled experiments in laboratory or greenhouse conditions. These experiments demand a considerable allocation of resources. Indeed, these studies often require a considerable number of plants to be used to test a large amount of rhizobia strains. The present application discloses a method of providing a customized inoculant formulation for a plant in a soil which does not require the need of setting up a field trial, and which thus results in a significant reduction in costs.

[0012] Thus, in an aspect, the present disclosure is directed to a method for providing a customized inoculant formulation for a plant grown in a soil comprising at least one preferred rhizobia strain, said method comprising the steps of: a. providing information / data of the plant, comprising at least said plant species; b. providing information / data of the soil, said information / data comprising at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (mpn) in the soil; and iv. information / data of at least one rhizobia present in the soil; c. predicting, by a machine learning model, the competitiveness of each rhizobia strain selected from a plurality of known rhizobia strains, wherein the competitiveness of each rhizobia strain is the ability to occupy the nodules of a plant in a given soil with multiple rhizobia strains present, and wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant comprising at least said plant species of the plurality of plants; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality of soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia strain present in said soil; and iv. a plurality of combinations of plant and soil, and associated plant data for each combinations of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the competitiveness of each rhizobia strain selected from the known rhizobia strains in i.; d. predicting, by a machine learning model, the effectiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or the one or more rhizobia strains predicted in step e. to have high competitiveness, wherein the effectiveness of a rhizobia strain is the ability to acquire nitrogen through the biological nitrogen fixation and / or its effect on plant biomass production in the plant grown in the soil, wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant comprising at least said plant species of the plurality of plants; and iii. information / data of a plurality of soils, and associated soil data for each soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia present in said soil; and iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the effectiveness of each rhizobia strain selected from the known rhizobia strains in i; f. selecting at least one preferred rhizobia strain; and g. formulating the customized inoculant comprising the at least one preferred rhizobia strain; thereby providing the customized inoculant formulation for said plant and soil.

[0013] It will be evident that steps a. to g. as described herein are performed, in some embodiments, sequentially. However, steps c. and d. as described herein might be performed at the same time, or step d might be performed before step c. as described in the section “the machine learning model” of the present disclosure.

[0014] The presently disclosed method can be at least partly implemented on a computer, such as general purpose computer, for example having one or more processing units, such as one or more local processing units and / or one or more external processing units, for example cloud computing. Computer implemented steps of the presently disclosed method can for example be steps a. to d. In some embodiments, step f. is a computer implemented step. In some embodiments, step g. is a computer implemented step.

[0015] Thus, in some embodiments, steps a., b., c., and d. are computer implemented steps. In some embodiments, steps a., b., c., d., and f. are computer implemented steps. In some embodiments, steps a., b., c., d. and g. are computer implemented steps. In some embodiments, steps a., b., c., d., f., and g. are computer implemented steps.

[0016] Thus, in an aspect, the present disclosure is directed to a computer-implemented method for providing a customized inoculant formulation for a plant grown in a soil comprising at least one preferred rhizobia strain, said method comprising the steps of: a. providing information / data of the plant, comprising at least said plant species; b. providing information / data of the soil, said information / data comprising at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (mpn) in the soil; and iv. information / data of at least one rhizobia present in the soil; c. predicting, by a machine learning model, the competitiveness of each rhizobia strain selected from a plurality of known rhizobia strains, wherein the competitiveness of each rhizobia strain is the ability to occupy the nodules of a plant in a given soil with multiple rhizobia strains present, and wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant comprising at least said plant species of the plurality of plants; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality of soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia strain present in said soil; and iv. a plurality of combinations of plant and soil, and associated plant data for each combinations of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the competitiveness of each rhizobia strain selected from the known rhizobia strains in i.; d. predicting, by a machine learning model, the effectiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or the one or more rhizobia strains predicted in step e. to have high competitiveness, wherein the effectiveness of a rhizobia strain is the ability to acquire nitrogen through the biological nitrogen fixation and / or its effect on plant biomass production in the plant grown in the soil, wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant comprising at least said plant species of the plurality of plants; and iii. information / data of a plurality of soils, and associated soil data for each soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia present in said soil; and iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the effectiveness of each rhizobia strain selected from the known rhizobia strains in i; h. selecting at least one preferred rhizobia strain; and i. formulating the customized inoculant comprising the at least one preferred rhizobia strain; thereby providing the customized inoculant formulation for said plant and soil.

[0017] In an aspect, the present disclosure is directed to a customized inoculant wherein the customized inoculant comprises a preferred rhizobia strain identified by the method as described herein .

[0018] In an aspect, the present disclosure is directed to a customized inoculant wherein the customized inoculant comprises a preferred rhizobia strain identified by the method as described herein. In an aspect, the present disclosure is directed to a method for providing / manufacturing a customized inoculant as defined herein comprising the method described herein, and formulating the customized inoculant formulation with the at least one preferred rhizobia strain selected as described herein.

[0019] In an aspect, the present disclosure is directed to a method to reduce the number of rhizobia strains to be tested to formulate a customized inoculant formulation for a plant of interest grown in a soil of interest, wherein said method comprises steps a. to d. of the method for providing a customized inoculant formulation for a plant grown in a soil as described herein, and testing the rhizobia predicted to have high effectiveness and / or competitiveness for the combination of said soil of interest and plant of interest.

[0020] In an aspect, the present disclosure is directed to a computer-implemented method to reduce the number of rhizobia strains to be tested to formulate a customized inoculant formulation for a plant of interest grown in a soil of interest, wherein said method comprises steps a. to d. of the method for providing a customized inoculant formulation for a plant grown in a soil as described herein, and testing the rhizobia predicted to have high effectiveness and / or competitiveness for the combination of said soil of interest and plant of interest.

[0021] Description of Drawings

[0022] Figure 1. Nodules detached from roots are placed in a 96-well plate. The plate is inserted into a scanner with a custom-made plate holder. A scan is obtained at high resolution, nodules and visible inside the wells. These images and input to i) Fiji Imaged or ii) llastik software. In i), nodules are manually marked, then area in pixels quantified along with x and y coordinates (position in a picture). In ii), software is trained to recognize nodules and nodule size quantification is done by image segmentation. First, a training set is processed, where two kinds of objects (“nodule” and “background”) are assigned. Objects are then segmented, area in pixels quantified along with x and y coordinates.

[0023] Figure 2. A 3D-printed custom-designed 12-pestle device compatible with 96-well plates. Figure 3. Heatmap shows the Iog2-transformed relative abundances (in 10,000) of the rhizobia strains (y-axis) through the faba bean cultivars (x-axis).

[0024] Figure 4. Abundance - occurrence plot depicts the colonisation profile of the rhizobia strains. The average relative abundance (in 10,000) of each rhizobia strain is shown on the Iog2-scaled x-axis. The number of the samples that an isolate is present is on the y- axis. The data points are in respect to their colonization profile.

[0025] Figure 5. Abundance - occurrence plot depicts the coloured by the isolates’ soil origin, strains. The average relative abundance (in 10,000) of each rhizobia strain is shown on the Iog2-scaled x-axis. The number of the samples that an isolate is present is on the y- axis. The data points are in with respect to the soil of origin of the isolates

[0026] Figure 6. Boxplot depicting the niche breadth of the isolates with their respective colonization profile. Boxplot whiskers extend 1.5 times the interquartile range from the upper and lower quartiles. The data points were also shown, and they reflected the soil origin of the isolates.

[0027] Figure 7. Co-occurrence plots show the presence-presence based interactions among the rhizobia isolates. Here, the isolates (circles) are coloured with respect to their soil origin and sized according to their number of interactions. The line between a pair of isolates denotes that they were present in many samples while the number of samples in which only one of them occurred is small.

[0028] Figure 8. Mutual exclusion plots show the presence-absence based interaction among the rhizobia isolates. The layout of these is the same as the co-occurrence plots (a, e, and i), the only difference is that the lines have a direction, the end of the line with an arrowhead denotes the isolate on that side is competitively inferior in comparison to the one at the other side of the line.

[0029] Figure 9. Co-abundance profiles of the isolates are shown with the histograms of the Pearson correlation between the pairs of the isolates. Pearson correlations were based on the relative abundances (in 10,000).

[0030] Figure 10. DOC analyses depict the relationship between community overlap and dissimilarity. The DOC curves are represented as the blue lines with their 95% confidence intervals in yellow. The vertical dashed line indicates the point where a negative DOC is first observed. The distribution density of sample pair overlap is in grey. Dissimilarity is based on Jensen-Shannon distance.

[0031] Figure 11. The estimates from LMMs shown with their 95% confidence intervals and statistical significance. An asterisk denotes the effect was statistically significant at FDR < 0.05.

[0032] Figure 12. The estimates from LMMs showing the effect of the isolates on the plant growth. The estimates were depicted with respect to an isolates’ group. Boxplot whiskers extend 1.5 times the interquartile range from the upper and lower quartiles.

[0033] Figure 13. The relationship between an isolate’s niche breadth and its effect on the plant growth. We fitted linear models separately on the colonization profile using the ‘geom_smooth’ function from the R package ‘ggplot2’.

[0034] Figure 14. Estimates from LMM showing the effect of the evenness of each group on the plant growth shown with their 95% confidence intervals. An asterisk denotes the effect was statistically significant at P < 0.05.

[0035] Figure 15. Prediction scores (R2values) from 100 random forests models to evaluate the contribution of the evenness on the variation on the plant biomass. Boxplot whiskers extend 1.5 times the interquartile range from the upper and lower quartiles.

[0036] Figure 16. Prediction scores (R2values) from 1000 random forests models to evaluate the contribution of the plant genetics on the variation on the Rhizobium evenness. Boxplot whiskers extend 1.5 times the interquartile range from the upper and lower quartiles.

[0037] Figure 17. Results of a field trial carried out in Germany in 2023. The yield per plot in kilograms is shown for faba bean plants treated with two different rhizobia inoculants (Formulation 1 and Formulation 2) compared to an uninoculated control (control).

[0038] Figure 18. Results the benchmarking of a customized rhizobia biofertilizer against a generic rhizobia biofertilizer. Shoot length.

[0039] Figure 19. Results the benchmarking of a customized rhizobia biofertilizer against a generic rhizobia biofertilizer. Dry biomass. Detailed description

[0040] Definitions

[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of skill in art to which the subject matter herein belongs. As used herein, the following definitions are supplied to facilitate the understanding of the present invention.

[0042] The term “comprise” is generally used in the sense of include, that is to say permitting the presence of one or more features or components. In addition, as used in the specification and claims, the language "comprising" can include analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of”.

[0043] As used in the specification and claims, the term "and / or" used in a phrase such as "A and / or B" herein is intended to include "A and B", "A or B", "A", and "B".

[0044] As used in the specification and claims, the singular forms "a", "an" and "the" include plural references unless the context clearly dictates otherwise. Similarly, terms such as “one or more” or “at least one” include both the singular and plural form of the respective feature.

[0045] It will be evident to the skilled person that the different sections of the present disclosure are provided to facilitate understanding and clarity. However, the teachings and embodiments described in these sections are not intended to be read in isolation. Rather, they are to be considered in conjunction with one another, and the various features, aspects, and embodiments disclosed herein are intended to be combined in any manner that would be understood as technically feasible and beneficial by the skilled person.

[0046] Customized inoculant formulation

[0047] The present disclosure is directed to a customized inoculant formulation for a plant grown in a soil, and methods to provide and manufacture thereof. The customized inoculant formulation of the present disclosure comprises at least one rhizobia strain.

[0048] Thus, in some embodiments, the customized inoculant formulation enhances the growth and development of the plant in the specified soil conditions by optimizing the interaction between the plant and soil microbiota. In some embodiments, the rhizobia strain included in the inoculant formulation promotes plant growth by facilitating nitrogen fixation, improving nutrient uptake, or enhancing root development. In some embodiments, the customized inoculant formulation promotes an increase in the total biomass of the plant, including both above-ground and below-ground components. In some embodiments, the rhizobia strain promotes an increase in the total biomass of the plant independently or in combination with other microbial or chemical agents.

[0049] In some embodiments, the plant is as described herein, such as the plant described in the section “the plant” of the present disclosure.

[0050] In some embodiments, the soil is as described herein, such as the soil described in the section “the soil” of the present disclosure.

[0051] Thus, in some embodiments, the customized inoculant formulation promotes the growth of said plant in said soil. In some embodiments, said rhizobia strain promotes the growth of said plant. In some embodiments, the customized inoculant formulation promotes an increase in the total biomass of said plant. In some embodiments, said rhizobia strain promotes an increase in the total biomass of said plant.

[0052] It is possible to test in a controlled environment whether the customized inoculant formulation and / or the preferred rhizobia described herein promotes plant growth. Such test comprises comparing plants, which have been contacted with the customized inoculant formulation and / or the preferred rhizobia, to plants which have not been contacted with customized inoculant formulation and / or the preferred rhizobia. Any suitable method to assess plant growth known in the art can be used.

[0053] In some embodiments, the plant growth is assessed by one or more of the methods selected from: i. Measurement of the heigh and / or length of the plant; ii. Measurement of the biomass of the plant, such as the dry or fresh biomass; iii. Measurement of the leaf area; iv. Measurement of the root length; v. Measurement of the root mass; vi. Counting of the flowers produced by the plant; vii. Counting of the fruits produced by the plant; viii. Measurement of the Chlorophyll content; or ix. Measurement of the stem diameter.

[0054] In some embodiments, the plant growth is assessed by measuring the height and / or length of the plant. In some embodiments, the plant growth is assessed by measuring the biomass of the plant. In some embodiments, the biomass of the plant is the dry biomass. In some embodiments, the biomass of the plant is the fresh biomass. In some embodiments, the plant growth is assessed by measuring the leaf area. In some embodiments, the plant growth is assessed by measuring the root length. In some embodiments, the plant growth is assessed by measuring the root mass. In some embodiments, the root mass is the dry root mass. In some embodiments, the root mass is the fresh root mass. In some embodiments, the plant growth is assessed by counting the flowers produced by the plant. In some embodiments, the plant growth is assessed by counting the fruits produced by the plant. In some embodiments, the plant growth is assessed by measuring the chlorophyll content. In some embodiments, the plant growth is assessed by measuring the stem diameter.

[0055] In some embodiments, the plant growth is assessed by measuring the total biomass of said plant. In some embodiments, the total biomass is the total dry biomass of the plant. In some embodiments, the total biomass is the total fresh biomass of the plant.

[0056] It will be evident to the skilled person that, in some embodiments, any combination of the parameters described herein might be used.

[0057] The plant

[0058] The method described herein allows the skilled person to provide a customized inoculant formulation for a specific plant grown in a specific soil by providing the information / data of said soil and said plant as described herein. Thus, in some embodiments, the plant for which the customized inoculant formulation is provided is representative of the plants to be grown in the soil.

[0059] In some embodiments, the customized inoculant formulation can be effectively used to promote the growth of different plants, e.g., plants of the same species but different cultivar.

[0060] The soil

[0061] The method described herein allows the skilled person to provide a customized inoculant formulation for a specific plant grown in a specific soil by providing the information / data of said soil and said plant as described herein.

[0062] Several parameters of a soil are affected by seasonality and by which type of crops have been grown in that soil previously. Thus, preferably, the information / data provided for the soil is indicative of the condition of the soil in which the plant will be grown.

[0063] Thus, in some embodiments, the soil has been obtained from a field before the plant of interest has been sown in the soil of interest.

[0064] In some embodiments, the soil is representative of the soils of an area of interest, and / or is representative of soils which share the same physical / chemical characteristic. Thus, in some embodiments, the customised inoculant formulation can be effectively applied in one or more soils which share, approximately, the same physical / chemical characteristic with the soil of interest described herein.

[0065] In some embodiments, the soil has been obtained no more than 5 months prior to the sowing of the plant, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, such as 3 days, such as 1 day prior to the sowing of the plant.

[0066] In some embodiments, the soil has been obtained no later than 4 months after to the sowing of the plant, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, such as 3 days, such as 1 day after to the sowing of the plant.

[0067] In some embodiments, the soil has been obtained between 5 months prior to the sowing of the plant and 5 months after sowing of the plant, between 4 months prior to the sowing of the plant and 4 months after sowing of the plant, such as 3 months prior to the sowing of the plant and 3 months after sowing of the plant, such as 2 months prior to the sowing of the plant and 2 months after sowing of the plant, such as 1 month prior to the sowing of the plant and 1 months after sowing of the plant, such as 3 weeks prior to the sowing of the plant and 3 weeks after sowing of the plant, such as 2 weeks prior to the sowing of the plant and 2 weeks after sowing of the plant, such as 1 week prior to the sowing of the plant and 1 week after sowing of the plant, such as 3 days prior to the sowing of the plant and 3 days after sowing of the plant, such as 1 day prior to the sowing of the plant prior to the sowing of the plant and 1 day after sowing of the plant.

[0068] Information / Data provided

[0069] The present disclosure is directed to a customized inoculant formulation for a specific plant grown in a specific soil. Thus, a necessary step in a method to provide or manufacture such customized inoculant formulation comprises providing information / data of the plant of interest, as well as providing information / data of the soil of interest.

[0070] It will be evident to the skilled person that the providing information / data as described herein is meant to encompass providing information / data to a computer carrying out the computer-implemented steps of the method described herein.

[0071] Likewise, in order to train the machine learning models described herein information / data should be provided: a. for each rhizobia strain used to train the model; b. for each plant used to train the model; c. for each soil used to train the model; and d. for each plant and soil combination used to train the model.

[0072] It will be evident to the skilled person that the associated plant and soil data provided to train the machine learning models, and the plant and soil information / data provided in step a. and b. of the method described herein may consists, in some embodiments, of the same parameters.

[0073] Plant data

[0074] When performing the method described herein, in step a. information / data of the plant of interest is provided. When performing the method described herein, in step c. and d. information / data of each plant of the plurality of plant is provided.

[0075] In steps a., c. and d. is provided at least the information / data of the plant species.

[0076] In some embodiments, providing information / data of a plant further comprises providing one or more of the information / data selected from: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; or d. the phenotype of the plant.

[0077] Thus, in some embodiments, providing information / data of a plant further comprises providing the following information / data: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; or d. the phenotype of the plant.

[0078] In some embodiments, providing information / data of a plant further comprises providing the cultivar of the plant.

[0079] In some embodiments, providing information / data of a plant further comprises providing the genus of the plant. In some embodiments, providing information / data of a plant further comprises providing the genotype of the plant. The skilled person will understand that the genotype of a plant can be provided in any suitable format known in the art.

[0080] In some embodiments, the genotype of a plant is provided / obtained by a method selected from: a list of SNPs generated by array hybridization; targeted sequencing e.g. using SPET or other oligonucleotide-based enrichment; or whole-genome resequencing.

[0081] In some embodiments, the genotype of a plant is provided / obtained by a list of SNPs generated by array hybridization. In some embodiments, the genotype of a plant is provided / obtained by targeted sequencing, such as SPET or other oligonucleotide- based enrichment methods. In some embodiments, the genotype of a plant is provided / obtained by whole-genome resequencing.

[0082] In some embodiments, the list of SNPs comprises at least 100.000 SNPs, such as at least 200.000 SNPs, such as at least 300.000 SNPs, such as at least 400.000 SNPs, such as at least 500.000 SNPs.

[0083] In some embodiments, providing information / data of a plant further comprises providing the phenotype of the plant. The skilled person will understand that the phenotype of the plant can be provided in any format known in the art.

[0084] In some embodiments, providing the phenotype of the plant comprises providing one or more of the data selected from: a. plant height; b. yield; c. flowering time of the plant; b. seed size; c. shape of seeds and plant branching structure; d. colour of flowers and seeds; e. insect tolerance; or f. pest tolerance. Thus, in some embodiments, providing the phenotype of the plant comprises providing the plant height. In some embodiments, providing the phenotype of the plant comprises providing the plant yield. In some embodiments, providing the phenotype of the plant comprises providing the flowering time of the plant. In some embodiments, providing the phenotype of the plant comprises providing the seed size. In some embodiments, providing the phenotype of the plant comprises providing the shape of seeds and plant branching structure. In some embodiments, providing the phenotype of the plant comprises providing the colour of flowers and seeds. In some embodiments, providing the phenotype of the plant comprises providing the pest tolerance, such as the insect tolerance.

[0085] Soil data

[0086] When performing the method described herein, in step b. providing the information / data of the soil of interest comprises at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and iv. information / data of one or more rhizobia present in the soil.

[0087] It will be evident to the skilled person that the pH of the soil, the concentration of Nitrogen in the soil, the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and the information / data of one or more naturally occurring rhizobia present in the soil can be determined by any method known in the art, e.g. such as those described herein.

[0088] For example, the Most Probable Number (MPN) of compatible strains in the soil is a statistical estimation of the population density of viable microorganisms (such as rhizobia or other beneficial microbial strains) capable of forming a specific interaction, such as nodulation with a particular host plant, under given conditions.

[0089] Thus, the skilled person will understand that any method that allows to quantify or estimate the abundance / density of viable microorganism might be used. Thus, in performing the method described herein, in step b. providing the information / data of the soil of interest comprises at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the Most Probable Number (MPN) of compatible strains in the soil; and iv. information / data of one or more rhizobia present in the soil.

[0090] In some embodiments, providing the information / data of one or more naturally occurring rhizobia present in the soil comprises providing the DNA fingerprint and / or the DNA sequence of the one or more naturally occurring rhizobia present in the soil. The DNA fingerprint and / or the DNA sequence of can be determined by any method known in the art.

[0091] In some embodiments, providing information / data of the soil further comprises providing one of more of the soil parameter selected from: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil; f. the soil class; or g. the place of sampling.

[0092] In some embodiments, providing information / data of the soil further comprises providing: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil; f. the soil class; and g. the place of sampling.

[0093] In some embodiments, the information / data of the soil further comprises: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil; f. the soil class; and / or g. the place of sampling.

[0094] In some embodiments, providing information / data of the soil further comprises providing the concentration of one or more naturally occurring rhizobia strain in the soil. In some embodiments, providing information / data of the soil further comprises providing the concentration of phosphorus in the soil. In some embodiments, providing information / data of the soil further comprises providing the concentration of potassium in the soil. In some embodiments, providing information / data of the soil further comprises providing the concentration of magnesium in the soil.

[0095] In some embodiments, the information / data of the soil further comprises the concentration of one or more naturally occurring rhizobia strain in the soil. In some embodiments, the information / data of the soil further comprises the concentration of phosphorus in the soil. In some embodiments, the information / data of the soil further comprises the concentration of potassium in the soil. In some embodiments, the information / data of the soil further comprises the concentration of magnesium in the soil.

[0096] A person skilled in the art will understand that the concentration of phosphorus, potassium, and / or magnesium in the soil can be calculated and expressed by any suitable method known in the art.

[0097] In some embodiments, providing information / data of the soil further comprises providing the percentage of organic material in the soil. In some embodiments, the information / data of the soil further comprises the percentage of organic material in the soil. The skilled person will understand that this can be determined by any suitable method known in the art. In some embodiments, the percentage of organic material in the soil is determined by dry combustion, see e.g., DIN EN 15936. In some embodiments, the percentage of organic material in the soil is expressed as total organic carbon (TOC).

[0098] In some embodiments, providing information / data of the soil further comprises providing the soil class. In some embodiments, the information / data of the soil further comprises the soil class. The skilled person will understand that any suitable soil classification known in the art can be used. In some embodiments, the soil classification is the Danish JB classification (1-12). Information regarding the Danish soil classification can be found in the book Atlas over Danmark (Series 1 , Volume 3), which is also available online e.g., https: / / rdgs.dk / publikationer / atlas-of-denmark-serie- 1-bind-3_-danish-soil-classification.pdf

[0099] In some embodiments, providing information / data of the soil further comprises providing the place of sampling. In some embodiments, the information / data of the soil further comprises the place of sampling.

[0100] In some embodiments, providing the place of sampling comprises providing the coordinates of the place of sampling, preferably wherein providing the information of the place of sampling comprises providing the GPS coordinates of the place of sampling.

[0101] In some embodiments, the information / data of the place of sampling comprises the coordinates of the place of sampling, preferably wherein information / data of the place of sampling comprises providing the GPS coordinates of the place of sampling.

[0102] In some embodiments, the coordinates of the place of sampling are GPS coordinates.

[0103] In some embodiments, the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil, is determined by a method comprising growing said plant in said soil in a speed breeding growth chamber.

[0104] The term “speed breeding growth chamber” as used herein refers to a controlled environment designed to accelerate the growth of plants or organisms. Such chambers, also known as growth chambers or plant growth chambers, are known in the art to provide optimal conditions for plant growth, such as temperature, humidity, light, and nutrient levels. By controlling these factors, it is possible to study the effects of various conditions on plant development and optimize crop growth.

[0105] In some embodiments, the plant is grown in conditions that accelerate nodulation.

[0106] Thus, in some embodiments, the plant is grown in a photoperiod between 18.5 and

[0107] 22.5 hours, such as between 18.5 and 22 hours, such as between 18.5 and 21.5 hours, such as between 18.5 and 21 hours, such as between 18.5 and 20.5 hours, such as between 18.5 and 20 hours, such as between 18.5 and 19.5 hours, such as between

[0108] 18.5 and 19 hours, such as between 19 and 22.5 hours, such as between 19 and 22 hours, such as between 19 and 21.5 hours, such as between 19 and 21 hours, such as between 19 and 20.5 hours, such as between 19 and 20 hours, such as between 19 and 19.5 hours, such as between 19.5 and 22.5 hours, such as between 19.5 and 22 hours, such as between 19.5 and 21.5 hours, such as between 19.5 and 21 hours, such as between 19.5 and 20.5 hours, such as between 19.5 and 20 hours, such as between 20 and 22.5 hours, such as between 20 and 22 hours, such as between 20 and 21.5 hours, such as between 20 and 21 hours, such as between 20 and 20.5 hours, such as between 20.5 and 22.5 hours, such as between 20.5 and 22 hours, such as between 20.5 and 21.5 hours, such as between 20.5 and 21 hours, such as between 21 and 22.5 hours, such as between 21 and 22 hours, such as between 21 and 21.5 hours, such as between 21.5 and 22.5 hours, such as between 21.5 and 22 hours, such as between 22 and 22.5 hours.

[0109] In some embodiments, the plant is grown at a temperature between 16 and 24°C, such as between 16 and 23°C, such as between 16 and 22°C, such as between 16 and 21°C, such as between 16 and 20°C, such as between 16 and 19°C, such as between 16 and 18°C, such as between 16 and 17°C, such as between 17 and 24°C, such as between 17 and 23°C, such as between 17 and 22°C, such as between 17 and 21 °C, such as between 17 and 20°C, such as between 17 and 19°C, such as between 17 and 18°C, such as between 18 and 24°C, such as between 18 and 23°C, such as between 18 and 22°C, such as between 18 and 21 °C, such as between 18 and 20°C, such as between 18 and 19°C, such as between 19 and 24°C, such as between 19 and 23°C, such as between 19 and 22°C, such as between 19 and 21°C, such as between 19 and 20°C, such as between 20 and 24°C, such as between 20 and 23°C, such as between 20 and 22°C, such as between 20 and 21°C, such as between 21 and 24°C, such as between 21 and 23°C, such as between 21 and 22°C, such between 22 and 24°C, such as between 22 and 23°C, such as between 23 and

[0110] In some embodiments, the plant is grown in a photoperiod between 50 and 80° / humidity, such as between 50 and 75% humidity, such as between 50 and 70% humidity, such as between 50 and 65% humidity, such as between 50 and 60% humidity, such as between 50 and 55% humidity, such as between 55 and 80% humidity, such as between 55 and 75% humidity, such as between 55 and 70% humidity, such as between 55 and 65% humidity, such as between 55 and 60% humidity, such as between 60 and 80% humidity, such as between 60 and 75% humidity, such as between 60 and 70% humidity, such as between 60 and 65% humidity, such as between 65 and 80% humidity, such as between 65 and 75% humidity, such as between 65 and 70% humidity, such as between 70 and 80% humidity, such as between 70 and 75% humidity, such as between 75 and 80% humidity.

[0111] Machine learning model

[0112] The term “machine learning” (ML) as used herein refers to a computer algorithm used to extract useful information from training data sets by building probabilistic frameworks, referred to as machine learning frameworks (also referred to as machine learning models), in an automated way. The machine learning may be performed using one or more learning algorithms such as linear regression, K-means, classification algorithm, reinforcement algorithm, etc. A “machine learning model” may for example be an equation or set of rules that makes it possible to predict an unmeasured from other, known values and / or to predict or select an action to maximize a future reward.

[0113] For example, the learning algorithm may comprise classification and / or reinforcement algorithms. The reinforcement algorithm may for example be configured to learn one or more policies or rules for determining a next set of parameters (action) based on the current set of parameters and / or previously used set of parameters. For example, starting from the current set of parameters and / or previous set of acquisition parameters the machine learning framework may follow a policy until it reaches a desired set of acquisition parameters. The policy represents the decision-making process of the model at each step e.g. it defines which parameter to change and how to change it, which new parameter to add to the set of parameters etc.

[0114] In some embodiments, the machine learning model is a deep learning model.

[0115] Deep learning is a subfield of machine learning that focuses on training artificial neural networks to learn and solve complex problems by mimicking somewhat the structure and function of the human brain. It is a form of representation learning, where the system automatically learns to identify patterns and features from data without relying on explicit programming.

[0116] The term "deep" in deep learning refers to the use of multiple layers in neural networks, often called "deep neural networks." These layers allow the network to progressively learn higher-level abstractions and representations from raw input data, making it capable of handling intricate and hierarchical patterns. Each layer in the network processes the output from the previous layer, leading to a cascade of transformations that eventually yield a final output or prediction.

[0117] In some embodiments, the machine learning model is trained via supervised learning.

[0118] Supervised learning involves training a model on a labelled dataset, where each data instance has an associated target or label that the model aims to predict. The goal is to learn a mapping between input features and corresponding output labels. During training, the model adjusts its parameters to minimize the difference between its predictions and the true labels. Examples of supervised learning tasks include image classification, object detection, sentiment analysis, and speech recognition.

[0119] It will be evident to the skilled person that steps described herein involving a prediction by machine learning models typically are carried out by a computer.

[0120] The term “known rhizobia strain” as used herein refers to a rhizobia bacteria which, in connection with its associated rhizobia strain data, has been used to train the machine learning model described herein to select an ideal rhizobia strain. Thus, the machine learning model predicts that one or more rhizobia strain selected from the plurality of known rhizobia is able to promote the growth of the plant of interest. The term “ideal rhizobia strain” as used herein refers to a rhizobia bacteria which has been shown to promote the growth of a plant, such as to promote the increase in total biomass of said plant, when grown in a specific soil.

[0121] The term “preferred rhizobia strain” as used herein refers to a rhizobia bacteria which has been selected following the prediction of the machine learning model and, optionally, a validation step.

[0122] The term “associated” as used herein refers to one or more information / data, such as one or more parameter, which define a particular entry. For example, the sentence “associated" rhizobia strain data for each known rhizobia strain refers to one or more information / data, such as one or more parameter which define each known rhizobia strain used to train the machine learning model described herein. Similarly, the sentence “associated" plant data for each plant refers to one or more information / data, such as one or more parameter which define each plant used to train the machine learning model described herein. One such associated plant data is e.g., said plant species. Likewise, the sentence “associated" soil data for each soil refers to one or more information / data, such as one or more parameter which define each soil used to train the machine learning model described herein. One such associated soil data is e.g., the pH of said soil.

[0123] The present disclosure describes two machine learning model: a machine learning model which predicts rhizobium competitiveness; a machine learning model which predicts rhizobium effectiveness.

[0124] The term “rhizobium competitiveness” as described herein refers to the ability of a given rhizobia strain to occupy a given plant genotype nodules (nodule occupancy, Yocc) in a given soil environment with multiple strains present.

[0125] The term “rhizobium effectiveness” (Ye / r) as described herein refers to the ability of a given rhizobia strain to acquire nitrogen through the biological nitrogen fixation and / or their effect on plant biomass production in a given plant genotype in a given soil environment. In some embodiments, rhizobium effectiveness expresses to the ability of a given rhizobia strain to acquire nitrogen through the biological nitrogen fixation. In some embodiments, rhizobium effectiveness expresses the effect of a rhizobia strain on plant biomass production in a given plant genotype in a given soil environment. In some embodiments, rhizobium effectiveness expresses to the ability of a given rhizobia strain to acquire nitrogen through the biological nitrogen fixation and the effect of said rhizobia strain on plant biomass production in a given plant genotype in a given soil environment.

[0126] The machine learning model described herein is trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant of the plurality of plants comprising at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality comprising at least: the pH of the soil; and the concentration of Nitrogen in the soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; and information / data of one or more rhizobia present in said soil; iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants as described herein; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils as described herein; and

[0127] When training the machine learning model which predicts rhizobium competitiveness, step iv. further comprises providing for each combination of plant and soil the rhizobium competitiveness of each rhizobia strain selected from the from the known rhizobia strains in i. When training the machine learning model which predicts rhizobium effectiveness, step iv. further comprises providing for each combination of plant and soil the rhizobium effectiveness of each rhizobia strain selected from the from the known rhizobia strains in i.

[0128] It will be evident to the skilled person that the associated plant and soil data provided to train the machine learning model, and the plant and soil information / data provided in step a. and b. of the method described herein may consists, in some embodiments, of the same parameters.

[0129] Further, it will be evident to the skilled person that step c and d might be performed sequentially or in parallel. Thus, in any of the methods described herein step c. might be performed before, after, or at the same time as step d.

[0130] In some embodiments, step c is performed before step d, and step d comprises: d. predicting, by a machine learning model, the effectiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or the one or more rhizobia strains predicted in step c. to have high competitiveness,

[0131] In some embodiments, step d is performed before step c, and step c comprises: d. predicting, by a machine learning model, the competitiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or the one or more rhizobia strains predicted in step c. to have high effectiveness,

[0132] Thus, one of the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant of the plurality of plants comprising at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; and information / data of at least one rhizobia present in said soil; iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the competitiveness of each rhizobia strain selected from the known rhizobia strains in i.;

[0133] Thus, one the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, , and associated plant data for each plant of the plurality of plants comprising at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia present in said soil; and iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii; and for each combination of plant and soil is provided the effectiveness of each rhizobia strain selected from the known rhizobia strains in i.;

[0134] In some embodiments, the machine learning model in step d. predicts the effectiveness of each rhizobia strain selected from a plurality of known rhizobia strains.

[0135] In some embodiments, the machine learning model in step d. predicts the effectiveness of each rhizobia strain selected from the one or more rhizobia strains predicted in step c. to have high competitiveness.

[0136] Thus, in some embodiments, in step d. the machine learning model predicts the effectiveness of each rhizobia strain selected from one or more rhizobia strains predicted in step c. to have high competitiveness.

[0137] In some embodiments, the one or more rhizobia strains predicted in step c. to have high competitiveness, comprises at least 5 rhizobia strains, such as at least 10 rhizobia strains, such as at least 15, at least 20, at least 30, at least 40, at least 50 rhizobia strains.

[0138] Plurality of known rhizobia strains

[0139] To train the machine learning models described herein, information / data regarding a plurality of known rhizobia strain is provided, as well as information / data regarding a plurality of plants, a plurality of soils, and a plurality of combinations of plant and soil.

[0140] In some embodiments, the associated rhizobia strain data for each known rhizobia strain comprises at least: i. said known rhizobia strain; and ii. its full or partial genomic sequence or genomic fingerprint;

[0141] For each known rhizobia, it is further provided a. its competitiveness in a plurality of combinations of plant and soil; and b. its effectiveness in a plurality of combinations of pland and soil.

[0142] In some embodiments, the plurality of known rhizobia strain comprises at least 200 known rhizobia strain, such as at least 300, such as at least 400, such as at least 500, such as at least 600, such as at least 700, such as at least 800, such as at least 900, such as at least 1000 known rhizobia strain.

[0143] In some embodiments, the associated rhizobia strain data comprises one or more of the data selected from: a. rhizobia strain genera; b. rhizobia strain species; b. rhizobia strain subspecies; c. rhizobia strain biovar; or d. rhizobia strain type or number.

[0144] In some embodiments, the associated rhizobia strain data comprises: a. rhizobia strain genera; b. rhizobia strain species; c. rhizobia strain subspecies; d. rhizobia strain biovar; and e. rhizobia strain type or number.

[0145] In some embodiments, the associated rhizobia strain data comprises rhizobia strain genera. In some embodiments, the associated rhizobia strain data comprises rhizobia strain species. In some embodiments, the associated rhizobia strain data comprises rhizobia strain subspecies. In some embodiments, the associated rhizobia strain data comprises rhizobia strain biovar. In some embodiments, the associated rhizobia strain data comprises rhizobia strain type. In some embodiments, the associated rhizobia strain data comprises rhizobia strain number.

[0146] It will be evident to the skilled person that, in some embodiments, the plurality of known rhizobia strain described in step c. and d. of the method described herein are the same.

[0147] To train the machine learning models described herein, information / data regarding a plurality of plants is provided, as well as information / data regarding a plurality of known rhizobia strains, a plurality of soils, and a plurality of combinations of plant and soil. The associated plant data provided for each plant comprises at least said plant species.

[0148] In some embodiments, the plurality of plants comprises at least 100 plants, such as 150 plants, such as 200 plants, such as 250 plants, such as 300 plants, such as 350 plants, such as 400 plants, such as 450 plants, such as at least 500 plants.

[0149] In some embodiments, the associated plant data provided for each plant comprises one or more of the information / data selected from: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; or d. the phenotype of the plant.

[0150] In some embodiments, the associated plant data provided for each plant comprises following information / data: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; and d. the phenotype of the plant.

[0151] In some embodiments, the associated plant data provided for each plant comprises the cultivar of the plant. In some embodiments, the associated plant data provided for each plant comprises the genus of the plant.

[0152] In some embodiments, the associated plant data provided for each plant comprises the genotype of the plant. The skilled person will understand that the genotype of the plant can be provided in any known format in the art.

[0153] In some embodiments, the genotype of the plant is provided as a list of SNPs. In some embodiments, the list of SNPs comprises at least 100.000 SNPs, such as at least 200.000 SNPs, such as at least 300.000 SNPs, such as at least 400.000 SNPs, such as at least 500.000 SNPs.

[0154] In some embodiments, the associated plant data provided for each plant comprises the phenotype of the plant. The skilled person will understand that the phenotype of the plant can be provided in any known format in the art.

[0155] In some embodiments, providing the phenotype of the plant comprises providing one or more of the data selected from: a. plant height; b. yield; c. flowering time of the plant; d. seed size; e. shape of seeds and plant branching structure; f. color of flowers and seeds; g. insect tolerance; or h. pest tolerance.

[0156] In some embodiments, providing the phenotype of the plant comprises providing: a. plant height; b. yield; c. flowering time of the plant; d. seed size; e. shape of seeds and plant branching structure; f. color of flowers and seeds; g. insect tolerance; and h. pest tolerance.

[0157] In some embodiments, the phenotype of the plant comprises one or more of: a. plant height; b. yield; c. flowering time of the plant; d. seed size; e. shape of seeds and plant branching structure; f. color of flowers and seeds; g. insect tolerance; or h. pest tolerance.

[0158] In some embodiments, the phenotype of the plant comprises: a. plant height; b. yield; c. flowering time of the plant; d. seed size; e. shape of seeds and plant branching structure; f. color of flowers and seeds; g. insect tolerance; h. pest tolerance.

[0159] In some embodiments, providing the phenotype of the plant comprises providing the plant height. In some embodiments, providing the phenotype of the plant comprises providing the plant yield In some embodiments, providing the phenotype of the plant comprises providing the flowering time of the plant. In some embodiments, providing the phenotype of the plant comprises providing the seed size. In some embodiments, providing the phenotype of the plant comprises providing the shape of seeds and plant branching structure. In some embodiments, providing the phenotype of the plant comprises providing the colour of seeds and flowers. In some embodiments, providing the phenotype of the plant comprises providing the pest tolerance, such as the insect tolerance.

[0160] It will be evident to the skilled person that, in some embodiments, the plurality of plants in step c. and d. of the method described herein are the same. of soils

[0161] To train the machine learning model described herein, information / data regarding a plurality of soils is provided, as well as information / data regarding a plurality of known rhizobia strains, a plurality of plants, and a plurality of combinations of plant and soil. The associated soil data provided for each soil comprises at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and iv. information / data one or more naturally occurring rhizobia present in the soil

[0162] It will be evident to the skilled person that the pH of the soil, the concentration of Nitrogen in the soil, the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and the information / data of one or more naturally occurring rhizobia present in the soil can be determined by any method known in the art, e.g. such as those described herein.

[0163] In some embodiments, providing the information / data of one or more naturally occurring rhizobia present in the soil comprises providing the DNA fingerprint and / or the DNA sequence of the one or more naturally occurring rhizobia present in the soil. The DNA fingerprint and / or the DNA sequence of can be determined by any method known in the art.

[0164] In some embodiments, the plurality of soils comprises at least 100 soils, such as 150 soils, such as 200 soils, such as 250 soils, such as 300 soils, such as 350 soils, such as 400 soils, such as 450 soils, such as at least 500 soils.

[0165] In some embodiments, the associated soil data provided for each soil comprises one of more of the soil parameter selected from: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil; f. the soil class; g. the place of sampling.

[0166] In some embodiments, the associated soil data provided for each soil further comprises: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil; f. the soil class; g. the place of sampling; and h. the historical and predicted temperature data for the place of sampling;

[0167] In some embodiments, the associated soil data provided for each soil further comprises providing the concentration of one or more naturally occurring rhizobia strain in the soil. In some embodiments, the associated soil data provided for each soil further comprises providing the concentration of phosphorus in the soil. In some embodiments, the associated soil data provided for each soil further comprises providing the concentration of potassium in the soil. In some embodiments, the associated soil data provided for each soil further comprises providing the concentration of magnesium in the soil.

[0168] A person skilled in the art will understand that the concentration of phosphorus, potassium, and / or magnesium in the soil can be calculated by any method known in the art. Similarly, the concentration may be expressed with any suitable measurement unit.

[0169] In some embodiments, the associated soil data provided for each soil further comprises the percentage of organic material in the soil. The skilled person will understand that this can be determined by any method known in the art. In some embodiment, the percentage of organic material in the soil is determined by dry combustion, see e.g., DIN EN 15936. In some embodiments, the percentage of organic material in the soil is expressed as total organic carbon (TOC).

[0170] In some embodiments, the associated soil data provided for each soil further comprises the soil class. The skilled person will understand that any soil classification known in the art can be used. In some embodiments, the soil classification is the Danish JB classification (1-12). Information regarding the Danish soil classification can be found in the book Atlas over Danmark (Series 1, Volume 3), which is also available online e.g., https: / / rdgs.dk / publikationer / atlas-of-denmark-serie-1-bind-3_-danish-soil- classification.pdf

[0171] In some embodiments, the associated soil data provided for each soil further comprises providing the place of sampling.

[0172] In some embodiments, the associated soil data provided for each soil further comprises the coordinates of the place of sampling, preferably wherein coordinates of the place of sampling are the GPS coordinates of the place of sampling . In some embodiments, the coordinates of the place of sampling are GPS coordinates.

[0173] It will be evident to the skilled person that, in some embodiments, the plurality of soils in step c. and d. of the method described herein are the same.

[0174] Combination of soils and plants

[0175] To train the machine learning model described herein, information / data regarding a plurality of combinations of plant and soil is provided, as well as information / data regarding a plurality of plants, a plurality of soils, and a plurality of known rhizobia strain. Said plurality of combinations of plant and soil, comprises: i. the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants as described herein; ii. the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils as described herein; and iii. for each combination of plant and soil is provided at least one ideal rhizobia strain, wherein the at least one ideal rhizobia strain is selected from the known rhizobia strains as described herein.

[0176] Said plurality of combinations of plant and soil, and associated data for each combination of plant and soil, comprises: i. the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants as described herein; ii. the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils as described herein; and iii. for each combination of plant and soil is provided at least one ideal rhizobia strain, wherein the at least one ideal rhizobia strain is selected from the known rhizobia strains as described herein.

[0177] It will be evident to the skilled person that, in some embodiments, the combination of plant and soil in step c. and d. of the method described herein are the same.

[0178] In some embodiments, the associated data for each combination of plant and soil, comprises for each combination of plant and soil at least one ideal rhizobia strain, wherein the at least one ideal rhizobia strain is selected from the known rhizobia strains as described herein.

[0179] In some embodiments, the plurality of combinations of plant and soil comprises at least 100 combinations, such as 150 combinations, such as 200 combinations, such as 250 combinations, such as 300 combinations, such as 350 combinations, such as 400 combinations, such as 450 combinations, such as at least 500 combinations.

[0180] In some embodiments, the plurality of combinations of plant and soil consists of a combination of each soil comprised in the plurality of soils provided for the training of the machine learning model and each plant comprised in the plurality of plants provided for the training of the machine learning model.

[0181] In some embodiments, for each combination of plant and soil the rhizobium it is provided the competitiveness of each rhizobia strain selected from the from the known rhizobia strains.

[0182] In some embodiments, for each combination of plant and soil the rhizobium it is provided the effectiveness of each rhizobia strain selected from the from the known rhizobia strains.

[0183] Validation step

[0184] At the present moment, the only viable approach to identify elite rhizobia strains which promote the growth of a plant in a specific soil involves conducting a field trial, which typically lasts 6 months, or controlled experiments in laboratory or greenhouse conditions. These experiments demand a considerable allocation of resources. Indeed, these studies often require a considerable number of plants to be used to test a large amount of rhizobia strains.

[0185] The term “elite rhizobia strain” as used herein refers to a highly efficient or effective strain of these bacteria, capable of promoting even better growth and nitrogen fixation in leguminous plants. Such strains are of great interest to agricultural researchers and farmers who seek to optimize crop yields and reduce the environmental impact of fertilizers.

[0186] The method described herein allows to predict, by a machine learning model, one or more known rhizobia strain able to promote the growth of a plant of interest grown in a soil of interest. Thus, the method disclosed herein allows a significant cut in costs to provide a customized inoculant formulation.

[0187] Thus, in an aspect, the present disclosure is directed to a method to reduce the number of rhizobia strains to be tested to formulate a customized inoculant formulation for a plant of interest grown in a soil of interest, wherein said method comprises steps a. to d. of the method for providing a customized inoculant formulation for a plant grown in a soil as described herein, and testing the rhizobia predicted to have high effectiveness and / or competitiveness for the combination of said soil of interest and plant of interest.

[0188] In some embodiments, it can be advantageous to validate the list of bacteria provided by the machine learning model, by inoculating the predicted rhizobia strain in said soil, and growing the plant of interest in said soil. Compared to typical field trials, where many rhizobia strain needs to be tested, the method described herein can be useful to significantly reduce the number of rhizobia strain to be tested, as only the rhizobia strain predicted by the machine learning model to promote the growth of the plant will be tested.

[0189] Thus, in an aspect, the present disclosure is directed to a method to reduce the number of rhizobia to be tested to formulate a customized inoculant formulation, and wherein said method comprises the method for providing a customized inoculant formulation for a plant grown in a soil as described herein throughout the present disclosure.

[0190] Thus, in some embodiments, the method described herein further comprises a validation step e. comprising the steps of: i. inoculating in the soil at least one of the one or more known rhizobia strains predicted to have high competitiveness in step c. and high effectiveness in step d.; ii. growing the plant in the soil; and iii. identifying rhizobia strains in the nodules of the plant.

[0191] Thus, in some embodiments, the method for providing a customized inoculant formulation for a plant grown in a soil comprising at least one preferred rhizobia strain comprises the steps of: a. providing information / data of the plant, comprising at least said plant species; b. providing information / data of the soil, said information / data comprising at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and iv. information / data one or more naturally occurring rhizobia present in the soil; c. predicting, by a machine learning model, the competitiveness of each rhizobia strain selected from a plurality of known rhizobia strains, wherein the competitiveness of each rhizobia strain is the ability to occupy the nodules of a plant in a given soil with multiple rhizobia strains present, and wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant comprising for each plant of the plurality of plants at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality of soils comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; and information / data of one or more rhizobia present in said soil; iv. information / data of a plurality of combinations of plant and soil, and associated data wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the competitiveness of each rhizobia strain selected from the known rhizobia strains in i.; d. predicting, by a machine learning model, the effectiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or one or more rhizobia strains predicted in step e. to have high competitiveness, wherein the effectiveness of a rhizobia strain is the ability to acquire nitrogen through the biological nitrogen fixation and / or its effect on plant biomass production in the plant grown in the soil, wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain; ii. information / data of a plurality of plants, and associated plant data for each plant comprising at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; and information / data of at least one rhizobia present in said soil; iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the effectiveness of each rhizobia strain selected from the known rhizobia strains in i.; e. a validation step comprising the steps of: i. inoculating in the soil at least one of the one or more known rhizobia strains predicted to have high competitiveness in step c. and high effectiveness in step d.; ii. growing the plant in the soil; and iii. identifying rhizobia strains in the nodules of the plant. f. selecting at least one preferred rhizobia strain g. formulating the customized inoculant comprising the at least one preferred rhizobia strain; thereby providing the customized inoculant formulation for said plant and soil.

[0192] Thus, in some embodiments, at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e.

[0193] In some embodiments, the validation step e. is performed in less than 5 months, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, preferably less than 3 months. In some embodiments, the result of the validation step e. is added to the information provided to train a machine learning model to be used in a method as described in any one of the preceding claims.

[0194] In some embodiments, at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e. with at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance.

[0195] In some embodiments, the at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e. which occur in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules analyzed.

[0196] In some embodiments, at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e. with at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance and the at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e. which occur in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules analyzed.

[0197] Inoculation of rhizobia strains

[0198] As described herein, it can be advantageous in some embodiments to validate the list of bacteria provided by the machine learning model, by inoculating the predicted rhizobia strain in said soil, and growing the plant of interest in said soil. This can be an useful strategy to significantly reduce the number of rhizobia strains to be tested in the validation step described herein. In some embodiments, the soil comprises at least one naturally occurring rhizobia strain. Thus, in some embodiments, the validation step e. of the method described herein allows to test whether the one or more rhizobia strain predicted in step c. and step d. of the method described herein are able to promote the growth of the plant of interest better than the rhizobia strains naturally occurring in the soil of interest.

[0199] In some embodiments, the one or more known rhizobia strain predicted in step c. and d. of the method described herein are inoculated in the soil during the validation e. described herein at a concentration of at least 1 x 10A5 cfu / ml , such at least 1 x 10A6 cfu / ml, such at least 1 x 10A7 cfu / ml, such at least 1 x 10A8 cfu / ml, such at least 1 x 10A9 cfu / ml, such at least 1 x 10A10 cfu / ml.

[0200] Growth of the plant during the validation step

[0201] The present application discloses a method which can be used to significantly reduce the number of rhizobia strains to be tested in the validation step described herein. Thanks to this limited amount of rhizobia strains to be tested, it becomes economically feasible, in some embodiments, to grow the plants at accelerated plant growth.

[0202] Thus, the validation step e. can be performed, in some embodiments, in less than 5 months, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, preferably less than 3 months.

[0203] In some embodiments, the plant is grown in a speed breeding growth chamber. The term “speed breeding growth chamber” as used herein refers to a controlled environment designed to accelerate the growth of plants or organisms. Such chambers, also known as growth chambers or plant growth chambers, are known in the art to provide optimal conditions for plant growth, such as temperature, humidity, light, and nutrient levels. By controlling these factors, it is possible to study the effects of various conditions on plant development and optimize crop growth.

[0204] In some embodiments, the plant is grown in conditions that accelerate nodulation.

[0205] Thus, in some embodiments, the plant is grown in a photoperiod between 18.5 and 22.5 hours, such as between 18.5 and 22 hours, such as between 18.5 and 21.5 hours, such as between 18.5 and 21 hours, such as between 18.5 and 20.5 hours, such as between 18.5 and 20 hours, such as between 18.5 and 19.5 hours, such as between 18.5 and 19 hours, such as between 19 and 22.5 hours, such as between 19 and 22 hours, such as between 19 and 21.5 hours, such as between 19 and 21 hours, such as between 19 and 20.5 hours, such as between 19 and 20 hours, such as between 19 and 19.5 hours, such as between 19.5 and 22.5 hours, such as between 19.5 and 22 hours, such as between 19.5 and 21.5 hours, such as between 19.5 and 21 hours, such as between 19.5 and 20.5 hours, such as between 19.5 and 20 hours, such as between 20 and 22.5 hours, such as between 20 and 22 hours, such as between 20 and 21.5 hours, such as between 20 and 21 hours, such as between 20 and 20.5 hours, such as between 20.5 and 22.5 hours, such as between 20.5 and 22 hours, such as between 20.5 and 21.5 hours, such as between 20.5 and 21 hours, such as between 21 and 22.5 hours, such as between 21 and 22 hours, such as between 21 and 21.5 hours, such as between 21.5 and 22.5 hours, such as between 21.5 and 22 hours, such as between 22 and 22.5 hours.

[0206] In some embodiments, the plant is grown at a temperature between 16 and 24°C, such as between 16 and 23°C, such as between 16 and 22°C, such as between 16 and 21°C, such as between 16 and 20°C, such as between 16 and 19°C, such as between 16 and 18°C, such as between 16 and 17°C, such as between 17 and 24°C, such as between 17 and 23°C, such as between 17 and 22°C, such as between 17 and 21 °C, such as between 17 and 20°C, such as between 17 and 19°C, such as between 17 and 18°C, such as between 18 and 24°C, such as between 18 and 23°C, such as between 18 and 22°C, such as between 18 and 21°C, such as between 18 and 20°C, such as between 18 and 19°C, such as between 19 and 24°C, such as between 19 and 23°C, such as between 19 and 22°C, such as between 19 and 21°C, such as between 19 and 20°C, such as between 20 and 24°C, such as between 20 and 23°C, such as between 20 and 22°C, such as between 20 and 21°C, such as between 21 and 24°C, such as between 21 and 23°C, such as between 21 and 22°C, such as between 22 and 24°C, such as between 22 and 23°C, such as between 23 and 24°C.

[0207] In some embodiments, the plant is grown in a photoperiod between 50 and 80% humidity, such as between 50 and 75% humidity, such as between 50 and 70% humidity, such as between 50 and 65% humidity, such as between 50 and 60% humidity, such as between 50 and 55% humidity, such as between 55 and 80% humidity, such as between 55 and 75% humidity, such as between 55 and 70% humidity, such as between 55 and 65% humidity, such as between 55 and 60% humidity, such as between 60 and 80% humidity, such as between 60 and 75% humidity, such as between 60 and 70% humidity, such as between 60 and 65% humidity, such as between 65 and 80% humidity, such as between 65 and 75% humidity, such as between 65 and 70% humidity, such as between 70 and 80% humidity, such as between 70 and 75% humidity, such as between 75 and 80% humidity.

[0208] Information / Data collected

[0209] In some embodiments, the validation step e. further comprises a step: i. quantifying one or more of the parameters selected from: a. the number of nodules per plant; b. the size of each nodule; c. proxy of nitrogen fixation in individual nodules d. the number of pods per plant; e. the pod weight; f. the total biomass of the plant.

[0210] Thus, in some embodiments, the validation step e. described herein comprises the steps of: i. inoculating in the soil at least one of the one or more known rhizobia strains predicted to have high competitiveness in step c. and high effectiveness in step d.; ii. growing the plant in the soil; iii. identifying rhizobia strains in the nodules of the plant; and iv. quantifying one or more of the parameters selected from: a. the number of nodules per plant; b. the size of each nodule; c. proxy of nitrogen fixation in individual nodules d. the number of pods per plant; e. the pod weight; f. the total biomass of the plant. These parameters can be used e.g., to help to identify one or more elite rhizobia stain.

[0211] In some embodiments, the validation step e. further comprises a step iv. comprising or consists of a step of quantifying number of nodules per plant.

[0212] In some embodiments, the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the size of each nodule.

[0213] In some embodiments, the validation step e. further comprises a step iv. comprising or consists of a step of quantifying proxy of nitrogen fixation in individual nodules.

[0214] In some embodiments, the method to quantify a proxy of nitrogen fixation in individual nodules is selected from: acetylene reduction assay (ARA), isotope 15N2 assimilation, isotopic acetylene reduction assay (ISARA), total nitrogen difference (TND), nodule size and / or biomass, or expression of a reporter, such as GFP, under the control of nif genes promoters. The skilled person will understand that any of the nif gene promoter can be effectively used to control the expression of a promoter.

[0215] In some embodiments, the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the number of pods per plant.

[0216] In some embodiments, the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the pod weight.

[0217] In some embodiments, the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the total biomass.

[0218] Identification of rhizobia strains in the nodules

[0219] It will be evident to the skilled person that any suitable method known in the art to identify rhizobia strains in a soil can be effectively used in the validation step described herein. In some embodiments, identifying rhizobia strains in the nodules of the plant comprises sequencing the rhizobia strains. In some embodiments, identifying rhizobia strains in the nodules of the plant comprises determining the DNA fingerprint and / or the DNA sequence of the one or more naturally occurring rhizobia present in the soil.

[0220] In some embodiments, identifying rhizobia strains in the nodules of the plant comprises sequencing an identifier unique for each rhizobia strain.

[0221] In some embodiments, the identifier unique for each rhizobia strain is selected from: a natural nucleotide sequence in the genome of the rhizobia strains; a synthetic nucleotide sequence in the genome of the rhizobia strains; a nucleotide sequence in the rRNA of the rhizobia strains; or a nucleotide sequence in a vector, such as a plasmid (extrachromosomal genome).

[0222] Selecting at least one preferred rhizobia strain

[0223] The present disclosure concerns a method that allows to select at least one preferred rhizobia strain for a specific soil and plant.

[0224] It will be evident to the skilled person that the step f. of the method described herein can be performed, in some embodiments, with a computer. Thus, in some embodiments, step f. is a computer-implemented step.

[0225] In some embodiments, the preferred rhizobia strain is selected in step f. from the one or more rhizobia predicted in to have high competitiveness in step c. and / or high effectiveness in step d. of the method described herein.

[0226] In some embodiments, the preferred rhizobia strain is selected in step f. from the one or more rhizobia predicted in to have high competitiveness in step c. and high effectiveness in step d. of the method described herein.

[0227] In some embodiments, at least one preferred rhizobia strain is selected from the rhizobia strains identified in the nodules of the plant during the validation step e. described herein. Thus, In some embodiments, the at least one ideal rhizobia strain is selected from rhizobia strains characterized by at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance in the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

[0228] Thus, In some embodiments, the at least one ideal rhizobia strain is selected from rhizobia strains having by at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance in the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

[0229] In some embodiments, the at least one ideal rhizobia strain is selected from rhizobia strains characterized by at least and occurs in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules of the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

[0230] In some embodiments, the at least one ideal rhizobia strain is selected from rhizobia strains having by at least and occurs in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules of the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

[0231] Thus, in some embodiments, the at least one ideal rhizobia strain is selected from rhizobia strains characterized by at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance in the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil, and / or wherein the at least one ideal rhizobia strain is selected from rhizobia strains characterized by at least and occurs in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules of the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

[0232] Thus, in some embodiments, the at least one ideal rhizobia strain is selected from rhizobia strains having at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance in the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil, and / or wherein the at least one ideal rhizobia strain is selected from rhizobia strains having at least and occurs in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules of the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

[0233] Formulation of the customized inoculant

[0234] The present disclosure concerns a method that allows to provide a customized inoculant formulation for a specific soil and plant comprising at least one preferred rhizobia strain identified as described herein.

[0235] It will be evident to the skilled person that step g. of the method described herein, i.e.: g. formulating the customized inoculant comprising the at least one preferred rhizobia strain;

[0236] Is meant to encompass, in some embodiments, defining or specifying the composition of the inoculant, including the at least one preferred rhizobia strain. As such, in some embodiments said step g. is a computer implemented method.

[0237] It will further be evident to the skilled person that, in some embodiments, step g. encompasses the actual preparation or production of the inoculant, wherein the specified components are combined to produce the inoculant, such as in a form ready for application. Thus, in some embodiments, step g. comprises preparing the inoculant.

[0238] In some embodiments, step g. comprises defining or specifying the composition of the inoculant, including the at least one preferred rhizobia strain; and preparing the inoculant.

[0239] The present disclosure concerns a method that allows to provide a customized inoculant formulation for a specific soil and plant comprising at least one preferred rhizobia strain identified as described herein and formulating the customized inoculant formulation with the at least one preferred rhizobia strain as described herein, such as the rhizobia selected in step f of the method described herein.

[0240] In an aspect, the present disclosure is directed to a customized inoculant formulation comprising a preferred rhizobia strain identified by the method described herein, as well as methods for providing or manufacturing said customized inoculant formulation.

[0241] In some embodiments, the customized inoculant formulation comprises the preferred rhizobia strain at a concentration of at least 1 x 10A5 cfu / ml , such at least 1 x 10A6 cfu / ml, such at least 1 x 10A7 cfu / ml, such at least 1 x 10A8 cfu / ml, such at least 1 x 10A9 cfu / ml, such at least 1 x 10A10 cfu / ml.

[0242] In some embodiments, the customized inoculant formulation is a soil inoculant.

[0243] In some embodiments, the customized inoculant formulation is a seed inoculant. Thus, in some embodiments, the customized inoculant formulation coats the seeds of the plant to be grown in a specific soil with the preferred rhizobia strain, as described herein, before sowing.

[0244] In some embodiments, customized inoculant formulation further comprises one or more components selected from: i. a carrier material, such as the material that that helps protect the preferred rhizobia strain and facilitates its application; ii. a nutrient source, such as nutrients that support the growth and survival of the preferred rhizobia strain; iii. a stabilizer, such as a substance that protects the preferred rhizobia strain from adverse environmental conditions, such as temperature fluctuations and UV radiation; iv. an adhesive and / or sticking agent, such as an adhesives to improve adherence to seeds or other planting materials; v. a surfactant, such as an agent that improves the spreading and coverage of the inoculant on seeds or in the soil; vi. a pH adjuster, such as an agent which ensures that the formulation provides the appropriate pH level for the preferred rhizobia strain; vii. an anti-desiccant, such as an agent which prevents the seeds from drying out during storage or planting.

[0245] In some embodiments, customized inoculant formulation further comprises a carrier material.

[0246] In some embodiments, customized inoculant formulation further comprises a nutrient source. In some embodiments, customized inoculant formulation further comprises a stabilizer. In some embodiments, customized inoculant formulation further comprises an adhesive and / or sticking agent. In some embodiments, customized inoculant formulation further comprises a surfactant. In some embodiments, customized inoculant formulation further comprises a pH adjuster. In some embodiments, customized inoculant formulation further comprises an anti-desiccant. In some embodiments, the customized inoculant formulation further comprises a polymer for seed coating. In some embodiments, the polymer is a biodegradable polymer. In some embodiments, the biodegradable polymer is a cellulose derivative.

[0247] Examples

[0248] Example 1: Microbiological soil analysis

[0249] Determining the Most Probable Number (MPN) of indigenous rhizobia strains capable of forming nodules on the specific host plant under investigation

[0250] Aim:

[0251] This is to determine functional rhizobia concentration in a given soil, meaning concentration of rhizobia that are capable to form nodule in a given plant.

[0252] Method:

[0253] Four sets of biological replicates were created for the inoculum, where each set involved mixing 1 gram of the target soil with 10 milliliters of sterile water. For each of the replicates, a series of dilutions ranging from 10'1to 10'9were prepared. The pots were filled with a mixture of vermiculite and leca, and then sterilized. Following this, seeds of the desired plant host or cultivar were placed in the pots and subseguently inoculated with 1 milliliter of the prepared dilutions.

[0254] To accelerate the nodulation process and achieve a time reduction of approximately 50% compared to plants grown under normal conditions, the plants were cultivated in a Speed breeding' (SB) growth chamber type. The SB growth chamber maintained a temperature of 28°C and provided 20 hours of light followed by 4 hours of darkness to support plant growth.

[0255] Once the plants displayed visible signs of nodulation, the roots were evaluated to determine if nodules were present (nodulation-positive) or absent (nodulation- negative). The Most Probable Number (MPN) was determined using the tables provided by Woomer, 1994.

[0256] Result:

[0257] Concentration of rhizobia capable of forming nodules in expressed by the highest dilution factor of soil resulting in nodulation of a given plant Conclusion:

[0258] The setup allows for quantification of functional rhizobia in rent plant-soil combinations.

[0259] Strain library (protocol)

[0260] Aim:

[0261] Creating strain library allows for collecting genetic diversity of natural rhizobia strains. Increasing the library size increases the potential for optimizing the efficiency of nitrogen fixation in a given environment in a given plant.

[0262] Method:

[0263] Nodules obtained from the plants and soil of interest were surface-sterilized and then allocated to designated positions in a 96-well plate by them. This was done to achieve two objectives: a) to evaluate the size of the nodules, and b) to isolate the corresponding rhizobia strains from each nodule by them.

[0264] To analyse size and exact position of each nodule, nodules are placed into 96-well plates by them, then scanned using a high-resolution scanner. Resulting images of nodules are processed using Imaged (Schneider, et.al. , 2012) to extract precise nodules sizes and retain nodule position in the 96-well plates. To account for the batch effect and biological variabilities in their setup, they calculate the contribution of each nodule as a fraction of total size. This enables comparison of nodule sizes between experiments. They developed an automated system using llastik (Berg, et.al., 2019) to decrease analysis time. In llastik, they use a set of images from different experiments to train pixel classifier to distinguish nodules from the background (Figure 1). In turn, they generate a list of objects characterized by size and position in the image. The latter allows for assigning nodules to specific positions in the 96-well plates. For batch processing, they apply the trained machine learning model (ML) to new images for rapid nodule size quantification.

[0265] To minimize processing time and prevent cross-contamination during nodule analysis, a 12-pestle device was custom-designed and 3D-printed by them to fit precisely into each well of the plate (Figure 2). This device enabled the simultaneous crushing of multiple nodules.

[0266] Using a 12-channel multichannel pipette, 100 microliters of sterile water were dispensed into each well by them and mixed. Then, 40 microliters of the resulting mixture were transferred to Rhizobium Defined Media (RDM) (Media described by Ronson and Primrose in 1979) to minimize the growth of non-rhizobia isolates.

[0267] Following the purification of the cultures, a high-throughput DNA extraction was carried out on 96-well plates by them.

[0268] To confirm the species of the isolates, PCR targeting the NodD gene of the given species was conducted using the extracted DNA by them. Subsequently, enterobacterial repetitive intergenic consensus PCR (ERIC-PCR) (as described by de Bruijn FJ. , 1992) were performed to generate a DNA fingerprint profile for each isolate. A ChemiDoc System was used to visualize and analyse DNA fragments in agarose gels. DNA bands profiles indicated the presence of distinct Rlv strains.

[0269] They determined the bacterial cell count using optical density (OD) in the strain library. This involves culturing serial dilutions of each rhizobia strain on plates and then placing them on a plate, which allows for the conversion of spectrophotometer readings of culture samples from each rhizobia strain to cell density.

[0270] The final strain library is stored at -80C to be used in future experiments by them.

[0271] Result: rhizobia strains are added into the library that allows for regrowing them as needed. The following data is stored: i) soil origin, ii) plant origin, iii) nodule size, iv) PCR amplification polymorphism.

[0272] Conclusion:

[0273] The method allows for isolating non-redundant rhizobia and preservation of information important for rhizobia selection as biofertilizers.

[0274] Characterisation of the background rhizobium population

[0275] Aim:

[0276] To acquire information about the background rhizobium population using sequencing. This parameter is used in the ML platform for prediction of elite rhizobium for biofertilizers.

[0277] Method:

[0278] Microbial DNA is isolated from a given soil or microbial DNA is isolated from plant root nodules. The resulting DNA is subjected to a targeted sequencing method, e.g. amplicon sequencing. The targeted genes include taxonomic marker genes and genes related to plant-rhizobium symbiosis. Alternatively, the microbial DNA is subjected to full metagenomic sequencing.

[0279] Result:

[0280] The sequencing data contains information on the background rhizobia population in a given soil.

[0281] Conclusion:

[0282] Characterization of the background rhizobia population using targeted sequencing provides input for the ML model.

[0283] Example 2: Physical and chemical soil analysis

[0284] Aim:

[0285] The soil analysis is conducted to obtain information for ML platform, subsequently used for prediction of optimal biofertilizer

[0286] Method:

[0287] A profile of each soil sample is created containing the following information: a. Place of sampling (including GPS coordinates) b. pH c. Total Nitrogen (N) content d. Phosphorus (P) content e. Potassium (K) content f. Magnesium (Mg) content g. Organic material % h. Soil class (JB) i. Soil texture %

[0288] Result:

[0289] A given soil is parameterized by factors listed in the method. This information is used in ML platform. Conclusion:

[0290] The approach allows for comprehensive physical and chemical soil characterization.

[0291] Example 3: Analysis of rhizobium competition and impact on plant performance

[0292] Aim:

[0293] This is to find a rhizobia strain or rhizobia strains that are the most competitive and most efficient in a given plant and in a given soil.

[0294] Methods: this section contains multiple methods contributing to the overarching aim of Example 3

[0295] Rhizobia labelling

[0296] Reporter plasmids, which include unique 12-nucleotide, error-correcting barcodes (IDs), referred to as Plasmid IDs, are used to label, track, and identify strains by sequencing to determine nodule occupancy. The Golden Gate cloning is used as a strategy for reporter plasmid construction to allow modular construction of the plasmids. The transformation of the Golden Gate cloning products into E. coli cells is done with ninety-six products at the same time. Subsequently, reporter plasmids are conjugated into rhizobia strains (Mendoza-Suarez et al., 2020).

[0297] Competition assays with multi strain inoculum

[0298] A pre-labelled multi-strain inoculum was prepared using 452 strains from the Rlv library. The bacterial cell count of the strain library was verified using optical density (OD). This involved culturing serial dilutions of each rhizobia strain on plates and converting spectrophotometer readings of culture samples from each rhizobia strain to cell density. The OD measurements were used to program a pipetting robot, ensuring the transfer of the necessary volume from each rhizobia strain culture.

[0299] The competitiveness and N2-effectiveness of each rhizobia strain in specific plant genotypes and / or soils was evaluated. The competition assays have reached a combination of multi-strain inoculum with up to 452 different strains and 213 different plant cultivars. Plant assays

[0300] The method of plant growth depends on the environment being tested. If required, the plants are cultivated under controlled laboratory conditions, or they are grown in a greenhouse facility with individualized irrigation to minimize cross-contamination. Every plant is logged into a database and labelled with a barcode to monitor its development and track the results.

[0301] When plants are harvested, the following data is collected and introduced to our database: a. Nodule number per plant b. Nodule size c. Total plant biomass

[0302] Identification of competitive strains from nodules

[0303] Root nodules were harvested, surface-sterilized, and pooled together. The DNA from the pooled nodules was extracted and processed for multiplex Illumina sequencing of the Plasmid-ID region. This allowed to determine the relative occupancy of each rhizobia strain.

[0304] Colonisation profiles of the isolates through the Faba bean panel

[0305] 608 individual plants, belonging to 212 Faba genotypes with 2-4 biological replicates, were analysed with the information of isolate composition. The library size of the samples ranged from 25 to 333,566, with a median of 113,866. There were 5 samples with very low sequencing depth (lower than 2000), which they did not include in the analyses (603 samples were analyzed). The number of isolates detected was 415, with several of those being very low in abundance. The data set was rather sparse, with 67% of the data points recording no presence of isolates (a 0 value). The sizes of the libraries in the samples showed significant variation, with a range from 25 to 333,566 and a median size of 113,866. Five samples with particularly low sequencing depth (below 2,000) were excluded from their analyses for quality control, leaving 603 samples to be analyzed on the x-axis of the heatmap in, where the Iog2-transformed relative abundances (in 100,000) of the Rhizobium strains placed on the y-axis. (Figure 3).

[0306] Consistent colonization patterns across plant genotypes were found for the different Rlv strains. They classified the strains in groups according to their colonization profiles into: Dominants, which are present ubiquitously and in high abundance; Specialists, which occur at high abundances but only in a small portion of the plants; Generalists, which are found frequently but in low abundance; and lastly, Transients, which occur rarely and in low abundance. (Figure 4).

[0307] Colonization profile groups correlated well with the soils where the isolates originated (Figure 5). The soil sources, UG (stands for the group of isolates from University of Gottingen from Germany), SJ and ND (stands for SEJET and Nordic Seed soils from Denmark, respectively) had most of the good colonising isolates, whereas the remaining soils consisted of the isolates that were unsuccessful in competition, therefore they were labelled as “Others”. Out of 86 Dominant strains, 80 were traced back to UG soil. Similarly, a significant proportion of Specialist strains (59 out of 80) also originated from UG soil. In contrast, SJ and ND isolates were mostly placed to the Generalists group, with 22 and 28 out of a total of 57 Generalist strains originating from SJ and ND soils, respectively. Further, more than 60% of isolates from SJ and ND were not robust colonizers and showed the characteristics of the “Transients”. Isolates that were categorised as “Dominants” (the most successful plant colonisers) have the highest niche breadth (Figure 6).

[0308] Community dynamics

[0309] The colonization classes showed varying community dynamics. After leaving out the non-robust colonisers, i.e. the Transients, the interactions among the strains from the remaining three groups was investigated based on their co-occurrence (Figure 7), mutual exclusion (Figure 8), co-abundance (Figure 9), and host-dependency (or universality) (Figure 10). The Dominants generally occurred together and did not show any mutual exclusion (Figure 8), whereas their abundance profiles were dissimilar, even the isolate pairs occurred together in a large number of faba bean genotypes did not show significant correlations in terms of their abundance patterns (Figure 9). The Generalists have shown both co-occurrence and co-abundance (Figure 9) without any significant mutual exclusions (Figure 8). It was not possible to detect any cooccurrence and co-abundance for the Specialists (Figure 7), but a large number of mutual exclusions among the isolate pairs was observed (Figure 8). Separate dissimilarity-overlap curve (DOC) analyses were also applied to the groups (Bashan et al. 2016). In this method, a negative slope between the number overlapping taxa (in this case rhizobia strains) and the dissimilarity of the subject pairs (faba bean genotypes) indicates the strains interact similarly in distinct individuals when they occur together, thereby high number of overlapping strains make the communities more similar (Verbruggen et al. 2018). It was observed that Generalists had a colonisation pattern in a host-independent fashion (fns = 0.55; Figure 10), whereas the DOC curves from the Dominants and the Specialists were flat (Figure 10). For the Dominants, the curve was a straight line, indicating an overlap-independent dissimilarity trend, suggesting non-universal dynamics for this group. Further, the Specialists were mostly mutually exclusive, therefore there should be few pairs with high community overlap, resulting in the absence of a negative slope.

[0310] Effect of the strains on

[0311] To determine if there is a link between the community composition of the isolates and the plant biomass, a linear mixed-effect model (LMM) was applied to the abundance of each rhizobia strain separately, using the bio-replicates as a random intercept and the plant genotype as a random slope within the random intercept within the batches. The model used is as follows:

[0312] Biomass ~ Abundance + (1 | Plant genotype) + (1 | Batch) + MDS2 + MDS3 + MDS4

[0313] Where MDS2-4 are the second to fourth dimensions of the GRM-based MDS analysis to account for the effect of the plant genotype. Relative abundances of the strains were log transformed first (Iog2(x + 1) where x represents relative abundance of an isolate). The estimates of the coefficients and p-value for the fixed effect (i.e., abundance) were calculated using the Imer function from the R package ImerTest. The strains that were classified as transients were not included in the analysis. After fitting the model for the remaining 223 isolates, p-values were corrected for multiple testing using Benjamini- Hochberg method. The effects from the individual strains on the biomass were negligible; only two isolates were found to have a significant influence on the plant growth (Figure. 11. The colonisation groups had shown distinct characters in terms of their growth promotion (Figure 12). In particular, the Generalists had a consistent negative effect on plant growth. In contrast, the Specialists and Dominants had mixed influence on the plants, spanning from negative to positive effects. It was also noticeable that the Specialists had a positive effect on the plant on average and the coefficient estimates of all three groups were significantly different from each other (Tukey’s HSD test at P < 0.05).

[0314] A negative correlation between niche breadth and the effect on the plant growth for the Generalists was observed (Figure 13). The Generalist isolates with a more uniform distribution through the Faba panel tended to have a more negative influence on the plant growth (r = -0.39, P = 0.004) (Figure 13). The relationship between the niche breadth and the plant growth promotion was profile-dependent: there was no correlation for these in the Specialists (r = 0.14, P = 0.22) and the Dominants had a positive correlation, but it was insignificant (r =0.19, P = 0.07) (Figure 13). Therefore, the colonisation success of these two groups was not associated with higher growth promotion, suggesting a more plant-dependent influence, in contrast to the Generalists, which was in line with the high genomic prediction accuracy of the cumulative relative abundance of the Dominants and the Specialists (Figure 13).

[0315] Even though the effect from single strains were low (Figure 11), the different colonization profiles groups had distinct impacts on the plants. This led them to evaluate if community-related parameters, such as evenness within the distinct colonisation groups, could explain the variation in plant growth promotion. Here they chose evenness over other parameters, such as Shannon’s diversity, since the number of isolates was unbalanced through the colonisation groups, particularly the Specialists occurred in much fewer plants in comparison to Generalists and Dominants (Figure 14). Therefore, even if very few Specialist isolates occurred in a plant, if their abundances were close to each other, that plant could get a high evenness value for the Specialists group. To associate the evenness with the plant biomass, three different LMM analyses were implemented (Figure 14). It was found that evenness of the Generalists and the Dominants had a significant impact on plant growth promotion, suggesting that a more even distribution of isolates from these groups could be linked to a positive effect on the plant performance (Figure 14). To consolidate this linear model-based result, a predictive analysis was performed using random forests, to evaluate if evenness of these groups could add predictive value. One hundred 80%- 20% train-test splits were performed to compare the prediction scores between Plant genotype + Batch effect + Evenness and Plant genotype + Batch effect (null model) (Figure 15). Indeed, the addition of the evenness information increased the predictivity by -10% (P = 0.0005) and the predictivity of a model with random evenness distributions was comparable to that of the null model (P = 0.41).

[0316] Result:

[0317] The above methods result in: i) unique labelling of each individual rhizobia strain, ii) characterizing each rhizobia strain by their competitiveness in different faba bean genotypes, iii) classification of rhizobia strains according to their colonization patters in the faba bean panel, iv) quantification of rhizobia classes for their ability to promote plant growth.

[0318] Conclusion:

[0319] Elite strain identification can be achieved empirically using the described experimental procedure. The described procedure allows for identification of parameters important for predicting of the elite strains using ML, obtaining training data for ML.

[0320] Example 4: Databases

[0321] Strain database

[0322] This database contains the list of rhizobia isolates from the strain library (see Example 3) and retain metadata on strain characteristics. The database in a format that allows to be integrated with other databases.

[0323] Soil database

[0324] This database contains the list of soil analyzed, with metadata on physical and chemical soil analysis (see Example 2), microbiological characteristics, including MPN and rhizobia background population (see Example 7) and sample origin (date of sampling and GPS coordinates). The database is in a format that allows to be integrated with other databases. Plant database

[0325] This database contains the list of 255 faba bean cultivars included in the ProFaba project. The list contains metadata on genotype (SNPs), cultivar pedigree and cultivar origin. Information regarding the ProFaba project can be found at https: / / www.suscrop.eu / projects-first-call / profaba. The database is in a format that allows to be integrated with other databases.

[0326] Plant, soil, bacteria match database

[0327] This dataset contains the list of elite strains for plant grown in specific soil.

[0328] The data is obtained as described in Example 3.

[0329] Field trial database

[0330] This dataset contains the list of field trials carried out. The database contains metadata on trial date, trial location, field management, plants used, rhizobia strains tested, soil identifier and plant field performance (e.g. seed yield, biomass, vigour). The database is in a format that allows to be integrated with other databases.

[0331] Example 5: Training of the machine learning model and prediction of elite strains

[0332] Aim:

[0333] Two machine learning models (ML model) are used for predicting the two parameters describing rhizobium-plant interaction:

[0334] Rhizobium competitiveness expressed as the ability of a given strain to occupy a given plant genotype nodules (nodule occupancy, Yocc) in a given soil environment with multiple strains present,

[0335] Rhizobium effectiveness (Ye / r) expressed as the ability of a given strain to acquire nitrogen through the biological nitrogen fixation or their effect on plant biomass production in a given plant genotype in a given soil environment.

[0336] Yocc and Ye / f are used to predict elite strains, subsequently used as inoculants. The training data is described in the Method section.

[0337] Method:

[0338] The method consists of two ML models:

[0339] I) As input data for training to predict rhizobium competitiveness, it is provided: • Soil characteristics including chemical composition (cc), physical properties (pp), historical soil temperature data (ht), rhizobium concentration defined as the most probable number of compatible strains (mpri) determined by trapping using a soil concentration series , e) DNA fingerprints or DNA sequences of trapped rhizobia or rhizobia present in the soil (Grhiz-background)

[0340] • Plant genotype as characterised by DNA markers distributed across the genome for the target cultivar (G plant).

[0341] • Rhizobium genotype data as characterised by DNA sequencing data for the candidate strain (G rhiz -candidate) .

[0342] • The nodule occupancy of the candidate strain (Y0Cc).

[0343] The following model is used:

[0344] Yocc = cc + pp + ht + mpn + Grhiz -background + Gplant + Grhiz-candidate

[0345] The data is split into training and test datasets and train a machine learning model, e.g. random forest, extreme gradient boosting or similar, to predict the nodule occupancy of a candidate rhizobium strain in a given plant genotype in a given soil with a given background rhizobium population. Cross-validation is used with the testing set to determine the prediction accuracy.

[0346] II) As input data for training to predict rhizobium effectiveness, it is provided

[0347] • Soil characteristics including chemical composition (cc), physical properties (pp), historical soil temperature data (ht), rhizobium concentration defined as the most probable number of compatible strains (mpn determined by trapping using a concentration series, e) DNA fingerprints or DNA sequences of trapped rhizobia or rhizobia present in the soil (Grhiz-background)

[0348] • Plant genotype as characterised by DNA markers distributed across the genome for the target cultivar (G plant).

[0349] • Rhizobium genotype data as characterised by DNA sequencing data for the Candidate Strain (Grhiz-candidate) .

[0350] The nitrogen fixation efficiency of the candidate strain (Yetf) as determined by acetylene reduction assays, or using other methods, on individual root nodules. The following model is used:

[0351] Yeff = cc + pp + ht + mpn + Grhiz -background + Gplant + Grhiz-candidate

[0352] The data is split into training and test datasets and train a machine learning model, e.g. random forest, extreme gradient boosting or similar, to predict the nitrogen fixation efficiency of a candidate rhizobium strain in a given plant genotype in a given soil with a given background rhizobium population. Cross-validation is used with the testing set to determine the prediction accuracy.

[0353] Result:

[0354] As the result of the ML model, rhizobia strains are ranked according to their Y0Cc and Yeff in a given plant genotype in a given soil.

[0355] Conclusion:

[0356] Among the rhizobia strains predicted to be highly competitive (high Y0Cc), one or more strains that are also predicted to be highly efficient (high Yetf) are selected as inoculants.

[0357] Example 6: Validation of elite strains

[0358] Aim:

[0359] Experiments under controlled conditions or field trials are used to validate the prediction of inoculant strain properties Y0Cc and Yeff and to produce additional training data to improve machine learning model performance.

[0360] Method:

[0361] When rhizobia strains have been selected as inoculants these are applied to the target plant genotype in the target soil under controlled conditions in growth chambers or in the greenhouse. Plant biomass, Yoccand Yeff are recorded. The same inocluatns are applied in the field where plant yield and Yoccare recorded.

[0362] Result:

[0363] The Yocc data from controlled conditions and field trials and the Yeff data from controlled conditions is added to the database and included as training data for the machine learning model. The plant biomass and yield data are used to validate the plant performance improvement resulting from rhizobia inoculation. Conclusion:

[0364] The validation procedure serves both to confirm predictions and to continuously generate additional data for improving machine learning prediction accuracy.

[0365] Example 7. Selection and field trial validation of elite strains

[0366] Aim:

[0367] The aim was to i) carry out identification of the best rhizobia strains for given plant genotypes and a given soil, ii) prepare customized inoculants and iii) assess their performance in the field.

[0368] Method:

[0369] Soil from an agricultural field in Northern Germany was sampled. The soil was subjected to i) physical and chemical analysis and ii) microbiological soil analysis with MPN. The field trail was planned with 3 specific genotypes of faba bean. These faba bean genotypes were used in a rhizobia competition assay. In this assay, the given plants were grown in presence of the field trial soil with addition of the library of Plasmid-ID tagged rhizobia strains. The plants were harvested after nodule emergence and inspected for GFP signal (strains from the library produce GFP while in nodules, native strains from the trial soil do not). The nodules were sampled into 96-well plates, surface sterilized and crushed to retrieve rhizobia from the nodules. Then the strains were cultured, DNA from the strains was isolated and DNA fingerprinting PCR was performed. High frequency fingerprints provided information on rhizobia competitiveness and those with high competitiveness were selected for a single inoculation plant assay. Single strains were added to the given 3 plant genotypes. After nodules were formed, the nodules were subjected to acetylene reduction assays and assessed for nitrogen fixation efficiency. Rhizobia strains that were highly competitive and effective were selected as elite strains. The selected strains were multiplied in amounts sufficient for in-field inoculation using a bioreactor.

[0370] For the field trial, seeds of the 3 given faba bean genotypes were mixed with the selected inoculants. Formulation 1 consisted of a single rhizobia strain. Formulation 2 consisted of 4 strains. Formulation 1 and 2 treated seeds constituted the experimental treatment. The experimental treatment was tested against controls consisting of the same 3 given faba bean genotype, but with no additional rhizobia added (standard practice). 5 replicates per treatment were used in a fully randomized field trial design. The field trial was managed by external partners.

[0371] Result:

[0372] During the trial the following data was collected: i) seed yield, ii) seed number, iii) number of plants, iv) seed protein content, v) plant above-ground biomass, vi) flowering time. Rhizobia strain inoculation treatment was then compared to untreated controls. In this case, they found a significant positive effect of the rhizobium inoculation with a selected elite strain (Figure 17).

[0373] Conclusion:

[0374] Elite strains were selected and an unbiased assessment of inoculant performance was carried out. The customized inoculation improved plant performance in the field.

[0375] Example 8. Benchmarking of a customized rhizobia biofertilizer againts a generic rhizobia biofertilizer

[0376] Aim:

[0377] Assessment of the performance of a customized rhizobia biofertilizer formulation for faba beans, tailored to given specific soil, compared to a generic biofertilizer.

[0378] Customized biofertilizer treatment

[0379] A Spanish soil sample was received and subjected to physico-chemical analysis, including tests for phytonutrients (4505 and 4522), pH (Rt), mineral nitrogen (N min), phosphorus, potassium, magnesium, organic matter, and soil class (JB). The soil was also assessed to determine the abundance of rhizobia compatible with faba beans.

[0380] Soils that shared similar pH and phytonutrient profiles were identified using a database of soils comprising the same parameters. Rhizobia strains that had been previously isolated from these soils were grown and used in equal quantity to inoculate the target faba bean genotype in the target soil. The candidate strains were fluorescently labeled, and whether the candidate or the native strains occupied the nodules of faba beans was assessed. In this particular experiment, the native strains of the soil outcompeted the candidate strains from our library.

[0381] Based on the genomic fingerprint of the native strains, the inventors selected the most competitive rhizobia occurring in the largest number of nodules. This analysis was followed by acetylene reduction assays of the top competitive strains to determine their nitrogen fixation efficiency. Based on the combined evaluation of competitiveness and efficiency, the native rhizobia strain ALA-A1 was selected as our inoculant. ALA-A1 was fermented to achieve at least the same abundance as found in the initial soil analysis, and the final formulation was coated on the faba bean seeds. From now on, we will refer to this treatment as the Customized biofertilizer.

[0382] Generic biofertilizer treatment

[0383] The commercial Rhizobium bacteria biofertilizer LegumeFix peat for Faba bean / Vetch / Lentil from the company Legume Technology was acquired through the distributor LegumiN. Faba bean seeds were coated following the company's recommendations. From now on, we will refer to this treatment as the Generic biofertilizer.

[0384] Non-treated seeds treatment

[0385] Untreated faba bean seeds were used as a negative control. From now on, we will refer to this treatment as the Non-treated seeds.

[0386] Experiment

[0387] Faba bean plants were grown in 14 cm3square pots filled with a mix of Spanish soil, leca, and vermiculite. Each pot had individualized irrigation systems to minimize crosscontamination. All plants received full standard nutrient fertilization except nitrogen. The Customized biofertilizer and Generic biofertilizer treatments had ten biological replicates each, and the Non-treated seeds had five biological replicates.

[0388] After 60 days, when the plants reached maturity, all treatments were harvested. Shoot length (Figure 18) and shoot dry biomass (Figure 19) were measured to assess the performance of each treatment. Conclusion:

[0389] Faba beans grown with the Customized biofertilizer showed increased shoot length and greater shoot dry biomass compared to faba beans grown with the Generic biofertilizer. This example demonstrates that it is possible to select a reduced set of rhizobia for rapid testing, and that a Customized biofertilizer can outperform standard, commercially available biofertilizers.

[0390] References

[0391] Bashan, Amir, Travis E. Gibson, Jonathan Friedman, Vincent J. Carey, Scott T. Weiss, Elizabeth L. Hohmann, and Yang-Yu Liu. 2016. “Universality of Human Microbial

[0392] Dynamics.” Nature 534 (7606): 259-62. doi:10.1038 / nature18301.

[0393] Verbruggen, Erik, Merlin Sheldrake, Luke D Bainard, Baodong Chen, Tobias Ceulemans, Johan De Gruyter, and Maarten Van Geel. 2018. “Mycorrhizal Fungi Show Regular Community Compositions in Natural Ecosystems.” The ISME Journal 12 (2): 380-85. doi: 10.1038 / ismej.2017.169.

[0394] Items

[0395] 1. A method for providing a customized inoculant formulation for a plant grown in a soil comprising at least one preferred rhizobia strain, said method comprising the steps of: a. providing information / data of the plant, comprising at least said plant species; b. providing information / data of the soil, said information / data comprising at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and iv. information / data one or more naturally occurring rhizobia present in the soil; c. predicting, by a machine learning model, the competitiveness of each rhizobia strain selected from a plurality of known rhizobia strains, wherein the competitiveness of each rhizobia strain is the ability to occupy the nodules of a plant in a given soil with multiple rhizobia strains present, and wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain; ii. a plurality of plants, and associated plant data for each plant comprising at least said plant species; and iii. a plurality of soils, and associated soil data for each soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; and information / data of one or more rhizobia present in said soil; iv. a plurality of combinations of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the competitiveness of each rhizobia strain selected from the known rhizobia strains in i.; d. predicting, by a machine learning model, the effectiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or the one or more rhizobia strains predicted in step c. to have high competitiveness, wherein the effectiveness of a rhizobia strain is the ability to acquire nitrogen through the biological nitrogen fixation and / or its effect on plant biomass production in the plant grown in the soil, wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain; ii. a plurality of plants, and associated plant data for each plant comprising at least said plant species; and iii. a plurality of soils, and associated soil data for each soil comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia present in said soil; and iv. a plurality of combinations of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the effectiveness of each rhizobia strain selected from the known rhizobia strains in i.; f. selecting at least one preferred rhizobia strain; g. formulating the customized inoculant comprising the at least one preferred rhizobia strain; thereby providing the customized inoculant formulation for said plant and soil. The method according to item 1, wherein the method further comprises performing a validation step e. comprising the steps of: i. inoculating in the soil at least one of the one or more known rhizobia strains predicted to have high competitiveness in step c. and high effectiveness in step d.; ii. growing the plant in the soil; and iii. identifying rhizobia strains in the nodules of the plant. The method according to item 2, wherein identifying rhizobia strains in the nodules of the plant comprises sequencing the rhizobia strains. The method according to any one of items 2 to 3, wherein identifying rhizobia strains in the nodules of the plant comprises sequencing an identifier unique for each rhizobia strain. The method according to item 4, wherein the identifier unique for each rhizobia strain is selected from: a natural nucleotide sequence in the genome of the rhizobia strains; a synthetic nucleotide sequence in the genome of the rhizobia strains; a nucleotide sequence in the rRNA of the rhizobia strains; or a nucleotide sequence in a vector, such as a plasmid (extrachromosomal genome).

[0396] 6. The method according to any one of items 2 to 5, wherein the at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e.

[0397] 7. The method according to any one of items 2 to 6 wherein the at least one preferred rhizobia strain selected in step f. is selected from rhizobia strains identified in the nodules of the plant in step e. with at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance.

[0398] 8. The method according to any one of items 2 to 7, wherein the at least one preferred rhizobia strain selected in step f. is selected from rhizobia strains identified in the nodules of the plant in step e. which occur in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules analyzed.

[0399] 9. The method according to any one of items 2 to 8, wherein the soil comprises at least one naturally occurring rhizobia strain.

[0400] 10. The method according to any one of items 2 to 9, wherein the plant is grown in conditions that accelerate nodulation.

[0401] 11. The method according to items 2 to 10, wherein the validation step e. further comprises a step: iv. quantifying one or more of the parameters selected from: a. the number of nodules per plant; b. the size of each nodule; c. proxy of nitrogen fixation in individual nodules d. the number of pods per plant; e. the pod weight; f. the total biomass of the plant. The method according to any one of items 2 to 11, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying number of nodules per plant. The method according to any one of items 2 to 12, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the size of each nodule. The method according to any one of items 2 to 13, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying proxy of nitrogen fixation in individual nodules. The method according to any one of items 11 to 14, wherein the method to quantify a proxy of nitrogen fixation in individual nodules is selected from: acetylene reduction assay (ARA), isotope 15N2 assimilation, isotopic acetylene reduction assay (ISARA), total nitrogen difference (TND), nodule size and / or biomass, or expression of a reporter, such as GFP, under the control of nif genes promoters. The method according to any one of items 2 to 15, wherein the validation step d. further comprises a step iv. comprising or consists of a step of quantifying the number of pods per plant. The method according to any one of items 2 to 16, wherein the validation step e. further comprises or consists of a step iv. comprising a step of quantifying the pod weight. The method according to any one of items 2 to 17, wherein the validation step e. further comprises or consists of a step iv. comprising a step of quantifying the total biomass. The method according to any one of items 2 to 18, wherein the plant is grown in a speed breeding growth chamber. The method according to any one of items 2 to 19, wherein the plant is grown in a photoperiod between 18.5 and 22.5 hours, such as between 18.5 and 22 hours, such as between 18.5 and 21.5 hours, such as between 18.5 and 21 hours, such as between 18.5 and 20.5 hours, such as between 18.5 and 20 hours, such as between 18.5 and 19.5 hours, such as between 18.5 and 19 hours, such as between 19 and 22.5 hours, such as between 19 and 22 hours, such as between 19 and 21.5 hours, such as between 19 and 21 hours, such as between 19 and 20.5 hours, such as between 19 and 20 hours, such as between 19 and 19.5 hours, such as between 19.5 and 22.5 hours, such as between 19.5 and 22 hours, such as between 19.5 and 21.5 hours, such as between 19.5 and 21 hours, such as between 19.5 and 20.5 hours, such as between 19.5 and 20 hours, such as between 20 and 22.5 hours, such as between 20 and 22 hours, such as between 20 and 21.5 hours, such as between 20 and 21 hours, such as between 20 and 20.5 hours, such as between 20.5 and 22.5 hours, such as between 20.5 and 22 hours, such as between 20.5 and 21.5 hours, such as between 20.5 and 21 hours, such as between 21 and 22.5 hours, such as between 21 and 22 hours, such as between 21 and 21.5 hours, such as between 21.5 and 22.5 hours, such as between 21.5 and 22 hours, such as between 22 and 22.5 hours. The method according to any one of items 2 to 20, wherein the plant is grown at a temperature between 16 and 24°C, such as between 16 and 23°C, such as between 16 and 22°C, such as between 16 and 21 °C, such as between 16 and 20°C, such as between 16 and 19°C, such as between 16 and 18°C, such as between 16 and 17°C, such as between 17 and 24°C, such as between 17 and 23°C, such as between 17 and 22°C, such as between 17 and 21 °C, such as between 17 and 20°C, such as between 17 and 19°C, such as between 17 and 18°C, such as between 18 and 24°C, such as between 18 and 23°C, such as between 18 and 22°C, such as between 18 and 21 °C, such as between 18 and 20°C, such as between 18 and 19°C, such as between 19 and 24°C, such as between 19 and 23°C, such as between 19 and 22°C, such as between 19 and 21°C, such as between 19 and 20°C, such as between 20 and 24°C, such as between 20 and 23°C, such as between 20 and 22°C, such as between 20 and 21°C, such as between 21 and 24°C, such as between 21 and 23°C, such as between 21 and 22°C, such as between 22 and 24°C, such as between 22 and 23°C, such as between 23 and 24°C. The method according to anyone of items 2 to 21, wherein the plant is grown in a photoperiod between 50 and 80% humidity, such as between 50 and 75% humidity, such as between 50 and 70% humidity, such as between 50 and 65% humidity, such as between 50 and 60% humidity, such as between 50 and 55% humidity, such as between 55 and 80% humidity, such as between 55 and 75% humidity, such as between 55 and 70% humidity, such as between 55 and 65% humidity, such as between 55 and 60% humidity, such as between 60 and 80% humidity, such as between 60 and 75% humidity, such as between 60 and 70% humidity, such as between 60 and 65% humidity, such as between 65 and 80% humidity, such as between 65 and 75% humidity, such as between 65 and 70% humidity, such as between 70 and 80% humidity, such as between 70 and 75% humidity, such as between 75 and 80% humidity. The method according to any one of items 2 to 22, wherein the validation step e. is performed in less than 5 months, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, preferably less than 3 months. The method according to any one of items 2 to 23, wherein the result of the validation step e. is added to the information provided to train a machine learning models to be used in a method as described in any one of the preceding items. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing one of more of the following soil parameters selected from: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil; f. the soil class; or g. the place of sampling. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the concentration of one or more naturally occurring rhizobia strain in the soil. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the concentration of phosphorus in the soil. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the concentration of potassium in the soil. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the concentration of magnesium in the soil. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the percentage of organic material in the soil. The method according to any one of items 25 to 30, wherein the percentage of organic material in the soil is determined by dry combustion. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the soil class. The method according to any one of the previous items, wherein providing information / data of the soil further comprises providing the place of sampling, preferably wherein providing the information of the place of sampling comprises providing the GPS coordinates of the place of sampling. 34. The method according to any one of the previous items, wherein providing information / data of the plant further comprises providing one or more of the information / data selected from: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; d. the phenotype of the plant.

[0402] 35. The method according to any one of the previous items, wherein the associated rhizobia strain data comprises one or more of the data selected from: a. rhizobia strain genera; b. rhizobia strain species; c. rhizobia strain subspecies; d. rhizobia strain biovar; and e. rhizobia strain type or number

[0403] 36. The method according to any one of the previous items, wherein the associated plant data comprises one or more of the data selected from: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; and d. the phenotype of the plant

[0404] 37. The method according to any one of items 34 or 36, wherein providing the phenotype of the plant comprises providing one or more of the data selected from: a. plant height; b. yield; c. flowering time of the plant; d. seed size; e. shape of seeds and plant branching structure; f. color of flowers and seeds; g. insect tolerance; and h. pest tolerance. The method according to any one of the previous items, wherein the associated soil data further comprises one or more of the data selected from: a. the concentration of phosphorus in the soil; b. the concentration of potassium in the soil; c. the concentration of magnesium in the soil; d. the percentage of organic material in the soil; e. the soil class; f. the place of sampling. The method according to any one of the previous items, wherein the at least one ideal rhizobia strain is selected from rhizobia strains having at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance in the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil. The method according to any one of the previous items, wherein the at least one ideal rhizobia strain is selected from rhizobia strains having at least and occurs in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules of the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil. The method according to any one of the previous items, wherein the soil has been obtained from a field before plant has been sown in the soil. The method according to any one of the previous items, wherein the soil has been obtained no more than 5 months prior to the sowing of the plant, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, such as 3 days, such as 1 day prior to the sowing of the plant. 43. A customized inoculant formulation wherein the customized inoculant formulation comprises a preferred rhizobia strain identified by the method according to any one of the preceding claims.

[0405] 44. The customized inoculant formulation according to item 43, wherein the customized inoculant comprises the preferred rhizobia strain at a concentration of at least 1 x 10A5 cfu / ml , such at least 1 x 10A6 cfu / ml, such at least 1 x 10A7 cfu / ml, such at least 1 x 10A8 cfu / ml, such at least 1 x 10A9 cfu / ml, such at least 1 x 10A10 cfu / ml.

[0406] 45. The customized inoculant formulation according to any one of items 43 to 44, wherein the customized inoculant further comprises a polymer for seed coating.

[0407] 46. The customized inoculant formulation according to item 45, wherein the polymer is a biodegradable polymer.

[0408] 47. The customized inoculant formulation according to item 46, wherein the biodegradable polymer is a cellulose derivative.

[0409] 48. A method for providing / manufacturing a customized inoculant formulation as defined in any one of items 43 to 47 comprising the method according to any one of the items 1 to 40.

Claims

Claims1. A method for providing a customized inoculant formulation for a plant grown in a soil comprising at least one preferred rhizobia strain, said method comprising the steps of: a. providing information / data of the plant, comprising at least said plant species; b. providing information / data of the soil, said information / data comprising at least: i. the pH of the soil; ii. the concentration of Nitrogen in the soil; iii. the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains (MPN) in the soil; and iv. information / data one or more naturally occurring rhizobia present in the soil; c. predicting, by a machine learning model, the competitiveness of each rhizobia strain selected from a plurality of known rhizobia strains, wherein the competitiveness of each rhizobia strain is the ability to occupy the nodules of a plant in a given soil with multiple rhizobia strains present, and wherein the machine learning model has been trained by a method comprising the steps of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality of known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant of the plurality of plants comprising at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality comprising at least: the pH of said soil; the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil;information / data of one or more rhizobia strain present in said soil; and iv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the competitiveness of each rhizobia strain selected from the known rhizobia strains in i.; d. predicting, by a machine learning model, the effectiveness of each rhizobia strain selected from: a plurality of known rhizobia strains; or the one or more rhizobia strains predicted in step c. to have high competitiveness, wherein the effectiveness of a rhizobia strain is the ability to acquire nitrogen through the biological nitrogen fixation and / or its effect on plant biomass production in the plant grown in the soil, wherein the machine learning model has been trained by a method comprising the step of: providing information / data comprising: i. information / data of a plurality of known rhizobia strains, and associated rhizobia strain data for each known rhizobia strain of the plurality known rhizobia strains; ii. information / data of a plurality of plants, and associated plant data for each plant of the plurality of plants comprising at least said plant species; and iii. information / data of a plurality of soils, and associated soil data for each soil of the plurality comprising at least: the pH of said soil; and the concentration of Nitrogen in said soil; the abundance or density of microorganisms in the soil, such as the most probable number of compatible strains in said soil; information / data of at least one rhizobia present in said soil; andiv. information / data of a plurality of combinations of plant and soil, and associated data for each combination of plant and soil, wherein: the plant in the plurality of combinations of plant and soil is comprised in the plurality of plants in ii.; the soil in the plurality of combinations of plant and soil is comprised in the plurality of soils in iii.; and for each combination of plant and soil is provided the effectiveness of each rhizobia strain selected from the known rhizobia strains in i.; f. selecting at least one preferred rhizobia strain; and g. formulating the customized inoculant comprising the at least one preferred rhizobia strain; thereby providing the customized inoculant formulation for said plant and soil.

2. The method according to claim 1, wherein the method further comprises performing a validation step e. comprising the steps of: i. inoculating in the soil at least one of the one or more known rhizobia strains predicted to have high competitiveness in step c. and high effectiveness in step d.; ii. growing the plant in the soil; and iii. identifying rhizobia strains in the nodules of the plant.

3. The method according to claim 2, wherein identifying rhizobia strains in the nodules of the plant comprises sequencing the rhizobia strains.

4. The method according to any one of claims 2 to 3, wherein identifying rhizobia strains in the nodules of the plant comprises sequencing an identifier unique for each rhizobia strain.

5. The method according to claim 4, wherein the identifier unique for each rhizobia strain is selected from: a natural nucleotide sequence in the genome of the rhizobia strains; a synthetic nucleotide sequence in the genome of the rhizobia strains; a nucleotide sequence in the rRNA of the rhizobia strains; or a nucleotide sequence in a vector, such as a plasmid (extrachromosomal genome).

6. The method according to any one of claims 2 to 5, wherein the at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e.

7. The method according to any one of claims 2 to 6 wherein the at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e. with at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance.

8. The method according to any one of claims 2 to 7, wherein the at least one preferred rhizobia strain selected in step f. is selected from the rhizobia strains identified in the nodules of the plant in step e. which occur in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules analyzed.

9. The method according to any one of claims 2 to 8, wherein the soil comprises at least one naturally occurring rhizobia strain.

10. The method according to any one of claims 2 to 9, wherein the plant is grown in conditions that accelerate nodulation.

11. The method according to claims 2 to 10, wherein the validation step e. further comprises a step of: iv. quantifying one or more of the parameters selected from: a. the number of nodules per plant; b. the size of each nodule; c. proxy of nitrogen fixation in individual nodules; d. the number of pods per plant; e. the pod weight; f. the total biomass of the plant.

12. The method according to any one of claims 2 to 11, wherein the validation step e. further comprises a step iv. comprising or consisting of a step of quantifying the number of nodules per plant.

13. The method according to any one of claims 2 to 12, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the size of each nodule.

14. The method according to any one of claims 2 to 13, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying proxy of nitrogen fixation in individual nodules.

15. The method according to any one of claims 11 to 14, wherein the method to quantify a proxy of nitrogen fixation in individual nodules is selected from: acetylene reduction assay (ARA), isotope 15N2 assimilation, isotopic acetylene reduction assay (ISARA), total nitrogen difference (TND), nodule size and / or biomass, or expression of a reporter, such as GFP, under the control of nif genes promoters.

16. The method according to any one of claims 2 to 15, wherein the validation step d. further comprises a step iv. comprising or consists of a step of quantifying the number of pods per plant.

17. The method according to any one of claims 2 to 16, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the pod weight.

18. The method according to any one of claims 2 to 17, wherein the validation step e. further comprises a step iv. comprising or consists of a step of quantifying the total biomass.

19. The method according to any one of claims 2 to 18, wherein the plant is grown in a speed breeding growth chamber.

20. The method according to any one of claims 2 to 19, wherein the plant is grown in a photoperiod between 18.5 and 22.5 hours, such as between 18.5 and 22 hours, such as between 18.5 and 21.5 hours, such as between 18.5 and 21 hours, such as between 18.5 and 20.5 hours, such as between 18.5 and 20 hours, such as between 18.5 and 19.5 hours, such as between 18.5 and 19 hours, such as between 19 and 22.5 hours, such as between 19 and 22 hours, such as between 19 and 21.5 hours, such as between 19 and 21 hours, such as between 19 and 20.5 hours, such as between 19 and 20 hours, such as between 19 and 19.5 hours, such as between 19.5 and 22.5 hours, such as between 19.5 and 22 hours, such as between 19.5 and 21.5 hours, such as between 19.5 and 21 hours, such as between 19.5 and 20.5 hours, such as between 19.5 and 20 hours, such as between 20 and 22.5 hours, such as between 20 and 22 hours, such as between 20 and 21.5 hours, such as between 20 and 21 hours, such as between 20 and 20.5 hours, such as between 20.5 and 22.5 hours, such as between 20.5 and 22 hours, such as between 20.5 and 21.5 hours, such as between 20.5 and 21 hours, such as between 21 and 22.5 hours, such as between 21 and 22 hours, such as between 21 and 21.5 hours, such as between 21.5 and 22.5 hours, such as between 21.5 and 22 hours, such as between 22 and 22.5 hours.21 . The method according to any one of claims 2 to 20, wherein the plant is grown at a temperature between 16 and 24°C, such as between 16 and 23°C, such as between 16 and 22°C, such as between 16 and 21 °C, such as between 16 and 20°C, such as between 16 and 19°C, such as between 16 and 18°C, such as between 16 and 17°C, such as between 17 and 24°C, such as between 17 and 23°C, such as between 17 and 22°C, such as between 17 and 21°C, such as between 17 and 20°C, such as between 17 and 19°C, such as between 17 and 18°C, such as between 18 and 24°C, such as between 18 and 23°C, such as between 18 and 22°C, such as between 18 and 21°C, such as between 18 and 20°C, such as between 18 and 19°C, such as between 19 and 24°C, such as between 19 and 23°C, such as between 19 and 22°C, such as between 19 and 21°C, such as between 19 and 20°C, such as between 20 and 24°C, such as between 20 and 23°C, such as between 20 and 22°C, such as between 20 and 21°C, such as between 21 and 24°C, such as between 21 and 23°C, such asbetween 21 and 22°C, such as between 22 and 24°C, such as between 22 and 23°C, such as between 23 and 24°C.

22. The method according to anyone of claims 2 to 21 , wherein the plant is grown in a photoperiod between 50 and 80% humidity, such as between 50 and 75% humidity, such as between 50 and 70% humidity, such as between 50 and 65% humidity, such as between 50 and 60% humidity, such as between 50 and 55% humidity, such as between 55 and 80% humidity, such as between 55 and 75% humidity, such as between 55 and 70% humidity, such as between 55 and 65% humidity, such as between 55 and 60% humidity, such as between 60 and 80% humidity, such as between 60 and 75% humidity, such as between 60 and 70% humidity, such as between 60 and 65% humidity, such as between 65 and 80% humidity, such as between 65 and 75% humidity, such as between 65 and 70% humidity, such as between 70 and 80% humidity, such as between 70 and 75% humidity, such as between 75 and 80% humidity.

23. The method according to any one of claims 2 to 22, wherein the validation step e. is performed in less than 5 months, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, preferably less than 3 months.

24. The method according to any one of claims 2 to 23, wherein the result of the validation step e. is added to the information provided to train a machine learning models to be used in a method as described in any one of the preceding claims.

25. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing one of more of the following soil parameters: a. the concentration of one or more naturally occurring rhizobia strain in the soil; b. the concentration of phosphorus in the soil; c. the concentration of potassium in the soil; d. the concentration of magnesium in the soil; e. the percentage of organic material in the soil;f. the soil class; g. the place of sampling.

26. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the concentration of one or more naturally occurring rhizobia strain in the soil.

27. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the concentration of phosphorus in the soil.

28. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the concentration of potassium in the soil.

29. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the concentration of magnesium in the soil.

30. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the percentage of organic material in the soil.

31. The method according to any one of claims 25 to 30, wherein the percentage of organic material in the soil is determined by dry combustion.

32. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the soil class.

33. The method according to any one of the previous claims, wherein providing information / data of the soil further comprises providing the place of sampling, preferably wherein providing the information of the place of sampling comprises providing the GPS coordinates of the place of sampling.

34. The method according to any one of the previous claims, wherein providing information / data of the plant further comprises providing one or more of the information / data selected from: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; or d. the phenotype of the plant.

35. The method according to any one of the previous claims, wherein the associated rhizobia strain data comprises one or more of the data selected from: a. rhizobia strain genera; b. rhizobia strain species; c. rhizobia strain subspecies; d. rhizobia strain biovar; or e. rhizobia strain type or number36. The method according to any one of the previous claims, wherein the associated plant data comprises one or more of the data selected from: a. the cultivar of the plant; b. the genus of the plant; c. the genotype of the plant; or d. the phenotype of the plant37. The method according to any one of claims 34 or 36, wherein providing the phenotype of the plant comprises providing one or more of the data selected from: a. plant height; b. yield; c. flowering time of the plant; d. seed size; e. shape of seeds and plant branching structure; f. color of flowers and seeds; g. insect tolerance; or h. pest tolerance.

38. The method according to any one of the previous claims, wherein the associated soil data further comprises one or more of the data selected from: a. the concentration of phosphorus in the soil; b. the concentration of potassium in the soil; c. the concentration of magnesium in the soil; d. the percentage of organic material in the soil; e. the soil class; or f. the place of sampling.

39. The method according to any one of the previous claims, wherein the at least one ideal rhizobia strain is selected from the rhizobia strains having at least 0.01 relative abundance, such as at least 0.02, such as at least 0.03, such as at least 0.04, such as at least 0.05, such as at least 0.06, such as at least 0.07, such as at least 0.09, such as at least 0.10, such as at least 0.15, such as at least 0.20, such as at least 0.3 relative abundance in the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

40. The method according to any one of the previous claims, wherein the at least one ideal rhizobia strain is selected from the rhizobia strains having at least and occurs in at least 50% of the nodules analyzed, such as in at least 60%, such as in at least 70%, such as in at least 80%, such as in at least 80% of the nodules of the plant nodules of said plant grown in said soil when said rhizobia strain is present in the soil.

41. The method according to any one of the previous claims, wherein the soil has been obtained from a field before plant has been sown in the soil.

42. The method according to any one of the previous claims, wherein the soil has been obtained no more than 5 months prior to the sowing of the plant, such as 4 months, such as 3 months, such as 2 months, such as 1 month, such as 3 weeks, such as 2 weeks, such as 1 week, such as 3 days, such as 1 day prior to the sowing of the plant.

43. A customized inoculant formulation wherein the customized inoculant formulation comprises a preferred rhizobia strain identified by the method according to any one of the preceding claims.

44. The customized inoculant formulation according to claim 43, wherein the customized inoculant comprises the preferred rhizobia strain at a concentration of at least 1 x 10A5 cfu / ml , such at least 1 x 10A6 cfu / ml, such at least 1 x 10A7 cfu / ml, such at least 1 x 10A8 cfu / ml, such at least 1 x 10A9 cfu / ml, such at least 1 x 10A10 cfu / ml.

45. The customized inoculant formulation according to any one of claims 43 to 44, wherein the customized inoculant further comprises a polymer for seed coating.

46. The customized inoculant formulation according to claim 45, wherein the polymer is a biodegradable polymer.

47. The customized inoculant formulation according to claim 46, wherein the biodegradable polymer is a cellulose derivative.

48. A method for providing / manufacturing a customized inoculant formulation as defined in any one of claims 43 to 47 comprising the method according to any one of claims 1 to 40 and formulating the customized inoculant formulation with the at least one preferred rhizobia strain selected in step h.

49. A method to reduce the number of rhizobia strains to be tested to formulate a customized inoculant formulation for a plant of interest grown in a soil of interest, wherein said method comprises steps a. to d. of the method for providing a customized inoculant formulation for a plant grown in a soil according to any one of claims 1 to 40, and testing the rhizobia predicted to have high effectiveness and / or competitiveness for the combination of said soil of interest and plant of interest.

Citation Information

Patent Citations

  • Methods and compositions for improving plant traits

    WO2018132774A1

  • Polymer compositions with improved stability for nitrogen fixing microbial products

    WO2020118111A1

  • Methods and systems for predicting crop features and evaluating inputs and practices

    WO2022165101A1

  • Methods and systems for assessing agriculture practices and inputs with time and location factors

    WO2022212156A1