Feature-vector-based prediction of interactions between plant and pathogen proteins

The feature-vector-based method efficiently predicts PPIs between plant and pathogen proteins, addressing the limitations of existing methods by providing scalable and accurate predictions, enabling the identification of pathogen-resistant cultivars and improving model transparency.

WO2026073923A1PCT designated stage Publication Date: 2026-04-09KWS SAAT SE & CO KGAA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing methods for predicting protein-protein interactions (PPIs) between plant and pathogen proteins are time-consuming, computationally expensive, and lack accuracy, particularly in the context of plant infection, and current databases and tools are not suitable for this specific application.

Method used

A feature-vector-based method using machine learning models to predict PPIs by transforming protein sequences into vectors, applying model interpretation software to identify relevant amino acids, and iteratively improving the model with empirical validation and retraining.

Benefits of technology

This approach allows for efficient, scalable, and accurate prediction of PPIs, facilitating the identification of pathogen-resistant plant cultivars and reducing the need for agrochemicals, while enhancing model transparency and accuracy through model interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025077993_09042026_PF_FP_ABST
    Figure EP2025077993_09042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for predicting protein-protein interactions between plant proteins and pathogen proteins. The method comprises: by one or more computing devices: a) transforming (102) one or more protein sequences (208) of a plant into a respective plant protein feature vectors (302; 306) and transforming (104) one or more protein sequences (210) of a pathogen into a respective pathogen protein feature vectors (304; 308); b) forming (106) a plurality of different feature vector pairs (313) respectively comprising one (312) of the pathogen protein feature vectors and one (310) of the plant protein feature vectors; c) providing (108) the vector pairs to a trained machine learning model (212), and in response receiving (110), for the pairs of feature vectors, an indication (322) if at least one output protein pair represented by one of the feature vector pairs is predicted by the trained machine-learning model to interact.
Need to check novelty before this filing date? Find Prior Art

Description

1 KWS.224.01WO / KWS0460FEATURE-VECTOR-BASED PREDICTION OF INTERACTIONS BETWEEN PLANT ANDPATHOGEN PROTEINSFIELD OF THE INVENTION

[0001] The invention relates to the field of predicting protein-protein interactions between plant and pathogen proteins.BACKGROUND

[0002] Within biological cellular processes, proteins fulfill various functions and very often form stable or unstable complexes with themselves or with other proteins. The experimental validation of these protein interactions by molecular biology methods is very time consuming and laborious. To increase the probability of success and / or to reduce the need for experimental validation, various methods to calculate the probability of two proteins interacting exist.

[0003] Protein-Protein-Interactions (PPIs) play a crucial role in pathogen infection processes when a microorganism, such as Cercospora beticola, secretes its proteins in order to modulate the host cell and to induce an environment that allows the colonization of the host. These interactions are induced by small, secreted proteins called effectors. These interactions can be on the one hand beneficial for the host when it is able to recognize the pathogen and induce countermeasures leading to resistance but also on the other hand harmful because some PPIs will suppress countermeasures and modulate the host cells in a way that enables the microorganism to colonize them. So, while some PPIs can help a plant defend itself against a pathogen, other PPIs will ultimately lead to plant disease and yield loss.

[0004] Existing PPI prediction tools and databases are not suitable for predicting PPIs relevant in the context of plant infection by pathogens and have not been used for a variety of reasons.2 KWS.224.01WO / KWS0460

[0005] Often, the prediction algorithms are so time-consuming and computationally expensive that they are not publicly available. Only the pre-computed prediction results are published in respective PPI databases.

[0006] For example, STRING (Search Tool for the Retrieval of Interacting Genes / Proteins) is a database of known and predicted protein-protein interactions (PPIs). The interactions include direct (physical) and indirect (functional) associations among many different prokaryotic and eukaryotic organisms, only some of them being plants; they stem from computational prediction, from knowledge transfer between organisms, and from interactions aggregated from other (primary) databases. A downside of this and similar databases is that they rely on a limited set of known or predicted interactions, which might not cover all PPIs of interest. Predictions may have high false positive rates due to the integration of indirect evidence.

[0007] The PrePPI database (Petrey D, Zhao H, Trudeau SJ, Murray D, Honig B. PrePPI: A Structure Informed Proteome-wide Database of Protein-Protein Interactions. J Mol Biol. 2023 Jul 15;435(14):168052. doi: 10.1016 / j.jmb.2023.168052. Epub 2023 Mar 17. PMID: 36933822; PMCID: PMC10293085) comprises a large set of predicted PPIs. However, the predictions are based on structural data. Structural data is often not available, and structure-based predictions can be computationally intensive. Furthermore, the database covers the human interactome and can therefore not be used for identifying or predicting plant-pathogen-PPIs.

[0008] AlphaFold is an artificial intelligence (Al) program developed by DeepMind, which performs predictions of protein structure. The predicted structures can then be used for predicting protein interactions. The program is designed as a deep learning system. Some researchers noted that the accuracy is not high enough for a large share of its predictions. Furthermore, structure-based predictions are computationally expensive and hence often not available for organizations or projects with limited computational resources.

[0009] Various other approaches for predicting PPI interactions, typically focusing on interactions with human proteins, have been described. An example is Kaundal R et al.: "deepHPI: a comprehensive deep learning platform for accurate prediction and visualization of host-pathogen protein-protein interactions" , Brief Bioinform. 2022 May3 KWS.224.01WO / KWS046013;23(3):bbacl25, doi: 10.1093 / bib / bbacl25, PMID: 35511057. A still further approach for predicting human-virus PPIs has been described in Pengfei Xie et al., "Emvirus: An embedding-based neural framework for human-virus protein-protein interactions prediction", Biosafety and Health, Volume 5, Issue 3, 2023, Pages 152-158, ISSN 2590- 0536, https: / / doi.Org / 10.1016 / i.bsheal.2023.04.003.SUMMARY OF THE INVENTION

[0010] It is an objective to provide for a system and method for predicting proteinprotein interactions between plant proteins and pathogen proteins. The objectives underlying the invention are solved by the features of the independent claims.

[0011] In one aspect a method for predicting protein-protein interactions between plant proteins and pathogen proteins is disclosed. The method comprises: a) transforming one or more protein sequences of a plant into a respective plant protein feature vector and transforming one or more protein sequences of a pathogen into a respective pathogen protein feature vector; b) forming a plurality of different feature vector pairs respectively comprising one of the pathogen protein feature vectors and one of the plant protein feature vectors; c) providing the vector pairs to a trained machine learning model, and in response receiving, for the pairs of feature vectors, an indication of whether at least one output protein pair represented by one of the feature vector pairs is predicted by the trained machine-learning model to interact.The method may be executed by one or more computing devices.

[0012] In the context of plant-pathogen-PPI-prediction, the above-mentioned featurevector based prediction approach might offer several benefits over other approaches, in particular protein-structure based prediction approaches. These benefits could include efficiency, scalability, and broader applicability across diverse datasets. For example, feature-vector based methods may leverage sequence information to create numerical representations of proteins, which can then be processed quickly by the machine learning model. This approach might be computationally less intensive compared to structure-based methods, which typically require detailed 3D structural data and4 KWS.224.01WO / KWS0460 complex modeling techniques. 3D structural data is often not available, and predicted 3D structures are often incorrect while their prediction is computationally expensive. In the context of PPI prediction, errors in a predicted 3D structure are particularly damaging, as an incorrect prediction in one of the two potential interaction partners is sufficient to render the entire prediction incorrect.

[0013] For example, screening large amounts of data using molecular dynamics approaches or docking programs are computationally very costly. Even AlphaFold multimer takes minutes to hours for each combination for a set of protein sequence pairs which can be evaluated within a few minutes or even seconds using the feature-vector based approach. The applicant has observed that based on the feature-vector based approach, it may be possible to screen thousands of protein combinations to find PPIs of interest within minutes, which can then be further validated in the lab and / or used for selecting plants for breeding projects.

[0014] Feature-vector based predictions might be more scalable as they could handle large datasets of protein sequences more easily than structure-based methods. This scalability may allow the analysis of vast amounts of protein data quickly, facilitating the discovery of new interactions and accelerating research progress.

[0015] Moreover, feature-vector based predictions might integrate diverse types of sequence-derived information, such as functional, physicochemical and / or structural properties, potentially capturing a wider array of interaction determinants than structure-based methods alone. The feature-vectors may comprise local, amino-acid specific features and optionally also more coarse grained or global features such as amino acid patterns or protein motifs, offering a more accurate prediction of PPIs.

[0016] The applicant has observed that the above-mentioned approach may be particularly suited for predicting plant-pathogen-PPIs. Hence, the method may be used for identifying and selectively breeding plant cultivars that are more robust against a pathogen, e.g. because they provide a particular strong PPI with a pathogen protein which triggers a defense reaction, or because they provide a particular weak (or even absent) PPI with a pathogen protein which is required by the pathogen for infecting or spreading in the plant.5 KWS.224.01WO / KWS0460

[0017] Currently, it is common practice for farmers to use agrochemicals (pesticides) against pathogens with have a negative impact on health and the environment. An alternative to pesticides is crop-rotation, but this often requires complex management to effectively implement different crop sequences and maintain soil health and may necessitate investment in specialized equipment for different types of crops. The use of disease-resistant varieties is often a more environmentally friendly option, but such varieties may not exist for every type of pathogen, or may have limited resistance to a particular pathogen. The feature-vector based prediction of plant-pathogen-PPIs may make the identification and / or development of pathogen-resistant cultivars significantly faster and easier, as the method may allow identifying plants having a particularly strong, desired PPI or lacking an undesirable PPI with respect to a pathogen protein.

[0018] The received indication can be, for example, an indication that the at least one output protein pair is predicted to interact. The indication may have the form of a binary output such as true (the protein pair is predicted to interact) or false (the protein pair is predicted not to interact). Preferably, the indication is provided in the form of a numerical value within a predefined value range that indicates the predicted likelihood that the proteins of the examined protein pair will interact. For example, the likelihood can be indicated as a score value, e.g. as a value between 0 and 1.0, or by a percentage value.

[0019] According to some examples, the method further comprises applying a model interpretation software to identify a sub-set of specific amino-acid residues in the plant protein sequence and pathogen protein sequence of the at least one output protein pair that had the highest impact on the prediction that the proteins of the at least one output protein pair will interact.

[0020] For example, the model interpretation software may be used to identify a sub-set of amino acids, e.g. amino acids that constitute or are part of the interaction site (also referred to as binding site), that constitute the reason for the machine learning model predicting that the currently analyzed pair of proteins will interact.

[0021] Using a model interpretation software for identifying amino acids which are considered by the machine learning model to be highly relevant for PPI interaction may have the advantage that also amino acids which are not be part of the interaction site6 KWS.224.01WO / KWS0460 but nevertheless have an important impact on PPI formation are identified. This is not possible in structure-analysis based methods which solely focus on identifying the interaction site. Hence, the model interpretation software is able to identify also "remote" amino acids relevant for PPI formation.

[0022] According to some examples, the model interpretation software is applied on the trained machine learning model and on at least one output protein pair.

[0023] For example, the application of the model interpretation software can comprise interrogating the machine-learning model, while it is processing or after it has processed the protein pair, to trace back which amino acid residues or amino acid residue patterns had the strongest effect on the prediction of the machine-learning model that this protein pair will interact. For example, the output protein pair may comprise amino-acid specific probability values for the respective amino acid being part of a PPI, and the application of the model interpretation software on the protein pair may comprise interrogating the machine-learning model, while it is processing or after it has processed the protein pair, to trace back which amino acid residues or amino acid residue patterns had the strongest effect on the prediction of the amino-acid specific probability values.

[0024] This may allow aligning the internal functions of the machine-learning model and how it computes its prediction with the prediction result, i.e., a qualitative and / or quantitative prediction that a particular protein pair will form a PPI. For example, the prediction output by the machine-learning model, e.g. in the form of an annotated output protein-protein pair, may comprise a likelihood of PPI formation for each amino acid residue of a protein sequence with respect to the other protein sequence of a protein-protein pair (e.g., the output pair for which a PPI was predicted), and the aligning or interrogation may comprise analyzing the contribution that a particular amino acid or amino acid sub-sequence or a discontinuous amino-acid pattern may have on the predicted likelihoods. Hence, while the trained machine learning model may be configured to compute likelihoods of PPI formation for the individual amino acids, the model interpretation software may be configured to predict relevance scores for a particular prediction result. It is important to note that an amino acid sub-sequence deemed by the model to be highly relevant for a prediction may be a different amino7 KWS.224.01WO / KWS0460 acid sequence than the amino acids predicted by the machine-learning model to be part of the binding site.

[0025] According to some examples, the model interpretation software is or comprises software separate from the machine learning model configured for predicting the PPI. The software PoSHAP may be an example for this approach.

[0026] According to some examples, the model interpretation software is or comprises an integral part of the machine learning model. For example, the integral part may be or comprise attention layers.

[0027] According to some examples, the trained machine-learning model is configured to compute, for each amino acid of the at least one output protein pair, a likelihood of protein-protein formation, and wherein the model interpretation software is configured to compute relevance scores that quantify the contribution of each amino acid residue to the predicted likelihood of protein-protein formation.According to some examples, the model interpretation software is selected from a group comprising: a machine learning model comprising long short-term memory - LSTM - layers; a software program or module implementing Layer-Wise Relevance Propagation - LRP; a software program or module implementing Local Interpretable Model-Agnostic Explanations - LIME; a software program or module implementing an attention mechanism, in particular one or more attention layers integral to a neural network structure comprising the machine learning model; a software program or module configured to make use of permutation importance; a software program or module configured to compute SHapley Additive exPlanations - SHAP - values for model interpretation and model description, wherein the SHAP values can in particular be KernelSHAP values or positional SHAP values - PoSHAP values; and a software program or module configured to perform Gradient-weighted Class Activation Mapping, Grad-CAM.8 KWS.224.01WO / KWS0460

[0028] According to some examples, machine-learning-model-independent software programs are used for identifying potential interaction-residues of a protein, such as molecular dynamics simulations at the atomic level (e.g. GROMACS: Pronk S, et al: "GROMACS 4.5: a high-throughput and highly parallel open-source molecular simulation toolkit", Bioinformatics, 2013 Apr l;29(7):845-54. doi: 10.1093 / bioinformatics / btt055. Epub 2013 Feb 13. PMID: 23407358; PMCID: PMC3605599), or docking programs that predict the orientation of two interaction partners together with the relevant amino acids (e.g, RosettaDock: Marze NA, Ret al.: "Efficient flexible backbone protein-protein docking for challenging targets", Bioinformatics, 2018 Oct 15;34(20):3461-3469, doi: 10.1093 / bioinformatics / bty355, PMID: 29718115; PMCID: PMC6184633). However, the applicant has observed that using model interpretation software for identifying amino acids having caused the machine-learning model to predict that two proteins will interact provides several important advantages:

[0029] Firstly, introducing an additional source of error is avoided. The software used for predicting if two proteins will interact and the model-independent software for identifying binding residues in a protein may have been created using different types of proteins and / or training data sets, may be based on different assumptions and have different biases which may result in an unpredictable, non-linear increase of biases and errors.

[0030] Second, the likelihood of a conflicting / contradictory prediction is avoided or at least reduced. Such a situation may occur if a PPI is predicted by the machine learning model, but the model-independent software for identifying binding residues does not identify any potential binding residues. These problems may be avoided by using a model interpretation software for identifying the amino acids most relevant for a predicted PPI.

[0031] A further important advantage may be that the model interpretation software may in some cases be able to determine that an unusually large or an unusually small number of amino acids were identified to be relevant for a predicted interaction site. In this case, the predicted PPI were identified as an artifact or as a less-reliable prediction, thereby avoiding costs associated with conducting breeding projects or expensive empirical tests of the predicted PPI.9 KWS.224.01WO / KWS0460

[0032] A still further advantage may be that the model interpretation software may in some cases be used for expanding the number of protein sequences and protein sequence pairs to be examined by permuting the interaction-causing amino acids identified by the model interpretation software. The expanded set of protein sequences and protein sequence pairs may be processed as described before, thereby potentially identifying protein variants with a modified interaction-site which has an even stronger, desired PPI, or an even weaker, undesired PPI than the protein-protein pair used for modifying the interaction-site amino acids of the plant protein and / or of the pathogen protein. In some cases, the PPI of one or more plant-protein-pairs of the expanded set of proteins may be empirically validated and the results may be used for re-training the machine-learning model. If certain amino acids or features are found to disproportionately influence the prediction result, developers can investigate and adjust the machine-learning model, e.g. by retraining the model on a less-biassed training data set, to ensure unbiassed and accurate predictions. Hence, the model interpretation software in this case may be used for improving the prediction accuracy of the machinelearning model by specifically expanding the training data with sequence variants that comprise variations selectively in sub-regions that are particularly relevant for a predicted PPI.

[0033] Depending on the implementation and use-case scenario, retraining of the machine learning model may be performed at many different points in time. The retraining process may be initiated automatically by the system or manually by a user. For example, retraining may be carried out whenever additional training data becomes available, such as after sequencing a further plant species, a new plant variety, or a novel pathogen. In another example, the predictions generated by the trained machine learning model may be continuously or periodically compared with experimentally measured protein-protein interaction values. If it is determined that the deviation between the predicted values and the measured values exceeds a predefined threshold, retraining may likewise be initiated. Such deviations may arise, for instance, if the proteomes of the investigated plant species or pathogen species differ substantially from the proteomes of the data used during initial training. In these cases, retraining may10 KWS.224.01WO / KWS0460 restore or improve the predictive accuracy of the model, thereby ensuring that the system remains adapted to newly emerging or significantly altered biological data.

[0034] As mentioned above, a model interpretation software may be implemented using many different approaches in order to interpret how inputs lead to specific outputs of a trained predictive machine learning model. One possible approach is the use of LSTM models, or model-specific interpretation strategies, such as layer-wise relevance propagation, LRP, (as described in, for example, Montavon, Gregoire & Binder, Alexander & Lapuschkin, Sebastian & Samek, Wojciech & Muller, Klaus-Robert (2019) Layer-Wise Relevance Propagation: An Overview. 10.1007 / 978-3-030-28954-6_10) or the attention mechanism (Bahdanau D, Cho K, Bengio Y. Neural Machine Translation by Jointly Learning to Align and Translate, ArXivl4090473, 2016 May 19) or permutation importance. Attention mechanisms, however, require that the model be constructed with attention layers. This may limit the flexibility of model architecture of the machine learning model used for predicting the PPIs.

[0035] According to one example, the model interpretation software may be a machine learning model comprising long short-term memory (LSTM) layers. LSTM networks are a class of recurrent neural networks specifically designed to capture long-range dependencies within sequential data by maintaining and updating an internal cell state through input, forget and output gates. In the context of protein sequences, an LSTM can therefore capture dependencies between amino acids that are distant in sequence but functionally related. A model interpretation software may be configured to interrogate the activations of the LSTM cells or to apply attribution techniques, such as gradientbased relevance scoring rules, to identify which residues at which positions contributed most significantly to the interaction prediction for a given protein pair.

[0036] Another possible example for a model-interpretation software is a software module implementing Layer-Wise Relevance Propagation (LRP). LRP is a backward decomposition technique that starts from the output of the predictive model (e.g. a score indicating the likelihood of protein-protein interaction for an individual amino acid and / or for the whole sequence) and propagates this "relevance" backwards through the network layers down to the input level. The propagation is performed according to layerspecific rules (e.g. E-rule, a|3-rule), thereby ensuring that the relevance assigned to the11 KWS.224.01WO / KWS0460 input features is consistent with the prediction of the model. When applied to protein sequence inputs, LRP allows to assign a numerical relevance score to each amino acid or token embedding, which directly indicates its contribution to the predicted PPI. Thus, the model interpretation software can provide a residue-level series of relevance scores identifying those parts of the protein sequences that had the highest impact on the model's output.

[0037] A further possible example for a model interpretation software is a software program or module that implements Local Interpretable Model-Agnostic Explanations (LIME). LIME is a method for explaining the predictions of complex machine-learning models. The LIME method comprises approximating the machine-learning model locally around a specific prediction with a simpler and interpretable surrogate model, typically a sparse linear model. To do this, the input provided to the surrogate model is perturbed (for example, by masking or altering features, e.g., by exchanging one or more amino acids in the plant or pathogen protein of a protein pair for which a PPI was predicted). The LIME method further comprises observing the resulting changes in the model's predictions, and then fitting the surrogate model to these perturbations. The resulting coefficients are interpreted as local importance weights that describe how much each feature contributes to the specific prediction. Since LIME is model-agnostic, it can be applied regardless of the underlying algorithm used.

[0038] For example, in the context of protein-protein interaction (PPI) prediction, the inputs are amino acid sequences or derived features from them. By perturbing subsequences or individual residues and observing how the predicted PPI probability changes, LIME assigns weights to individual amino acids or amino acid features. High positive values indicate residues that strongly support the predicted interaction, while strongly negative values identify residues that argue against it. In this way, LIME values can be used to pinpoint amino acids or motifs that are particularly relevant for binding and may represent biologically meaningful interaction hotspots.

[0039] As LIME builds a local surrogate model around a single prediction and thus provides explanations that are locally faithful but not necessarily consistent across the entire model, the LIME can be considered a relatively simple and computationally efficient approach though sensitive to the way perturbations are generated. SHAP, in12 KWS.224.01WO / KWS0460 contrast, is grounded in Shapley values from cooperative game theory and distributes the difference between the actual prediction and a baseline fairly among all features. SHAP values obey desirable properties such as consistency and additivity, making them globally comparable across predictions. However, this may come at the cost of higher computational complexity. Hence, in practice, using a software module implementing the LIME for providing the model interpretation software may be particularly useful when quick, local explanations are needed for why a model predicted an interaction for a specific amino acid sequence, while SHAP-value based approaches may be preferable when consistent and globally interpretable feature attributions across many predictions are required.

[0040] A further possible example for a model interpretation software is the use of an attention mechanism, typically attention layers in a neural network structure. Attention mechanisms compute, for each element in a sequence, a weight that reflects its relative importance in the prediction task. The attention mechanism is configured to "align" source and target tokens. When applied to protein sequence analysis, attention layers can learn which amino acids in one protein should be weighted most strongly when evaluating a possible interaction with residues of the other protein. The model interpretation software may therefore read out these attention weights and present them as an interpretable mapping between input residues and predicted interaction likelihood. It is noted, however, that attention mechanisms may require that the machine-learning model used to compute the PPI predictions itself be constructed with attention layers.

[0041] A still further example for a model interpretation software is software that is configured to make use of permutation importance. Permutation importance is a modelagnostic technique that measures the impact of individual features on the performance of the model. The principle is to permute (shuffle) the values of a given feature across samples, thereby destroying its correlation with the target output, and then to reevaluate the predictive performance of the model. A significant drop in performance upon permutation indicates that the feature is of high importance. In the present context, amino acid positions or derived features can be permuted, and the resulting loss in accuracy or predictive power can be quantified. This allows the model interpretation13 KWS.224.01WO / KWS0460 software to assign a global importance score to features, thereby identifying amino acids or motifs that are critical for the overall predictive capability of the trained model. The permutation importance-based approach for identifying amino acids which are of particular relevance for the machine-learning model for determining whether or not two proteins interact should not be confused with a later performed, optional step of permuting the amino-acid subset that was identified by the model interpretation software with the purpose of optimizing a PPI. While the purpose of using a model interpretation software may be to identify a small set of amino acids which were the basis for the machine-learning model to perform a particular prediction for a particular protein-protein pair, the purpose of permuting the amino acids of this already identified sub-set may be to identify sequence variants with increased or decreased binding strength.

[0042] Each of these approaches may provide a different pathway for implementing a model interpretation software. LSTM-based analysis may allow sequential dependencies to be highlighted; LRP may provide residue-level attributions consistent with the model's output; attention mechanisms may allow direct visualization of focus weights within the model, provided the architecture supports it; and permutation importance may offer a model-agnostic evaluation of global feature importance. In all cases, the outcome of applying the model interpretation software is the identification of amino acid residues or regions in the plant and / or pathogen protein sequences that are most relevant for a predicted protein-protein interaction, thereby increasing transparency of the prediction process and enabling targeted downstream applications such as sequence optimization, e.g. by further computationally permuting the identified sub-sequence for further sequence optimization, and / or protein engineering, and / or selective breeding or other use case scenarios.

[0043] According to preferred examples, the model interpretation software uses SHAP values for model interpretation and model description. For example, the software PoSHAP can be used as model interpretation software. This software is described in greater detail in "Positional SHAP (PoSHAP) for Interpretation of machine learning models trained from biological sequences", Quinn Dickinson, Jesse G. Meyer, January 28, 2022, PLOS Computational Biology, https: / / doi.org / 10.1371 / journal.pcbi.1009736. The14 KWS.224.01WO / KWS0460 software utilizes SHapley Additive exPlanations (SHAP) to generate positional model interpretations and includes three long short-term memory (LSTM) regression models that predict peptide properties, including interaction (or "binding") affinity, and collisional cross section (CCS) measured by ion mobility spectrometry. The SHAP values are able to explain the output of any machine learning model. The SHAP values determine how each input provided to a machine-learning model (configured to predict the presence or likelihood of a PPI formation) alters the model's prediction. SHAP values are derived from SHapley values, which reflect the contributions of the inputs to the predicted output plus a baseline.

[0044] SHAP values may have the advantage of offering a rigorous and model-agnostic approach, grounded in cooperative game theory, and are able to quantify both the marginal and joint contributions of amino acids to a predicted protein-protein interaction. SHAP may have the advantage of being universally applicable to essentially any predictive architecture and provides locally faithful explanations of single predictions. However, the exact computation of SHAP values can be computationally demanding, particularly when applied to high-dimensional protein sequences, as the exact computation of Shapley values scales combinatorially with the number of input features. Approximation strategies such as KernelSHAP or PoSHAP may be used to mitigate this limitation.

[0045] For example, in order to evaluate the contribution of the individual amino acids of the plant and / or of the pathogen protein in a protein pair, the following steps may be performed: the protein pair of interest is selected. For example, this pair may be the one for which the highest likelihood of PPI formation was predicted. Then, variants of the plant protein sequence or of the pathogen protein sequence are computed, e.g. by replacing one or more of the amino acids of the respective protein sequences by other amino acids. Based on these plant and / or pathogen protein sequence variants, a plurality of pathogen-plant-protein sequence pair variants is formed and provided as input into the machine learning model for predicting the likelihood of a PPI of the protein sequence variant pair. In these sequence pair variants, only the protein sequence or only the pathogen sequence or both protein sequences may have been modified compared to the initially selected protein pair of interest. The set of protein sequence pair variants may15 KWS.224.01WO / KWS0460 then be provided sequentially to the machine learning model for having this model predict the likelihood of PPI formation of this protein sequence pair variant. Some of these sequence pair variants may have a higher likelihood of PPI formation than the initially selected protein pair, others may have a lower predicted likelihood.

[0046] The multiple protein sequence pairs variants, their predicted likelihood of PPI formation, and the machine-learning model for the PPI prediction may be provided as input to the model interpretation software for enabling the model interpretation software to identify the sub-set of amino acids (including their position) which contributed the most the model predicting the presence of a PPI.

[0047] The plurality of pathogen-plant-protein sequence pairs variants and the trained machine learning model having predicted the PPIs may be input into SHAP's KernelExplainer method of the PoSHAP program. The contribution of each amino acid at each position of each of the two proteins of each pair are stored in an array and the mean SHAP value of each amino acid at each position is calculated. The KernelExplainer tracks the contribution of each amino acid in each sequence pair variant to the predicted PPI likelihood, and outputs, e.g., via a heatmap, the sub-set of amino acids in the plant protein and the pathogen protein having the strongest impact on the PPI prediction of the model. A heatmap is a graphical representation of data where individual values or value ranges are encoded in and shown as colors. For example, the heatmap may have the form of a matrix plot or table where individual cells represent amino acids of a protein sequence and wherein each cell in the heatmap displays a value with a color. In a heat map, different colors indicate varying levels of intensity or quantity. Instead of colors, different grey levels or hatchings may also be used.

[0048] PoSHAP is an attractive option for PPI prediction models, because the SHAPvalue based approach of PoSHAP can dissect interactions between inputs, for example when inputs are correlated, and can be used with any arbitrary model. By providing clear, efficient, and accurate explanations for the predictions of the machine-learning model, PoSHAP enhances the transparency and reliability of the predicted PPIs, and in particular allows to identify the sub-set of amino acids in each of the proteins of the examined protein pair which have been responsible for the prediction of the machine-learning16 KWS.224.01WO / KWS0460 model that this protein pair will interact. SHAP values are calculated on a per input basis, thereby helping to understand the model's "reasoning" behind any PPI prediction.

[0049] As mentioned above, another example for a model interpretation software is Layer-wise Relevance Propagation. Layer-wise relevance propagation is an approach and software framework which allows to decompose the prediction of a neural network computed over a sample input, e.g. an image or a sequence of symbols or features, down to relevance scores for the single input dimensions of the sample such as subpixels of an image, sub-sets and sub-sequences of symbols or features. Layer-wise Relevance propagation has been described for example in "Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers", Alexander Binder, et al., arXiv:1604.00825, https: / / doi.org / 10.48550 / arXiv.1604.00825. Still further examples of a model interpretation software is the Grad CAM approach transferred from image processing to protein feature vector processing (see "Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization", Ramprasaath R. Selvaraju, et al., arXiv:1610.02391 , https: / / doi.org / 10.48550 / arXiv.1610.02391.

[0050] Preferably, the model interpretation software is model-agnostic, meaning it can be applied to basically any machine learning model. This versatility is crucial for its adoption in various domains and for different types of models. PoSHAP is a modelagnostic model interpretation software.

[0051] In some examples, applying the model interpretation software comprises computing relevance scores for the amino acid residues in the plant protein sequence and the pathogen protein sequence. The relevance scores are scores which indicate the contribution of the amino acid residues to the likelihood of the output protein pairs to interact. The method further comprises using the relevance scores for identifying the sub-set of the amino acid residues.

[0052] The calculation of the relevance scores may have the advantage of providing finegrained information on the relevance of the individual amino acids of each of the two proteins of the pair for the predicted PPI, thereby easing the selection of the amino acids to be modified in order to increase or decrease the strength of the PPI. For example, if a relevance score indicates that a particular amino acid at a particular position in the plant17 KWS.224.01WO / KWS0460 protein is the most relevant for the predicted PPI, then it may be advisable to exchange specifically this amino acid in order to prevent the formation of the predicted PPI.

[0053] According to some examples, the method comprises visualizing the relevance scores in the form of a heatmap overlaid on the plant protein sequence and / or the pathogen protein sequence of the respective at least one output protein pair. This may provide a particularly intuitive way of highlighting the sub-set of amino acids based on which the machine-learning model has predicted the existence of a PPI. This may ease the decision which ones of the highlighted amino acids should be modified experimentally for increasing or decreasing the strength of the predicted PPI.

[0054] According to some examples, the method comprises using experimental protein engineering for replacing one or more of the amino acids in the identified sub-set for increasing or decreasing the likelihood of interaction of the respective at least one output protein pair.

[0055] Protein engineering is the process of developing proteins by modifying amino acid sequences of existing proteins to alter protein structure and / or function. This may be achieved through various techniques such as directed evolution and rational design. Often, genetic engineering techniques are used for protein engineering. Genetic engineering is the direct manipulation of an organism's genes using biotechnology. It involves altering the genetic material of an organism to achieve desired traits or outcomes.

[0056] The accuracy of the identification of the sub-set of specific amino acid residues having the highest impact on the prediction may be tested empirically, e.g. by conducting Yeast Two-Hybrid (Y2H) Assays, Co-Immunoprecipitation (Co-IP) tests, Pull-Down Assays or Fluorescence Resonance Energy Transfer (FRET) tests. The results of the tests may be used for improving the trained machine-learning model.

[0057] For example, the empirically verified information if the two proteins interact may be assigned as labels to the two proteins and may be added to the training data set used to generate and train the machine learning model. In case the empirical test revealed that the predicted PPI was a false positive prediction, a retraining on the supplemented training dataset may ensure that the model learns to better assess the likelihood that the two proteins of this pair will interact, and will no longer compute this false positive18 KWS.224.01WO / KWS0460 prediction for these and similar protein sequence pairs. The information regarding interaction between two proteins are provided, according to some examples, as binary, e.g., 0 for "empirically verified not-interacting" or 1 for "empirically verified interacting". According to other embodiments, the labels are or comprise (e.g. in addition to the binary label) a floating point number representing the empirically observed interaction strength.

[0058] Preferably, all empirically verified protein pairs, whether positive or negative, are treated in the training or re-training of the machine learning model, as equally important.

[0059] According to some examples, the method further comprises applying the model interpretation software during the training of the machine-learning model after each training epoch to monitor evolution of residue relevance during model training.

[0060] For example, if the model interpretation software initially assigns a high relevance score to an amino acid sub-sequence known to be incapable of forming a protein-protein interaction (PPI) (e.g., because it is located inside the protein), the training process is considered to be on the right track if, in later epochs, a different subset of amino acids, known to be a common PPI motif, is assigned a high relevance score. If, however, the relevance score is assigned to different amino acids in each epoch— i.e., it never converges— or if there is a more or less homogeneous distribution of relevance scores across all or most of the amino acids in a sequence, the training process is considered potentially incomplete, biased, or erroneous. In such cases, the training may be aborted and / or repeated with a modified or enriched training dataset.

[0061] According to some embodiments, the machine-learning model is trained and / or retrained also based on feedback from the model interpretation software to enhance a prediction accuracy of the machine-learning-model.

[0062] The feedback from the model-interpretation software can be, or can comprise, for example, one or more additional pairs of plant and pathogen protein sequences, wherein the plant protein sequence and / or the pathogen protein sequence of these additional pairs are variants which have been computationally generated by automatically permuting amino acids of the sub-set of amino acids identified by the model interpretation software as the ones having the highest impact on the machine- learning-model's prediction.19 KWS.224.01WO / KWS0460

[0063] For example, the automated permutation of amino acids may comprise creating a variant of the plant protein sequence (and / or of the pathogen protein sequence) of the output at least one protein by replacing one or more amino acids of the sub-set of amino acids in said protein sequence. As a consequence, one or more plant-protein-protein sequence pair variants are computed starting from the at least one output plant pathogen protein pair. Then, this "virtual", computationally generated plant-pathogen- protein sequence pair, also referred herein as plant-pathogen protein sequence pair variant, is used as input of the machine learning model for predicting the likelihood that said sequence pair variant will form a PPI. If the likelihood of this protein pair sequence variant is above a first likelihood threshold, this computationally generated protein sequence pair is assigned a label indicating that this protein pair forms a PPI. If the likelihood of this protein pair sequence variant is below a second likelihood threshold, this computationally generated protein sequence pair is assigned a label indicating that this protein pair does not form a PPI. The one or more protein sequence pair variants, annotated automatically using the machine learning model as described above, may be used for enriching the training data set with additional information, i.e., with additional protein sequence pairs computationally annotated as forming or not forming a PPI. Then, the machine-learning model is retrained on the enriched training data set. The first and second thresholds may be predefined and may depend on the nature of the two proteins of each pair and their typical binding strength.

[0064] The feedback information used for enriching the training data set may in addition comprise empirical data obtained by empirically validating the amino acids involved in the PPI, and / or by empirically validating the existence or strength of the predicted PPI. For example, the protein sequence variants and protein sequence pair variants generated computationally by permuting amino acids in the sub-set of amino acids may be used for synthesizing these protein variants via a protein engineering technique, and measuring the tendency of the protein pair variant to form a PPI in vitro.

[0065] Hence, the feedback may comprise protein pair variants generated by modifying one or more of the amino acids identified as the ones having the highest impact on the learning-model's prediction, and empirical data obtained when validating PPIs of these protein pair variants. Thereby, an iterative loop of using the machine learning model for20 KWS.224.01WO / KWS0460 identifying protein pairs with a high likelihood of forming a PPI, and using a model inspection software for identifying particularly relevant amino acids whose permutation will highly likely have an impact on the PPI likelihood predicted by the trained machine learning model, can be created. These protein pair variants can be used as input for the machine learning model for automatically computing labels, and the automatically labeled protein sequence pair variants can be used for expanding and enriching the training data set and for repeating the training on the enriched training data set, thereby iteratively improving the accuracy of the machine learning model.

[0066] These cyclic iterations, which can be executed automatically or semi- automatically, may lead to an automatic improvement of the accuracy of the machine learning model. A pseudocode of this automatic or semi-automatic model improvement approach can be described as follows: i. Use an existing set of plant and pathogen protein sequence pairs annotated as forming a PPI or not forming a PPI; ii. Train a machine-learning model on the training data to provide a machine learning model configured to predict if a plant-pathogen-protein sequence pair provided as input will likely form a PPI or not; iii. Execute the trained machine learning model on a plurality of plant-pathogen- protein pairs provided as input to the trained machine learning model ; iv. Output at least one plant-pathogen protein pair predicted to form a PPI; v. Apply the model interpretation software on the plant protein sequence and / or the pathogen protein sequence of the output protein sequence pair for identifying, in at least one of said two protein sequences, a sub-set of amino acids having the highest impact on the prediction result output by the trained machine learning model ; vi. Generate one or more protein sequence pair variants of the at least one output plant-pathogen protein sequence pair by permuting one or more amino acids of the sub-set of amino acids identified in the plant protein sequence and / or by permuting one or more amino acids of the sub-set of amino acids identified in the pathogen protein sequence of the at least one output protein sequence pair; and21 KWS.224.01WO / KWS0460 vii. Automatically annotating the one or more protein sequence pair variants by applying the trained machine learning model on said protein sequence pair variants, the annotation comprising an indication of the protein sequence pair variant will form, according to the prediction of the trained machine learning model, a PPI or not; viii. Supplementing the training data with the annotated protein sequence pair variants; and ix. retraining the trained model on the supplemented training data set, thereby providing a new (more accurate) version of the trained machine learning model .

[0067] According to other embodiments, the model interpretation software may be used for removing wrongly annotated protein sequence pairs from the training data set to provide a cleared training data set, and executing the training or re-training of the machine learning model on the cleared training data set. For example, the steps i-v mentioned above may be executed for identifying a sub-set of amino acid sequences which were the reason for the machine learning model to predict that the at least one output pair will form a PPI. Then, the identified sub-sequence may be compared to a set of sequence patterns or motives known not to take part in a PPI.

[0068] For example, some sequence patterns may be known to be inside of the folded protein, or may be known to be removed or modified during a posttranslational processing step such that the amino acids comprised in these patterns are known not to form or enhance a PPI. If the sub-set of amino acids identified by the model interpretation software matches one of these patterns, this is an indication that the training data comprises wrongly annotated sequence pairs. Hence, the model interpretation software and the sub-set of amino acids identified by the model interpretation software can be used for automatically or semi-automatically identifying and removing wrongly annotated sequence pairs from the training data set. Then, the machine learning model may be retrained on the cleared training data set being free of the wrongly annotated pairs, thereby improving the accuracy of the trained machine learning model.22 KWS.224.01WO / KWS0460

[0069] One or more of the above-mentioned approaches may allow to increase the accuracy of the machine-learning model iteratively with the help of the model interpretation software.

[0070] According to some examples, the application of the model interpretation software is integrated into an automated pipeline for continuous monitoring and updating of the machine-learning model as new protein interaction data becomes available.

[0071] According to some examples, the method further comprises: creating a plurality of plant protein sequence variants of the plant protein of the at least one output protein pair by computationally permuting the sub-set of amino-acid residues identified by the model interpretation software in said plant protein sequence; and repeating steps a) to c) with the plant protein sequence variants (instead of the initially used plant protein sequences) for identifying at least one pair of one of the plant protein variants and one of pathogen proteins which is predicted by the trained machine-learning model to interact.

[0072] Permuting the sub-set of amino-acid residues may mean to delete one or more of these amino acids, replace one or more of these amino acids by other amino acids, change the order of the amino acids, and / or to add additional amino acids in-between the sub-set of amino acid residues. The permutation may be performed randomly or based on an assessment of the biochemical, functional and / or structural properties of the amino acids of the sub-set. The purpose of the permutation may be to create an expanded set of protein pairs which may comprise one or more new protein pairs having an even stronger, desirable PPI, or having an even weaker, undesired PPI, than the protein pair which was used as a basis for the permutation. Another purpose of the permutation may be to feed the generated protein sequence pair variants into the trained machine learning model and to use the output for automatically supplementing or clearing an existing training data set, and for re-training the machine learning model on the supplemented or cleared training data set.

[0073] Preferably, at least the amino acids of the plant protein of at least one plant- protein-pathogen-protein pair predicted by the machine learning model to show a PPI is permutated, as it is typically the plant, not the pathogen, whose genome can be23 KWS.224.01WO / KWS0460 manipulated. However, in some examples, the amino acid subset of the pathogen protein may be permutated instead of or in addition to the amino acid subset of the plant protein. Hence, according to some embodiments, the method may comprise the following step in addition to or instead of the above-mentioned permutation of the amino acids of the plant protein: creating a plurality of pathogen protein sequence variants of the pathogen protein of the at least one output protein pair by computationally permuting the sub-set of amino-acid residues identified by the model interpretation software in said pathogen protein sequence; and repeating steps a) to c) with the plant protein sequence variants and the pathogen protein sequence variants for receiving an indication of at least one output protein variant pair of the plant proteins variants and pathogen protein variants predicted by the trained machine-learning model to form a stronger (desired) or weaker (undesired) PPI than the PPI predicted for the at least one output protein pair.

[0074] An important benefit of the variant creation based on the output of the model interpretation software is that also amino acids which are considered by the model to be highly relevant for PPI interaction - even though these amino acids may not be part of the interaction site - are taken into account and may be optimized by permuting also these amino acids. In structure-based PPI optimization approaches, which solely focus on the amino acids known to be part of the interaction site, those "remote" amino acids are often neglected, and options for PPI optimization are missed. This is avoided by the use of a model interpretation software for identifying amino acids relevant for PPI formation.

[0075] The computationally optimized protein variant pair, i.e., the protein variant pair identified to form a stronger or weaker PPI than the protein pair used for computing the protein sequence pair variants, is then output, e.g., to a user via a GUI or to an application program.

[0076] The computationally optimized protein sequence variant pair can be output, for example, for use in one or more of:• generating a plant whose genome encodes the plant protein sequence of the output plant-pathogen-protein variant pair using genetic engineering or• selectively breeding a plant comprising the plant protein of the at least one output protein variant pair;24 KWS.224.01WO / KWS0460• selectively breeding a plant lacking the plant protein of the at least one output protein variant pair;• identifying one or more individuals carrying a marker for use in marker-assisted breeding;• genetically engineering a plant for generating a plant whose proteome comprises or lacks plant protein of the at least one output protein variant pair;• using chemical mutagenesis methods for generating a plant whose proteome comprises or lacks the plant protein of the at least one output protein variant pair;• empirically validating the PPI formed by the at least one output protein variant pair;• conducting a functional analysis of the plant protein, e.g. by expression of the plant protein or a predicted variant thereof in relevant plant tissues (e.g. leaf) via transient gene transfer;• screening plant populations for plants whose proteome comprises or lacks the plant protein of the at least one output protein variant pair.• predicting which varieties in a population might be most susceptible for a (new) pathogen;

[0077] According to some examples, the method further comprises:• receiving an initial set of protein sequences of the plant;• transforming the initial set of plant protein sequences into a set of plant protein sequences having the same length; and• using the set of plant protein sequences having the same length as the one or more plant protein sequences to be transformed into the plant protein feature vectors;• receiving an initial set of protein sequences of the pathogen;• optionally: analyzing the initial set of pathogen protein sequences for filtering out non-secreted proteins, thereby providing a sub-set of pathogen protein sequences; and25 KWS.224.01WO / KWS0460 using the set of pathogen protein sequences having the same length as the one or more pathogen protein sequences to be transformed into the pathogen protein feature vectors.

[0078] These steps may be advantageous because many machine learning methods, in particular neural network architectures, require an input vector of a predefined length. The pre-processing steps mentioned above may ensure that all feature vectors entered into the model are of the same length, as the length of the feature vectors are determined by the length of the protein sequences. At a minimum, all plant and pathogen protein sequences should be of the same length, respectively. Preferably, both plant vectors and pathogen protein sequences should be of the same length. The vectors of the same length produced from these same-length protein sequences are then used as the input of the machine-learning model.

[0079] There are several ways of making the vectors the same length. For example, the protein sequences can be brought to the same length before they are transformed into vectors. For example, a maximum length can be defined, e.g., such that it is greater than 80% of the lengths of the protein sequences in the initial set, or as large as the protein sequence in the initial set with the maximum length. All vectors shorter than this predefined length are filled (padded) with predefined values up to the predefined length. However, it is also possible to define a much shorter maximum length and to truncate or filter out all protein sequences that are longer. Combinations of these methods (filtering, truncation and padding) are also possible. Therefore, depending on the type of preprocessing procedure used, the number of protein sequences in the initial set may be equal to or greater than the number of protein sequences ultimately transformed into the feature vectors.

[0080] According to some examples, e.g., after having executed the above-mentioned steps, all pathogen and protein sequences that are used for being transformed into feature vectors have a uniform length which may correspond to the longest one of all examined plant proteins and pathogen proteins (possibly after an optional filter step to remove protein sequence artifacts and / or non-secreted pathogen proteins).

[0081] According to some examples, e.g., after having executed the above-mentioned steps, all pathogen and protein sequences that are used for being transformed into26 KWS.224.01WO / KWS0460 feature vectors have the same length. This length may be, for example, 100-10.000 amino acids, e.g., 1.000-4.000 amino acids, e.g., 3.000 amino acids.

[0082] For each amino acid, more than 500, e.g., 1024-4096, or even more than 5000 features, e.g., 8192 features, may be extracted. For example, 2048 features may be extracted from each amino acid of each of the amino acid sequences. These initially extracted features are used for generating the feature vectors which are fed into the trained machine learning model. Various layers of the trained machine learning model may be configured to perform further feature extraction steps and to take also the further features into account when predicting whether two proteins will interact.

[0083] Typically, the initial set of protein sequences used for initially training the machine learning model consists of all known protein sequences of the plant or pathogen under investigation, or a predefined subset of these proteins that are of interest for various reasons. In some examples, optional filtering steps are be performed to reduce the number of protein sequences processed and hence to speed-up the data processing. For example, only pathogen protein sequences comprising a sequence pattern that indicates that it is a secreted protein may transformed into a feature vector and all other pathogen protein sequences may be filtered out, since cytosolic pathogen proteins are usually not directly involved in infection.

[0084] According to some examples, the trained machine-learning model comprises a neural network configured to receive a pair of a protein feature vector and a pathogen protein feature vector as input.

[0085] The use of neural networks configured to process feature vectors may be advantageous because feature vector processing, as described above, may be performed very quickly and with comparatively little computational effort. Neural networks are particularly good at recognizing features and combinations of features that are likely to form protein-protein interactions. This is particularly important when the feature vectors do not contain any structural information, as may be the case in some examples, so that a reliable prediction of whether the two proteins will interact must be calculated from comparatively little information for a pair of proteins.

[0086] In addition, neural networks have the advantage that once the framework for training the network has been established, the trained model can be improved relatively1 KWS.224.01WO / KWS0460 easily. All that is needed is to enrich the training data set with additional data and start the training process again on the extended training data set.

[0087] According to some examples, the neural network comprises a Siamese neural network. The Siamese neural network comprises a first and a second sub-network. The first sub-network is configured to process the plant protein feature vector and the second sub-network is configured to process the pathogen protein feature vector. The Siamese neural network is configured to employ parameter sharing between the subnetworks.

[0088] Using Siamese networks may have the advantage that this network architecture shows excellent learning capabilities and is able to capture the essence of inputs provided in the form of feature representations. It has been observed that the Siamese network architecture is able to process the sequence pairs very fast. The Siamese network is configured to shared weights between the two sub-networks, and this weight sharing allows for the reduction of the number of parameters and for good performance characteristics. Furthermore, the Siamese network has been observed to perform well with limited data by leveraging the ability to recognize similarity between new sequences / feature vectors and previously seen ones.

[0089] According to some examples, the neural network comprises a convolutional neural network (CNN). The CNN will extract features from the protein sequences which are important for interaction. CNN are particularly suited to identify patterns, in particular spatial, local patterns, which are relevant in a given context, e.g., because these feature patterns are predictive for PPI interactions. The CNN comprises convolutional layers, pooling layers, and possibly fully connected layers, and uses them to extract feature maps from the feature vectors obtained initially by transforming the protein sequences into feature vectors. The output of the CNN is a feature map that represents and indicates feature patterns considered to be relevant for PPI prediction. The feature map output by the CNN may be indicative of one or more single amino acids or groups of amino acids identified by the CNN as being highly predictive for the presence or absence of a PPI with respect to the other protein in a protein pair. The feature map is then output by the CNN for further processing by subsequent layers or sub-networks of the overall neural network architecture. For example, the CNN can be28 KWS.224.01WO / KWS0460 combined with a Siamese network to leverage the strengths of both architectures: the CNN is used as the feature extractor of the Siamese network to extract meaningful features from the input data. Each branch of the Siamese network may share the same CNN architecture and weights, ensuring that the feature extraction process is identical for both inputs.

[0090] For example, the feature vectors obtained from the plat and pathogen protein sequence in an initial tokenization and feature extraction step are passed through the same CNN to extract feature maps from each of the two input sequences. The feature maps will contain compressed, high-level representations of the input data. For example, for each (pruned and / or truncated) amino acid sequence having n amino acids whose respective feature vector is processed by the Siamese network, the CNN may be configured to generate and output a feature map having the dimensions m x n, wherein m represents the number of features extracted for an individual amino acid for computing the initial feature vectors. It should be noted that the features in the feature map output by the CNN are typically not identical to the features in the feature vectors derived initially from the protein sequences and having been provided as input to the Siamese network. This is because the initial feature vectors input to the Siamese network will be processed and supplemented or replaced by other features computed by the CNN.

[0091] Once the CNN has extracted a feature map for each of the two protein sequences, the Siamese network compares these feature maps to assess a likelihood of the two sequences forming a PPI.

[0092] After the Siamese network outputs a PPI formation likelihood score or binary classification (PPI-formation: yes / no), this result can be used as input to another neural network for further processing. For instance, the likelihood score of PPI formation might inform a downstream decision-making process, or the extracted features from both branches can be concatenated and passed to a downstream neural network for classification tasks.

[0093] According to one embodiment, the Siamese network is operatively coupled to a Bi-LSTM and / or a Transformer network and is configured to provide its output as input to the Bi-LSTM and / or to the Transformer for further processing. The Bi-LSTM and / or29 KWS.224.01WO / KWS0460Transformer is configured for capturing relationships between single amino acids or groups of amino acids based on the feature maps extracted by the CNNs used by the Siamese network. Preferably, a Transformer is used. The transformer network is able to and is configured to capture relationships between amino acids far apart in the sequence, but which are still important for the formation of a PPI. This could for example be amino acids on either side of a binding pocket, that might be close in 3D space, but far in sequence.

[0094] The Bi-LSTM and / or the Transformer downstream of the Siamese network is configured to compute and output a score value being indicative of the likelihood of the two proteins forming a PPI.The transformer network is able to capture relationships between AA far apart in the sequence, but still important for the final goal of interaction. This could for example be amino acids on either side of a binding pocket, that might be close in 3D space, but far in sequence.

[0095] Using transformer networks may have the advantage that transformers are scalable and can effectively capture relationships between distant elements in a sequence. Hence, transformers are able to recognize the impact of individual amino acids even in cases where these amino acids may not be part of the interaction side but may nevertheless have an impact on the formation of a PPI. Applicant has observed that transformers are particularly suitable for predicting the presence and / or likelihood of a PPI for a pair of plant and pathogen protein sequences.

[0096] In general, the feature vectors initially extracted from the protein sequences that are provided as input to the Siamese network can be extracted, for example, by LLMs trained on proteins sequences, wherein the LLMs may comprise LSTM, CNN, and / or attention layers.

[0097] According to some embodiments, the machine-learning model comprises, downstream of the Siamese network, a transformer or a LSTM network architecture, and optionally attention layers. The network architecture downstream of the Siamese network may comprise a fully connected deep neural network for computing the likelihood of PPI formation for a pair of protein sequences.30 KWS.224.01WO / KWS0460

[0098] According to some examples, the protein sequences are transformed into feature vectors using an encoder part of an encoder-decoder neural network.

[0099] This may have the advantage of integrating feature extraction and feature-based prediction into the same software, the machine-learning model, which may ease the use and deployment of both functions.

[0100] According to some examples, the feature vectors are vectors being indicative of chemical and / or electrical and / or structural properties of the amino acids constituting a protein.

[0101] An advantage compared to predominantly structure-based predictions is increased performance. An advantage compared to other protein-sequence based tools such as clustal omega (Sievers F, et al.: "Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega", Mol Syst Biol. 2011 Oct 11;7:539. doi: 10.1038 / msb.2011.75) may be that these other tools often rely only on phylogeny, but often relationship information is not available. According to some examples, the feature vectors are free of phylogenetic information. In addition, or alternatively, the feature vectors are free of explicit structural information. For example, the feature vectors may only comprise chemical and / or electrical properties of the amino acids, and secondary features derived therefrom. For example, the chemical and electrical information may relate to the number, type and position of atoms of an amino acid and atom-specific polarity information. Nevertheless, even with this limited amount of data, the feature-vector based plant-pathogen-PPI prediction was observed to provide highly accurate results in very short time.

[0102] According to some examples, the transformation of one of the plant or pathogen protein sequences into a respective feature vector comprises: representing the protein sequence as a sentence consisting of a set of different characters representing proteinforming amino acids, the sentence comprising n tokens (also referred to as "words"), wherein each token is a k-mer having a fixed size of k characters, the k-mers being derived by sliding a window having a length of k characters along the protein sequence; and computing the feature vector as a function of the sentence with the n tokens. The feature vectors are sequences of feature sets. This means that for each token, a plurality of features (i.e., a feature set) is extracted. For example, a protein sequence comprising31 KWS.224.01WO / KWS0460100 amino acids may be represented as a feature vector comprising 100 vector elements (100 feature sets). These features can comprise, for example, information regarding protein properties such as chemical and / or physical properties, and structural properties. The feature values are in some embodiments extracted such that amino acids closer in 3D space are closer in vector space, chemically or physically similar amino acids are closer in vector space than chemically or physically dissimilar amino acids, and amino acids in similar k-mers in the protein are closer together in vector space than amino acids in dissimilar k-mers.

[0103] For example, k may be a number between 4 and 9, in particular between 5 and 8, in particular 6. The number n is determined implicitly by the length I of the protein sequence and k according to n=l / k if I is a multiple of k. Otherwise, n=L / k + 1.

[0104] According to another example, the transformation of a respective one of the plant or pathogen protein sequences into a feature vector comprises representing each amino acids of the protein sequence as token and then representing the protein sequence as a series of token,; and computing the feature vector as a function of the sequence of tokens, wherein each of the tokens in the token sequence is transformed into a feature set comprising features of the amino acid represented by said token. The feature vector is a sequence of the feature sets. The feature sets derived from a single token (single amino acid) may also be referred to as token embeddings.

[0105] This "other example" may be implemented as the token-based approach described above, whereby the k-mer is formed for k=l, and hence the size of the sliding window is a single amino acid.

[0106] In some examples, the transformation of protein sequences into feature vectors is not limited to the use of a single, fixed window size. Instead, the feature extraction may be performed using a combination of windows of multiple different sliding window sizes (also referred to as "token sizes" or "word sizes"), ranging from a predefined or user- entered minimum window size to a predefined or user-entered maximum window size. For example, the minimum window size may be one amino acid (k=l), while the maximum window size may be 8 or 9 amino acids, thereby generating feature representations both for individual amino acids and for amino acid motifs or subsequences of greater length. The resulting feature vectors may therefore comprise a32 KWS.224.01WO / KWS0460 combination of feature sets derived from tokens of different sizes. This may allow the machine learning model to capture and evaluate both local dependencies at the singleresidue level and broader sequence motifs that span multiple residues. This multi-scale feature extraction approach may enable the machine-learning model to analyze the predictive power of amino acid patterns of different sizes, thereby improving the accuracy and robustness of protein-protein interaction prediction.

[0107] In some examples, the transformation of a respective one of the plant or pathogen protein sequences into a feature vector comprises: Receiving a minimum window size and a maximum window size, wherein in particular the minimum window size is 1 and the maximum window size is a value smaller 12, in particular smaller 10. For each window size of all window sizes from the minimum to the maximum window size: perform the representation of the protein sequence and the computation of the feature vector in accordance with the examples described above, wherein the sliding window length is equal to the window size; and use the feature vectors computed for the sliding windows of the different sizes ranging from the minimum window size to the maximum window size for performing the predicting by the machine-learning model.

[0108] Importantly, the use of multiple window sizes does not prevent the computation of residue-level relevance scores. Model interpretation software such as SHAP, layerwise relevance propagation, or gradient-based attribution methods may be applied to the trained model to determine how each input feature— whether derived from a single amino acid or from a larger token— contributes to the final prediction. Where features are extracted from windows larger than one amino acid, the model interpretation software assigns a relevance score to the corresponding token, which may then be decomposed or distributed across the amino acids that constitute the token. Since each amino acid participates in one or more overlapping tokens, the combined relevance values derived from all tokens in which a residue appears can be aggregated to yield a residue-specific relevance score. In this way, even when the machine learning model processes feature sets derived from multiple window sizes, it remains possible to compute and visualize the contribution of each individual amino acid to the predicted likelihood of protein-protein interaction.33 KWS.224.01WO / KWS0460

[0109] In embodiments of the invention, the resolution of the extracted feature information may correspond to the chosen window size, so by choosing a combination of different window sizes, a combination of patterns of different sizes may be examined and analyzed in the training and injunction phase. For example, with a window size of k = 7, the extracted feature vectors each represent the local context of seven consecutive amino acids, such that the subsequent machine learning model can directly assign contributions at this level of granularity. Accordingly, explanations of the model output can initially be obtained only at the resolution of the selected window size. However, by choosing an appropriate network architecture, e.g. a prediction architecture comprising convolutional neural networks (CNNs), the contributions of individual amino acids can be resolved at a finer scale. By applying explainability methods such as perturbation analysis or layer-wise relevance propagation (LRP) within the network structure of the machine learning model used to predict the PPIs, the relevance attributed to a window can be redistributed to the underlying amino acids forming that window. In this way, the contribution of individual residues to the model output can be determined, even if the original feature extraction step was performed with windows larger than one amino acid.

[0110] In addition, or alternatively, a combination of windows of different window sizes, including windows of size 1, may be used.

[0111] Thus, although feature extraction with a window size k yields an initial resolution of k, the combination of CNN-based architectures with attribution methods allows a decomposition of the model prediction down to the level of individual residues (k = 1) even in case all evaluated window sizes should have a window size larger than 1. This may enable the identification of amino acids of particular importance for the predicted protein-protein interaction.

[0112] The approach of combining feature extraction across multiple window sizes may improve the representational richness of the feature vectors without compromising the ability of the model interpretation software to identify a subset of specific amino acids most relevant for a prediction. On the contrary, by considering amino acid features at different granularities, the system may be able to identify not only single residues but also motifs and patterns of residues that play an important role in protein-protein34 KWS.224.01WO / KWS0460 interactions, thereby enabling a more complete and biologically meaningful interpretation of the trained model's predictions.

[0113] The sliding window method may be used by different network architectures, which are, according to some embodiments, part of the trained machine learning model, for modifying and / or enriching the initially created feature vectors with further features extracted when applying the sliding window approach: for example, the sliding window approach may be an inherent part of the convolution operation of a CNNs, wherein a convolution filter (kernel) acts as a sliding window that moves across the initial feature vector obtained for the protein. Thereby, the CNN kernel may detect local features (e.g., the presence of amino acid patterns being predictive for PPI formation) by applying the filter to each region of the initial feature vector. This allows CNNs to "catch" local patterns regardless of their position in the input. Instead of or in addition to a CNN, a RNN or LSTM may employ a sliding window approach for extracting additional features from the initial feature vectors. These enriched and / or modified feature vectors may be provided as input to a Siamese network as described above.

[0114] Treating each individual amino acid as a unit ("token' or 'word') within a sequence of amino acids (a sentence) may have the advantage of allowing the machinelearning model during the training phase to flexibly decide on and learn the granularity of the feature patterns which are predictive for the existence of a PPI, without having a human to define a certain token-size (word-size) arbitrarily.

[0115] Applicant has observed that using a token size of a single amino acid allows the machine learning model to decide if smaller or larger token / amino acid sub-sequences ("chunks") are indicative of feature patterns which are predictive the formation or nonformation of a PPI, thereby allowing to provide a trained machine-learning model able to accurately predict PPI formation. Using a token-size of a single amino acid may ensure that all amino acids (up to the size limit of the protein sequence) and their representations within inner network layers of the machine-learning model are connected to the initial network having received the respective feature vector of this protein sequence as input. The decision as far the size of predictive amino acid feature patterns is concerned, is let to the model to decide according to the available data. Using tokens at the scale of a single amino acid increases computational demands but captures35 KWS.224.01WO / KWS0460 more detailed information. The CNN or other type of neural network used for applying a sliding window approach for extracting further features will typically require the number of input neurons corresponding to the number of tokens extracted from each protein sequence. When using single amino acid tokens, the number of input neurons matches the padded length of the sequence.According to one embodiment, an existing tool is used for transforming an initial protein sequence into a feature vector which can then be either fed into a further network for further feature extraction or can be fed directly into a trained machine learning model for performing the PPI prediction. For example, pre-trained, Transformer-architecture based, self-supervised methods can be used for transforming the protein sequences into feature vectors. Examples for these methods and tools are ProteinBert or ProstT5 (ProstT5: Bilingual Language Model for Protein Sequence and Structure Michael Heinzinger, et al., bioRxiv 2023.07.23.550085; doi: https: / / doi.org / 10.1101 / 2023.07.23.550085) which are pre-trained on millions of proteins, capable of extracting feature-representations of the amino acids of a protein that includes latent information on protein function and structure. These network structures and existing tools can be used for identifying amino acids that are closer together in 3D space but not in sequence and can compute vector representations capturing the spatial relationship of these amino acids such that the amino acid distance in 3D space is transformed into a respective distance in the vector space.

[0116] The machine learning model for performing the PPI prediction can in particular comprise a Siamese network.

[0117] According to some examples, the one or more plant protein sequences are multiple protein sequences, e.g. at least 500, e.g., at least 1.000 or at least 5.000 or at least 10.000 protein sequences. In addition, or alternatively, the one or more pathogen protein sequences are multiple protein sequences, e.g. at least 500, e.g., at least 1.000 or at least 5.000 or at least 10.000 protein sequences.

[0118] Hence, due to the high performance of the feature-vector based, prediction, the method may be used for computing an a I l-against-a II PPI search of the whole proteome of a plant and the whole proteome of a pathogen, thereby identifying also completely unknown PPIs and completely new options for developing pest-resistant plants, even in36 KWS.224.01WO / KWS0460 cases where the plant and pathogen proteins involved in an infection are completely unknown.

[0119] Also described herein is a method for providing the machine learning model by training the model on training data. The training comprises:• providing pairs of plant proteins and pathogen proteins being examples of interacting plant-protein-pathogen-protein pairs; and providing pairs of plant proteins and pathogen proteins being examples of non-interacting plant-protein- pathogen-protein pairs; for example, these positive and negative protein pairs, also referred to as "training protein pairs" or "training pairs", including their respective labels, may be derived from the literature or from various empirical tests or a combination thereof. The labels assigned to each pair may indicate if the pair forms a PPI or not. In some examples, the labels may only be indicative of the presence of a PPI. In other examples, the labels may also be indicative of the strength of the PPI and / or an indication of the amino acids constituting the interaction site; for example, the training pairs can be formed by automatically generating all combinatorially possible combinations of one of the plant protein sequences and one of the pathogen protein sequences of the training data set;• transforming the provided pairs of plant and pathogen protein sequences into training pairs of plant protein feature vectors and pathogen protein feature vectors; preferably, this is performed by transforming the plant and pathogen protein sequences into protein sequences having the same length, e.g., by filtering, truncating and / or padding, transforming each of these pre-processed protein sequences having the same length into a respective feature vector (referred also as plant or pathogen training feature vector), and then combining these precomputed training feature vectors into training feature vector pairs which correspond to the provided plant and pathogen protein pairs; in other words, the transformation of the protein pairs into sequence vector pairs may reuse precomputed feature vectors of each of the proteins in order to avoid repeating the same transformation operation multiple times for each protein;37 KWS.224.01WO / KWS0460• using the training protein feature vector pairs and an indication if the training pair is an example of an interacting or non-interacting protein pair for training the machine-learning model and for obtaining the trained machine-learning model.

[0120] For example, the machine-learning model may comprise one or more neural networks, in particular a Siamese network, which uses back-propagation to learn, based on the features of the amino acids comprised in each protein pair, which ones of these amino acids or patterns of amino acids having similar features correlate with a label indicating the existence of a PPI, and which features and feature patterns correlate with a label indicating the absence of a PPI.

[0121] Then, the trained machine-learning model may be used to predict the existence and / or likelihood of existence of a PPI between pairs of proteins which have not been part of the training data set. These new proteins and protein pairs used after the training phase (i.e., the "test phase" or "application phase") may also be referred to as "test protein sequences" or "new protein sequences". At the test phase, the trained machine learning model may be applied to new protein sequence pairs and may compute an output, e.g., a score, which is indicative if the one or more test protein pairs provided as input will form a PPI.

[0122] The output may comprise an indication of one or more of the test protein sequence pairs and respective test protein pairs which are predicted to form a PPI. These one or more output protein pairs may be used for many different purposes, for example:• selectively breeding a plant comprising the plant protein of the at least one output protein pair predicted to interact or not to interact with the pathogen protein of the at least one output protein pair; For example, a large number of plant protein sequences and pathogen sequences may have been input into the trained model, and the model may evaluate all combinatorially possible combinations of plant and pathogen protein sequences as described above and may return the one of these evaluated protein pairs having the highest predicted likelihood of forming a PPI. The output may also comprise a set of the z plant and pathogen protein pairs having the highest predicted likelihood of forming a PPI, whereby z may be any integer greater than 1, typically a number smaller 20; Then, a population of plants may be genetically screened for variants having one38 KWS.224.01WO / KWS0460 of the plant protein sequence variants comprised in the z output protein pairs; The plants identified in this screening step may be plants which have a protein variant which will highly likely form a strong PPI of a desired PPI type in the context of an infection;• selectively breeding a plant not comprising the plant protein of the at least one output protein pair predicted to interact or not to interact with the pathogen protein of the at least one output protein pair; in application scenarios where it is desired to identify plants which do not form an undesired PPI, one option may be not to breed any plant comprising a protein variant predicted by the model to have a high likelihood of forming a PPI (see example above); another option according to one example is that the machine learning model is configured and operated such that it returns a single or a subset of protein pairs which have the lowest predicted likelihood of forming a PPI; selectively these plants are then used in the breeding project;• genetically engineering a plant for generating a plant whose proteome comprises or does not comprise the plant protein comprised in the at least one output protein pair; for example, if the removal of a plant protein which enables or facilitates infection via a predicted PPI is not essential for the plant, one option may be to selectively breed existing or genetically modified plants whose genome does not comprise a gene encoding this plant protein;• using chemical mutagenesis methods (e.g., Tilling) for generating a plant whose proteome comprises or does not comprise the plant protein comprised in the at least one output protein pair;• identifying candidate genes or candidate mutations adapted to increase pathogen resistance via the presence or absence of the plant protein comprised in the at least one output protein pair, and empirically testing the candidate genes or mutations or plants comprising the candidate genes or mutations; for example, a model interpretation program may be used for identifying a sub-set of amino acids having caused the machine learning model to predict the presence of a PPI; then, protein pair variants may be created computationally by permuting amino acids of this sub-set in the plant protein; then, plant variants respectively39 KWS.224.01WO / KWS0460 expressing a different one of these plant protein variants may then be created using genetic engineering techniques, and these plant variants may then be tested empirically for increased pathogen resistance;• conducting a functional analysis of the plant protein, e.g. by expression of the plant protein or a predicted variant thereof in relevant plant tissues (e.g. leaf) via transient gene transfer;• screening plant populations for plants whose genome results in the presence or absence of the plant protein comprised in the at least one output protein pair;• Predicting which varieties in a population might be most susceptible for a (new) pathogen.

[0123] In some examples, the plant protein of the at least one output protein pair is an R-gene conveying resistance of the plant against the pathogen.

[0124] In other examples, the plant protein of the at least one output protein pair is an S- gene conveying susceptibility of the plant against the pathogen.

[0125] When performing a screening or genetic engineering step, it is typically desired to selectively breed plant individuals expressing an R-gene, in particular R-genes having a strong PPI, and it is typically desired to avoid breeding plant individuals expressing an S- gene, in particular S-genes having a strong PPI.

[0126] In a further aspect, described herein is a computer system configured for: a) transforming one or more protein sequences of a plant into a respective plant protein feature vector and transforming one or more protein sequences of a pathogen into a respective pathogen protein feature vector; and b) forming a plurality of different feature vector pairs respectively comprising one of the pathogen protein feature vectors and one of the plant protein feature vectors; c) providing the vector pairs to a trained machine learning model, and in response receiving, for the pairs of feature vectors, an indication if at least one output protein pair represented by one of the feature vector pairs is predicted by the trained machine-learning model to interact.

[0127] For example, the steps of data transformation and of providing the feature vectors to the trained model may be executed by a data transformation software. Plant and pathogen protein sequences may be pre-processed to bring them into a uniform length, and existing tool such as ProteinBert or ProstT5 may be used for transforming the40 KWS.224.01WO / KWS0460 pre-processed protein sequences into feature vectors which may be fed into a neural network architecture which may comprise a series of subnetworks or layer, e.g., CNNs for extracting further, local features and feature patterns, a Siamese network and downstream networks (e.g., a transformer) for computing a PPI formation likelihood score.

[0128] According to some examples, the data transformation software may be an integral part of an integrated pipeline of multiple neural networks, e.g. a network for feature extraction, and a combination of a Siamese network and a transformer network constituting the network actually performing the PPI prediction. In some examples, the data transformation software may comprise software functions for transforming protein sequences into feature vectors and for computing protein sequence pairs and / or respective feature vector pairs.

[0129] Optionally, the data transformation software may comprise a function for computing all combinatorially possible protein pairs, for transforming a plurality of protein sequences into protein sequences having the same length, for training the machine learning model and / or for executing a model interpretation software.

[0130] Optionally, the data transformation software may also comprise software for computing protein variants by permuting individual amino acids and / or for screening plant and pathogen protein databases.

[0131] The data transformation software may be a single integrated software program or a collection of programs stored on the same or on a distributed computer system or a network of computer systems. In some examples, the data transformation software may comprise or be operatively coupled to the machine learning model.

[0132] In some embodiments, the computer system may comprise an interface. For example, the plant and pathogen protein sequences may be provided via the interface to the data transformation software and / or the prediction results computed by the trained model (and optionally by a model interpretation software) may be returned via the interface to a computer (e.g. a computer referred to as client computer) or to a software (e.g. referred to as client software) having provided the protein sequences via the interface. The interface can be, for example, a network interface in case the data transformation software and the machine learning model are stored on different41 KWS.224.01WO / KWS0460 computer systems connected to each other via a network. Alternatively, the interface may be an internal interface of a software program, e.g., a neural network software program, comprising a functionality for transforming protein sequences into feature vectors and / or for computing protein or feature vector pairs.

[0133] According to some examples, the computer system further comprises a graphical user interface (GUI) configured to display the output generated by the machine-learning model and / or generated by the model interpretation software on a display device, e.g. an LCD display. For example, the GUI may comprise a graphical representation of the predictions of the machine-learning model for one or more plant-protein-pathogen- protein pairs. These representations may include bar charts, summary plots, histograms, scores, line charts, and dependence plots. In addition, or alternatively, the GUI may comprise a graphical representation of the predictions of the model interpretation software for one or more plant-protein-pathogen-protein pairs or for a single protein. These representations may include a color-coded or greylevel-coded protein sequence, where the amino acids identified as being the most relevant for the predicted PPI are highlighted.

[0134] According to some examples, the computer system also comprises the trained machine learning model and / or a model interpretation software.

[0135] The expression 'permutation of amino acids' as used herein refers to the process of changing the nature and / or order of amino acids in a protein or peptide sequence or sub-sequence. This involves changing the order of the amino acids while keeping the same set of amino acids, and / or replacing one or more amino acids by other amino acid(s) which may or may not be comprised in the permutated sequence. In some cases, also the number of amino acids may be changed, e.g., by removing or adding one or more amino acids. Permutations, also referred to as "sequence variations", potentially result in different structural and functional properties of the protein.

[0136] A 'pathogen' or 'plant pathogen' as used herein may be defined as any organism, such as a bacterium, virus, or fungus, that could cause disease in a plant. For example, the bacterium Xanthomonas campestris could infect various plants and might disrupt their normal physiological processes. The fungus Puccinia graminis might infect cereal42 KWS.224.01WO / KWS0460 crops and could lead to significant yield losses. Cercospora beticola is a fungal pathogen responsible for the Cercospora leaf spot (CLS) disease of beet crops.

[0137] A 'protein sequence' as used herein refers to the order of amino acids in a protein molecule. The sequence may be provided in the form of a sequence of amino acid names or symbols or may be provided in the form of a nucleotide sequence which implicitly encodes the amino acid sequence. This sequence may determine the protein's structure and function.

[0138] A 'feature vector' as used herein is a representation, in particular a numerical representation, of certain characteristics or features extracted from a protein sequence. These vectors could be used as input for machine learning models to predict interactions. For example, a feature vector for a protein might include values representing chemical, electrical, structural or other properties of individual amino acids, and may optionally also comprise more coarse-grained or global properties of the amino acid sequence such as specific amino acid patterns or structural motifs. A feature vector can be a multidimensional vector, in particular a matrix. According to some examples, the dimensions of the matrix may be defined by the length of the amino acid sequence and the numbers of features extracted per amino acid. The term "feature vector" in the context of this application is to be interpreted broadly as a numerical, one-dimensional or multidimensional representation of certain characteristics or features extracted from a protein sequence. It is not limited to a strictly linear data structure such as an array.

[0139] A 'computing device' as used herein is an electronic device that processes data according to specified instructions. It might include servers, desktop computers, or specialized hardware used to run machine learning algorithms for predicting proteinprotein interactions. The computing device can be a monolithic or distributed computer system, e.g., a distributed cloud computer system. According to another example, the computing device may be a high-performance computing device, or a distributed computer system comprising at least one server computer and one or more client computers.

[0140] A 'token' as used herein refers to a single unit of a protein sequence, such as a sub-string of amino acids or a single amino acid, that is processes to extract vector representations (embeddings) for the tokens. In the context of text processing, the43 KWS.224.01WO / KWS0460 embeddings capture the semantic meaning of words based on their context within the text. In the context of protein sequences, the embeddings capture the chemical, physical, electrical and / or structural properties of amino acids or amino acid sub-strings based on their context within a protein sequence. According to embodiments, the number of input neurons of a neural network used for processing the protein pairs for extracting features and for predicting if the pair will form a PPI is equal to the number of tokens comprised in each one of the input protein sequences (which have the same length, e.g. as a result of a filtering, padding and / or pruning process). In the case of single amino acid sized tokens, the number of input neurons is equal to the number of amino acids in each of the input protein sequences.

[0141] A 'machine learning model' as used herein may be a mathematical representation or algorithm trained on a set of data to recognize patterns and make predictions or decisions based on new, unseen data. These models are built through various machine learning techniques, such as supervised learning, unsupervised learning, or reinforcement learning. In particular, the machine learning model may comprise a neural network for PPI prediction which may be provided via supervised learning. In some examples, the machine learning model may be a complex software comprising additional functions, e.g., for feature extraction, feature vector computation, vector-based prediction, model interpretation, and / or other functions, e.g., functions for graphically representing the prediction results on a GUI.

[0142] A 'model interpretation software' as used herein may, for example, refer to software designed to provide insights into how machine learning models make their predictions. For example, a model interpretation software may be configured to identify a sub-set of amino acids of an amino acid sequence which are of particular relevance for the prediction of a protein-protein interaction computed by a machine-learning model, e.g., a Siamese network. The subs-set of amino acids may have been considered relevant by this model, e.g., because they the amino acids comprised in this sub-set may constitute the interaction site or another site that has an important impact on the structure or biochemical properties of a protein's interaction site (also referred to as "binding site"). A model interpretation software may allow to understand and validate the prediction made by a model with respect to the existence and / or likelihood of a PPI,44 KWS.224.01WO / KWS0460 and this information may be used for increasing or decreasing the strength of the predicted interaction, e.g. via various genetical engineering methods for modifying these particularly relevant amino acids. A model interpretation software may be configured to identify the amino acids which were the reason for the machine-learning model to output the indication, thereby making "black box" nature of machine learning models more transparent and interpretable, and enabling scientists to modify and optimize protein sequences. For example, a model interpretation software may be implemented using a recurrent neural network architecture or with other approaches. The model interpretation software may be software separate from the machine learning model configured for predicting the PPI, or it may be an integral part of it (e.g., in case the model interpretation software comprises attention layers and the attention layers are used as model interpretation software.

[0143] The expressions "stronger / enhanced" or "weaker / decreased" PPI or binding strength as used herein should be understood broadly in the sense that a stronger / enhanced PPI may have an increased interaction strength, an increased interaction duration and / or an increased likelihood of the two proteins interacting with each other. Likewise, a weaker / decreased PPI or binding strength may have a decreased interaction strength, a decreased interaction duration and / or a decreased likelihood of the two proteins interacting with each other. A plant pathogen pair 'lacking" a particular PPI may be a plant pathogen pair where the plant protein and / or the pathogen protein exists in the form of an allele which prevents the formation of a PPI between these two proteins completely or which weakens the PPI to such a degree that the predicted likelihood of PPI formation is below a predefined threshold, e.g., below 1%.

[0144] It is understood that one or more of the aforementioned examples may be combined as long as the combined examples are not mutually exclusive.BRIEF DESCRIPTION OF THE DRAWINGS

[0145] In the following, examples are described in greater detail making reference to the drawings in which:

[0146] Fig. 1 is a flow chart of a method for predicting plant-pathogen-PPIs,45 KWS.224.01WO / KWS0460

[0147] Fig. 2 is a block diagram of a computer system configured to predict plant- pathogen-PPIs,

[0148] Figs. 3A, B are illustrations of data objects and software modules involved in the prediction of plant-pathogen-PPIs,

[0149] Fig. 4 comprises plots illustrating the application of the plant-pathogen-PPI prediction,

[0150] Figs. 5A, B are illustrations of a PPI prediction pipeline; and

[0151] Fig. 6 illustrates an example of a software architecture of the data processing software and the machine learning model.DETAILED DESCRIPTION

[0152] In the following, similar elements are denoted by the same reference numerals.

[0153] As mentioned above, the interactions between pathogen proteins and host proteins have a crucial impact on the development of disease symptoms, the identification of those interacting proteins and furthermore the identification of the interacting residues can be very helpful for understanding the infection process and may allow selection of the best lines for the development of new varieties with enhanced tolerance up to complete resistance to pathogens. This may allow minimizing the use of pesticides.

[0154] Various improved approaches for protecting plants against pathogens are described herein relying on the use of a machine learning model configured to predict whether a pair of a plant and a pathogen protein will interact based on an evaluation of feature vectors derived from the protein sequences of the said pair.

[0155] Figure 1 shows a flow chart of a method for predicting plant-pathogen-PPIs.

[0156] PPIs in general are important for many biological processes, both within an organism and across organisms. Plants are attacked by a variety of pathogens. Often this attack is carried out or mediated by proteins from the pathogen which interact with proteins in the plant. To reduce or eliminate the use of pesticides, plants are often bred for resistance to pathogens. However, this is often very challenging and sometimes not46 KWS.224.01WO / KWS0460 feasible: often the interaction partners are not known, or only one of the interacting partners is known. Existing PPI databases comprise only a limited number of PPIs and often do not cover the plant or pathogen of interest. The software used for creating these databases is often not publicly available, inaccurate or too slow, in particular in situations when only one or even none of the interaction partners is known and a large number of interaction candidates have to be evaluated.

[0157] The feature-vector based approaches described herein allow to screen for possible interaction partners, even if none or only one is known. Once the interaction partners responsible for a plant-pathogen interaction mediating the infection have been identified, the model interpretation software described herein may be used for identifying the interaction sites and for computing and screening a set of plant protein variants (corresponding to mutations in the respective gene) which may have an even stronger desired PPI, or an even weaker undesired PPI. The identified interaction sites and the respective amino acid motifs may also be used as markers for marker-based selective breeding.

[0158] The feature-based plant-pathogen-PPI prediction may be executed on a computer system 200 as illustrated, for example, in figures 2, 3 and 5 in greater detail.

[0159] Initially, one or more plant protein sequences and one or more pathogen protein sequences of a plant-pathogen pair of interest may be obtained, e.g. from the literature or protein databases or experimentally or by a combination thereof. Depending on the case, these protein sequences may cover the whole proteome of the plant and / or of the pathogen, or a sub-set thereof, e.g. may only relate to secretory proteins of the pathogen or to plant and pathogen protein sequences for which a role in the infection process is suspected or at least not ruled out. If one interaction partner is already known, it is possible that only the protein sequence of this reaction partner and a plurality of protein sequences of the other organism are provided.

[0160] In step 102, the one or more protein sequences of the plant are transformed into a respective plant protein feature vector and in step 104 the one or more protein sequences of the pathogen are transformed into a respective pathogen protein feature vector. Preferably, the transformation is performed such that the resulting plant and pathogen feature vectors all have the same size, e.g., via length-based sequence filtering,47 KWS.224.01WO / KWS0460 padding and / or truncation. The feature extraction may be performed by a separate feature extraction software.

[0161] Preferably, transforming protein sequences into feature vectors is performed by tokenizing the protein sequence into tokens and computing token-based embeddings. For example, if the tokens respectively consist of a single amino acid, the embeddings are amino-acid-based feature vectors. For example, the feature extraction, also referred to as embedding computation, can be performed via a transformer neural network architecture.

[0162] According to one example, the transformer neural network architecture can be a self-supervised deep language model designed and optimized for protein sequences.

[0163] For example, the transformer-based programs ProteinBERT or ESM2 may be employed.

[0164] The ProteinBERT approach has been described in Brandes, Nadav, et al."ProteinBERT: a universal deep-learning model of protein sequence and function." Bioinformatics 38.8 (2022): 2102-2110 which is incorporated herein by reference in its entirety.

[0165] The ESM2 approach has been described Rives, Alexander, et al. "Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences." Proceedings of the National Academy of Sciences 118.15 (2021): e2016239118 which is incorporated herein by reference in its entirety.

[0166] According to other embodiments, the embedding computation is performed via the word2vec or doc2vec approach. The doc2vec approach has been described in Q Le, T Mikolov, "Distributed representations of sentences and documents", International conference on machine learning, PMLR (2014) 1188-1196 which is incorporated herein by reference in its entirety. The word2vec approach has been described in Mikolov T, Chen K, Corrado G, Dean J: Efficient estimation of word representations in vector space [Preprint], arXiv (2013) 13013781, https: / / doi.org / 10.48550 / arXiv.1301.3781 which is incorporated herein by reference in its entirety.

[0167] Feature extraction and vector computation may be performed, for example, by so called embedding layers of a neural network. Then, in step 106 a plurality of different feature vector pairs are formed. The pairs respectively comprise one of the pathogen48 KWS.224.01WO / KWS0460 protein feature vectors and one of the plant protein feature vectors. Preferably, this step comprises forming all combinatorially possible plant-pathogen feature vector pairs. In step 108, the feature vector pairs are fed as input into a trained machine learning model 212, which is preferably a Siamese network, wherein one sub-network is configured to receive the plant protein feature vector and the other sub-network is configured to receive the pathogen feature vector of the pair. The model predicts and outputs, for the pairs of feature vectors, an indication if at least one output protein pair represented by one of the feature vector pairs is predicted by the trained machine-learning model to interact. For example, this indication may comprise a numerical score value computed for each feature vector pair provided as input, wherein the score indicates the likelihood that the protein pair corresponding to the feature vector pair will form a PPI. In response to the provision of the feature vector, the indication of whether the plant-pathogen pair will interact is received in step 110. For example, it is possible that a client device, which submits the protein sequence pairs to a server hosting the machine learning model for feature extraction and predicting the indication, whereby the client device is configured to receive the computed indication via a network (e.g., the internet or an intranet) from the server and display it via a GUI or provide it to another application program.

[0168] The results output by the machine learning model may be used for various applications, e.g., for selective breeding, wherein a plant population is screened for plants whose genome comprises a plant protein variant predicted to form a desired PPI with high likelihood or predicted not to form an undesired PPI is selected for a breeding project. This screening approach would allow the identification of plants which are resistant or have a higher tolerance for the pathogen. Preferably, the prediction of the model is empirically validated before a plant is actually selected and used for the breeding project.

[0169] In some examples, the method may comprise additional steps illustrated with reference to figure 5 to predict and optionally also improve, via genomic engineering techniques, the interaction sites of the most interesting plant-protein pair, wherein an improvement may be the enhancement of a desired PPI or a weakening / decrease of an undesired PPI.49 KWS.224.01WO / KWS0460

[0170] Figure 2 is a block diagram of a computer system 200 configured to predict plantpathogen PPIs. The computer system has one or more processors 202 and includes a volatile and / or non-volatile storage medium 206 on which one or more plant protein sequences 208 and one or more pathogen protein sequences 210 are stored. For example, the protein sequences 208, 210 may initially be stored in one or more sequence databases generated or collected internally by the organization performing the prediction of the PPIs. Alternatively, the protein sequences 208, 210 may be initially stored and provided externally by other database providers, such as universities and research groups. These proteins may then be read from the storage medium or retrieved via a network connection and loaded into the main memory.

[0171] The computer system 200 may be a monolithic or distributed computer system whose individual hardware and software components are interconnected, e.g. networked, so that they are operationally coupled and can form a functional unit.

[0172] The computer system comprises a machine learning model (which may also be referred to as the "predictive model") 212 which is generated in the course of a machine learning process, typically by supervised learning. In addition, the computer system comprises software 204 for transforming protein sequences into feature vectors. In some implementations, this software 204 may also be an integral part of the machine learning model 212. For example, the machine learning model may be or comprise a neural network, whereby the neural network comprises so-called embedding layers configured to extract a feature vector from each plant or pathogen protein sequence provided as input. Optionally, the computer system can comprise software 216 for interpreting the operation of the machine learning model 212. This model interpretation software 216 may be configured, for example, to feed various protein sequences or feature vectors derived therefrom as input to the machine learning model 212 and to analyze the predicted probabilities of occurrence of a PPI output by the model in response thereto. This analysis may be performed to identify amino acids that are likely to be causative of the model's prediction that a PPI will occur.

[0173] The interface 214 is used to forward protein sequences 208, 210 to the model 212 during the test phase and optionally also during the training phase, and to forward the prediction result(s) generated by the machine learning model 212 and / or by the model50 KWS.224.01WO / KWS0460 interpretation software 216 to an output interface. The output interface can be, for example, a graphical user interface.

[0174] According to one embodiment, the computer system 200 is a client-server computer system comprising a server computer system and one or more client computer systems. For example, the client computer systems can be smartphones, tablet computers, desktop computers, notebooks or other forms of mobile or stationary computing devices typically assigned to an individual user. Each client computer may host a client software configured to receive protein sequences 208, 210, to submit the protein sequences via the interface 214 over a network, e.g., the internet, to the server computer system. The server computer system may comprise the machine learning model 212, optionally also software for training or re-training the model, and optionally also further software such as the model interpretation software 216, and the data transformation software 204. The server computer system is configured to receive the protein sequences 208, 210 via the internet from one of the one or more client computer systems, use the machine learning model for computing an indication if at least one pair of the protein sequences received from the client forms a PPI, and optionally to use the model interpretation software to identify one or more amino acids residues having the highest impact on the prediction of the model 212 that two proteins will form a PPI. The server computer system is further configured to send the prediction results computed by the machine learning model 212, and optionally also the prediction results computed by the model interpretation software 216 via the network to the one of the client computers having submitted the protein sequences.

[0175] Figures 3A, 3B are illustrations of data objects and software modules involved in the prediction of plant-pathogen PPIs.

[0176] Initially, one or more plant protein sequences 208 of the plant species of interest and one or more pathogen protein sequences of the pathogen of interest may be obtained. For example, it may be possible to retrieve the sequences representing the whole proteome of the plant or the pathogen, or a sub-set of these protein sequences, e.g. only the secreted proteins of the pathogen. The sequences may be derived from the literature, an external protein database, or may be obtained experimentally by sequencing a plant or a pathogen.51 KWS.224.01WO / KWS0460

[0177] Then, the protein sequences are transformed into feature vectors, also referred to as embedding vectors. Processing a feature vector instead of a protein sequence may have the advantage that a feature vector may comprise much more information than comprised in the amino acid sequence alone. For example, the feature vector may also comprise information concerning the size, polarity, atomic composition, side chains, etc.

[0178] For example, this may be performed by embedding layers of a neural network 204 as described in Pengfei Xie et al., 2023 mentioned above. In particular, the algorithm 'doc2vec' may be used for transforming protein sequences into feature vectors, whereby the tokens (words) in this case are k-mers of amino acids or individual amino acids. Doc2vec learns the similarities between words from a large document corpus and uses a similar vector to describe tokens with similar contexts. Hence, the transformation of protein sequences into feature vectors may be performed by a trained neural network 204, which has learned how to transform the tokens ('words') in a protein sequence into a feature vector that can represent the chemical, electrical and / or functional and optionally also structural features of the entire protein in a detailed and accurate manner.

[0179] According to one example, the neural network 204 may in particular have a transformer architecture and may be configured to translate a protein sequence into a feature vector capturing the relationship between the amino acids.

[0180] Preferably, the feature vectors of the plant and pathogen protein sequences are transformed into feature vectors having the same length. This can be done on the level of the feature vectors, e.g., via filtering, padding and / or truncating the feature vectors. Alternatively, the pre-processing step may already be performed on the level of the protein sequences by filtering out, truncating or padding the protein sequences.

[0181] After this step, a set of one or more plant protein feature vectors 306 and a set of one or more pathogen protein feature vectors 308 all having the same length is provided.

[0182] Then, one or more feature vector pairs are selected. For example, the computer system 200 may comprise software configured to loop through all plant protein feature vectors 306 and pathogen feature vectors 308 available to form all combinatorially possible feature vector pairs. Each vector pair, e.g., the pair 313 comprising the plant protein feature vector 310 and the pathogen feature vector 312, are fed as input into a52 KWS.224.01WO / KWS0460 respective sub-network 316, 318 of a Siamese network 314. For each examined protein pair, the plant protein feature vector goes into sub-network 316 and the pathogen protein feature vector goes into sub-network 318. The sub-networks are identical and the weights are shared. Each sub-network can consist of or comprise a convolutional neural network (CNN), an LSTM, a transformer or another network architecture. Preferably, a transformer architecture is used as it is fast and can be parallelized easily.

[0183] The Siamese network 314 identifies and extracts the important features that are particularly predictive of the presence and / or likelihood of occurrence of a PPI.

[0184] The results of the Siamese network are input into a downstream neural network 320. The downstream neural network processes the features of the two proteins provided by the Siamese network. The downstream neural network may be, for example, a CNN or a recurrent neural network (RNN) architecture, in particular a Bidirectional Long Short-Term Memory, with attention. According to preferred examples, the downstream network 320 is a language model, in particular a large language model.

[0185] Finally, the network 320 predicts and outputs an indication 322 of whether the protein pair whose vector pair was provided as input will form a PPI. Preferably, this indication may be provided in the form of a score, e.g. a score between 0 and 1, wherein 0 may indicate that no PPI will be formed, and 1 that a PPI is predicted to form with very high likelihood.

[0186] Figure 4 shows plots illustrating the application of the feature-based plant- pathogen-PPI prediction method for sugar beet plants with respect to the pathogen Cercospora beticola. Cercospora beticola is a fungal pathogen responsible for a significant plant disease known as Cercospora leaf spot (CLS). This disease primarily affects beet crops, including sugar beets and table beets. Infected plants often experience reduced photosynthetic capability, leading to lower crop yields and reduced sugar content in sugar beets.

[0187] Indirect control of Cercospora beticola is done via the selection of beet cultivars with healthy leaves and the cultivation of the beets with at least a 3-year crop rotation. Markedly better control of the infestation may be achieved with a combination of resistant cultivars. Less susceptible Cercospora-resistant beet cultivars have been offered on the market since 2000. The resistance of these cultivars is based upon several genes53 KWS.224.01WO / KWS0460 and is quantitatively passed down, wherein the exact number of the genes that are responsible for the resistance is not known.

[0188] During the infection of the host, a pathogen often uses effector proteins that manipulate host cell processes to promote infection and pathogen survival. However, when recognized by specific plant resistance proteins, these proteins can activate plant defense mechanisms. Those effector proteins of pathogens are therefore also referred to as AVR (avirulence) proteins.

[0189] CRBM4 (Cercospora Resistance Gene in Beet) is a resistance gene found in beet plants. This gene encodes a protein involved in recognizing specific AVR proteins produced by Cercospora beticola. CRBM4 is part of the plant's immune system. It helps detect the presence of AVR proteins and triggers a defense response to inhibit pathogen growth and spread. When a beet plant with the CRBM4 gene encounters Cercospora beticola, the CRBM4 protein can recognize an AVR protein produced by the pathogen. This recognition event activates a cascade of immune responses in the plant, leading to the deployment of various defense mechanisms such as the production of reactive oxygen species, strengthening of cell walls, and activation of pathogen-related (PR) genes. The successful recognition of an AVR protein by CRBM4, which is based on the formation of a PPI between CRBM4 and an AVR protein, confers increased tolerance or even resistance to Cercospora leaf spot disease.

[0190] However, as the exact number and nature of all AVR proteins of Cercospora beticola are still not known, and as experimental validation of the ability of an AVR candidate to form a PPI with a plant protein, e.g., CRBM4, is time-consuming and expensive, the feature-vector based prediction method was used for identifying the most promising AVR-candidate able to form a PPI with the plant protein CRBM4 or any other plant protein.

[0191] In order to achieve this, the whole proteome of the sugar beet plant was derived from a database. All sugar beet protein sequences were transformed into amino acid sequence of the same length, namely 2000 amino acids, using some filtering steps in combination with padding. In addition, the whole proteome of the Cercospora beticola pathogen was predicted by standard bioinformatic means (FunAnnotate (Palmer, J. & Stajich, J. Funannotate. GitHub. https: / / github.com / nextgenusfs / funannotate (2023 ) and54 KWS.224.01WO / KWS0460 resulting putative proteins were analyzed for a potential characteristic as being part of either the secretome (SignalP, Nielsen, H., et al. A Brief History of Protein Sorting Prediction. Protein J 38, 200-216 (2019). https: / / doi.org / 10.1007 / sl0930-019-09838-3) and furthermore as being part of the effectorome (EffectorP, Sperschneider J and Dodds PN (2021) EffectorP 3.0: prediction of apoplastic and cytoplasmic effectors in fungi and oomycetes. MPML Abstract). Additionally, the cellular localization was predicted using e.g. ApoplastP (Sperschneider J et al. (2017) ApoplastP: prediction of effectors and plant proteins in the apoplast using machine learning. New Phytologist. doi:10.1111 / nph.14946) in order to identify potential secreted proteins with the characteristics of being an effector and being located in the apoplast (The apoplast is the network of cell walls and intercellular spaces within a plant, through which water and solutes can move freely). Their sequences of this subset of predicted proteins was used for further performance checks. These Cercospora protein sequences were transformed into amino acid sequences of the above-mentioned length using the same pre-processing steps.

[0192] The gene AVR-CRBM4 is a gene of Cercospora beticola which is detected by the sugar beet gene CRBM4. The AVR-CRBM4 gene was tested against the proteome of Sugar beet to identify further interacting genes. In order to accomplish this task, all combinatorially possible pairs of a sugar beet protein and a Cercospora beticola pathogen were formed and each protein sequence pair was input into a trained machine learning model that had been trained on a labelled plant-pathogen PPI training data set using a supervised learning approach. The machine learning model extracts a feature vector pair from each protein pair provided as input and returns a predicted likelihood of the two proteins provided as input forming a PPI.

[0193] Although this approach involved a whole-plan-proteome-vs.-whole-pathogen- proteome screen, Plot 402 selectively shows a histogram of the predicted probabilities that one of the pathogen proteins interacts with the specific plant protein CRBM4. More than 12.000 pathogen proteins were predicted to form a PPI with the protein CRBM4 with a likelihood of > 0%. However, only 46 proteins had a predicted likelihood of over 50%. Hence, the approach was able to provide a reasonably small number of AVR candidates that can be further evaluated experimentally. In particular, the computed PPI-55 KWS.224.01WO / KWS0460 formation likelihoods allow to pick the few pathogen proteins with the highest likelihood of forming a PPI for empirical validation.

[0194] Figures 5A, B are illustrations of a pipeline for PPI and interaction site prediction. The pipeline may be implemented on the computer system 200 or on a different computer system.

[0195] The computer system may comprise various software programs and / or functions 204, 504, 505, 506, and 212 for processing plant and protein sequences for predicting the formation and / or likelihood of PPI formation. The use of a model interpretation software for identifying and / or optimizing amino acid residues responsible for the predicted PPI. These software programs and functions have been described already with reference to figure 2. In the example implementation depicted in figure 5, the feature extraction and feature vector creation is implemented by embedding layers of a neural network 204. In the depicted example, the length of the plant and pathogen protein feature vectors is brought to the same unified length by a respective software functionality 506 of the network 204. However, in other embodiments, the feature extraction and feature vector combination may be performed by a different approach. Likewise, the length unification may be performed on the protein sequence level or on the feature vector level. In any case the feature vectors must have the same length before they are fed into the machine learning model 212 configured to predict the indication of whether one or more of the protein pairs input to the model will form a PPI.

[0196] In the depicted example, a vector pair selection engine 505, which may be part of the machine learning model 212 or may be implemented as a separate piece of software, computes all combinatorially possible plant-pathogen feature vector pairs and provides them as input to the model 212. The engine 505 may be implemented as a software loop and may be configured to select unique pairs of plant and pathogen protein feature vectors and feed these pairs into the model 212 until all combinatorially possible plant- pathogen-protein combinations have been examined or until another termination criterion is reached (e.g. until at least one protein feature vector pair with a PPI formation likelihood over a predefined minimum threshold has been identified.

[0197] The machine learning model 212 may comprise a sequence of machine learning models. For example, it may comprise a network 204 for creating the feature vectors and56 KWS.224.01WO / KWS0460 implementing the feature extraction function 504. It may comprise a Siamese network 314 for identifying the features having the highest relevance for PPI prediction. The model 212 may comprise a further neural network trained to predict whether a protein pair will form a PPI as a function of the output provided by the Siamese network. For each examined protein pair, the machine learning model, in particular, the last subnetwork 320, may output a score. Each score may indicate a predicted likelihood of PPI formation for each plant-pathogen pair whose feature vector pairs were provided as input. After processing all protein pairs provided to the model 212 as input, the output of the model may indicate the one(s) of the protein pairs having the highest likelihood of forming a desired PPI or the lowest likelihood of forming an undesired PPI.

[0198] In some examples, the indication 322 output by the model may allow the selection of a protein pair 508 of particular interest. For example, this protein pair 508 may be a protein pair having the highest likelihood of forming a (desired) PPI, or may be a protein pair having the lowest likelihood of forming a (undesired) PPI, or may be a protein pair having the highest likelihood of forming a PPI and in addition fulfilling further conditions. The further condition could be that the pathogen protein comprised in this pair is already suspected to be an AVR protein, or that the plant protein comprised in this pair is already suspected to be a resistance protein. The protein pair 508 may be used directly for experimental validation or as a positive or negative selection marker in plant breeding projects for obtaining a more pathogen resistant plant.

[0199] The at least one 508 protein pair may be subject to further evaluation. The pair 508 may be subject to further computational and / or empirical analysis, e.g., for predicting a sub-set of amino acids forming the interaction site of the PPI of the pair 508. The pair 508 may also be used as the basis for increasing or decreasing the strength of the PPI, preferably by a combination of computer-based predictions and genetic engineering techniques.

[0200] In some examples, for predicting the amino acids that are responsible for the formation of the PPI, and which often form the interaction site, the computer system 200 comprises a model interpretation software 216. The software is configured to analyze the trained machine learning model 212 for identifying, e.g. for the protein pair 508 of particular interest, or for any other analyzed protein pair for which the formation of a PPI57 KWS.224.01WO / KWS0460 was predicted, a sub-set of specific amino-acid residues in the plant protein sequence and / or in the pathogen protein sequence that had the highest impact on the prediction that said protein pair will form the PPI. For example, by comparing a large number of protein-protein feature pairs for which a high likelihood of a PPI formation was predicted with the respectively predicted likelihood, the model inspection software 216 may identify a subset of amino acids that were the reason for the model 212 predicting the formation of a PPI. The model interpretation software 216 may be an integral part of the machine learning model 212. In other examples, it is a separate software program.

[0201] According to some examples, the model interpretation software 216 comprises or is operatively coupled to a protein sequence variant generation software 514 or software function. The protein sequence variant generation functionality is configured to compute protein sequence variants (and hence also protein sequence pair variants) 510 from any given input protein pair 508 by permutating (i.e., permuting) one or more amino acids, in particular amino acids predicted to be responsible for the formation of the PPI, and feeding the variants into the pipeline of software functionalities 204, 505, 212 for computing the PPI formation likelihoods 322 also for the protein pair variants. At least one of the protein sequence pair variants (e.g., the variant with the strongest or weakest predicted PPI) may then be fed - together with the machine learning model 212 - into the model interpretation software for enabling the software 216 to predict the subset of amino acids that are responsible for PPI formation and that have been the cause for the model 212 to predict the formation of a PPI. Thereby, it is possible to identify protein sequence variants which have an increased or decreased likelihood of forming a PPI compared to the protein pair 508 used as basis for computationally creating the protein sequence pair variants. Optionally, also the protein sequence pair variants and the model may be fed into the model interpretation software 216 for identifying the sub-set of amino acids having caused the model 212 to predict the formation of a PPI.

[0202] Hence, the combination of computing protein sequence variants of an already interesting protein sequence pair 508 and a model interpretation software 216 for identifying interaction sites allows computationally identifying protein variants with an improved (e.g. stronger or weaker) PPI-interaction site that may not even exist in nature. Starting from e.g. a protein pair of interest 508 that was formed from existing or known58 KWS.224.01WO / KWS0460 plant and protein sequences, by identifying the subset of amino acids responsible for the formation of the PPI and selectively permuting the amino acids.

[0203] Permuting amino acids selectively in the sub-set of amino acids identified by the software 216 as being responsible for PPI formation may be beneficial because the number of protein sequence variants that need to be computed and computationally evaluated for identifying a protein pair with improved PPI formation capabilities is greatly reduced, thereby tremendously reducing processing time and CPU consumption.Permuting amino acids randomly at an arbitrary position in a sequence will often result in protein variants whose ability to form a PPI is not changed at all relative to the original protein sequence pair. To the contrast, by modifying the sub-set of amino acids identified by the software 216 to be the reason for the model 212 to predict the presence or absence of a PPI may allow to optimize (strengthen or weaken) a PPI with comparatively low computational effort: every change of an amino acid of the identified sub-set will highly likely have an impact on the strength of PPI formation. By feeding the protein pair variants 510 into the pipeline, a plant-protein pair variant 516 with an even better (higher or lower) PPI formation capability than the original protein pair 508 may be identified. The protein pair variant 516 may for example comprise a plant protein sequence variant which is not known, i.e., is not comprised in any protein database and also not known from in-house genomic screening projects.

[0204] The variant generation software 514 may be configured to selectively compute plant sequence variants, or pathogen protein variants, or both, in order to compute protein pair variants 510 from the at least one protein pair of interest 508. Preferably, at least the plant protein sequence of the pair 508 is modified for creating plant protein variants, as it is typically the plant, not the pathogen, whose genomic properties can be controlled by humans. Hence, when an important interaction (represented by the protein pair 508) has been identified, the available genetic diversity will be further increased by the variant generation software 514 which is capable of predicting alleles with even better PPI formation characteristics.

[0205] Preferably, the protein sequence pair 516 with optimized PPI formation capabilities, or at least the optimized plant protein sequence comprised therein, is empirically validated. Upon a successful validation, e.g., upon having confirmed using one59 KWS.224.01WO / KWS0460 of various PPI-analysis approaches that the predicted particularly strong or particularly weak PPI strength is actually observed in the real proteins, the plant protein sequence is used as a genomic marker for selecting individual plants for a plant breeding project to increase the resistance of the plant against the pathogen. In case no plant comprising the plant protein sequence of the "optimized" pair 516 exists, plants comprising the plant protein sequence can be created using genomic engineering techniques such as CRISPR- Cas9 or undirected mutagenesis approaches.

[0206] The use of a protein sequence analysis pipeline comprising a feature-vector based PPI prediction followed by the use of a model interpretation software for identifying the sub-set of amino acids that caused the model to predict the presence, absence (or generally: a likelihood of) a PPI may have the advantage that plant protein variants having improved pathogen resistance can be found reliably and with very low computational effort. The output of the model interpretation software does not only provide a clue on the nature and position of a subset of amino acids most relevant for the prediction, but also allows to efficiently create protein sequence variants that differ from each other specifically at these relevant sites, thereby easing and accelerating the identification of protein sequences with even better PPI formation capabilities.

[0207] In some embodiments, the protein sequence variants computed by the function 514 are used for supplementing the training data set of the machine learning model 212 and for retraining the model on the supplemented training data set. Thereby, the accuracy of the model 212 is improved, without introducing other biases that may be introduced by using a separate, model-independent software for predicting the amino acids relevant for PPI formation.

[0208] Figure 6 shows the software architecture of a computer system according to one example. According to the depicted example, the data transformation software 204 configured to perform the protein sequence transformation may comprise a software function 602, 604 for preprocessing and tokenizing the protein sequences. The preprocessing comprises in particular bringing all input protein sequences to a unified length, e.g., by filtering, padding and / or truncating. The pre-processed protein sequences having uniform length are then transformed into feature vectors (vector embeddings").60 KWS.224.01WO / KWS0460The software functions 602, 604 may comprise a large language model (LLM) trained on protein sequences for performing the amino-acid-sequence-to-feature-vector transformation. For example, the LLM may comprise one or more LSTM, CNN, and / or attention layers configured to perform tokenization and feature extraction. For example, the tokenization, feature extraction and embedding steps can be performed using pretrained models such as ProteinBert, providing about 1024-8196 features per amino acid., Hence, a feature vector, which may also be referred to as feature matrix, extracted from a 100 amino-acid-long protein sequence may have a dimensionality of 100 x 8196.

[0209] In some examples, the preprocessing for providing feature vectors of uniform length is performed on the sequence vectors, not the protein sequences.

[0210] The feature vectors extracted from a plant and pathogen protein pair by the preprocessing and feature extraction functions 602, 604 are input into a Siamese network 314 as described already, for example, with reference to figure 3. Optionally, two CNNs 606, 608 are used for processing the feature vectors provided by the software functions 602, 604, for extracting further features, in particular structural feature patterns, from the initially provided feature vectors, and for outputting the further extracted features to a Siamese network which is or has been trained to predict the formation of a PPI.

[0211] Downstream of the Siamese network, a fully connected deep neural network 320 (and optionally also another network architecture, e.g. a transformer) is used according to some example implementations for computing the PPI formation probabilities from the output of the Siamese network. The network downstream of the Siamese network for computing the PPI probabilities may be, for example, a fully connected CNN.

[0212] Preferably, the Siamese network is a multilayer Siamese network with more than 2 layers, in particular 3 or more layers. The use of a multi-layer Siamese network 314 (rather than, for example, a 1-D convolutional network) may have the advantage that the network 314 is able to process a greater number of features and thus provide prediction results with greater accuracy.61 KWS.224.01WO / KWS0460

[0213] The two feature vectors of the plant and pathogen protein pair, which have been optionally processed and amended or enriched with other features by two CNNs 606, 608, are passed through two identical sub-networks of the Siamese network configured to share weights. The two feature vectors computed by the optional CNNs may comprise high-level characteristics of each input protein sequence. The Siamese network compares the feature vectors to produce, alone or in combination with a downstream neural network, a likelihood score for PPI formation. For example, the Siamese network may concatenate the two modified feature vectors output by the optional CNNs 606, 608 and pass the concatenated vectors to a further network 320 downstream of the Siamese network. This further network 320 may evaluate the concatenated feature vectors and any other output generated by the Siamese network for further analysis and for computing a final indication if the evaluated protein sequence pair will likely form a PPI or not. Typically, the output of the network 320 may comprise a score 322 being indicative of a likelihood of PPI formation.

[0214] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed examples.62 KWS.224.01WO / KWS0460REFERENCE SIGNS LIST102-110 steps200 computer system202 processor(s)204 software for transforming protein sequences into feature vectors206 volatile or non-volatile storage medium208 protein sequences of a plant210 protein sequences of a pathogen212 machine learning model214 interface216 model interpretation software302 plant protein feature vectors304 pathogen protein feature vectors306 (padded) plant protein feature vectors308 (padded) pathogen protein feature vectors310 plant protein feature vector312 pathogen protein feature vector313 feature vector pair314 Siamese network316 sub-network of Siamese network for plant protein feature vector318 sub-network of Siamese network for pathogen protein feature vector320 downstream neural network322 predicted PPI likelihood score402 histogram of predicted PPI probabilities for CRBM4504 feature extraction software505 vector pair selection engine506 length unification software508 protein pair of particular interest510 sequence variants of protein pair of particular interest512 sub-set of amino acids responsible for PPI predictionKWS.224.01WO / KWS0460 sequence variant generation software protein sequence pair with optimized PPI , 604 software functions , 608 optional CNNs

Claims

64 KWS.224.01WO / KWS0460CLAIMS1. A method for predicting protein-protein interactions between plant proteins and pathogen proteins, comprising: by one or more computing devices: a) transforming (102) one or more protein sequences (208) of a plant into a respective plant protein feature vector (302; 306) and transforming (104) one or more protein sequences (210) of a pathogen into a respective pathogen protein feature vector (304; 308); b) forming (106) a plurality of different feature vector pairs (313) respectively comprising one (312) of the one or more pathogen protein feature vectors and one (310) of the one or more plant protein feature vectors; c) providing (108) the vector pairs to a trained machine learning model (212), and in response receiving (110), for the pairs of feature vectors, an indication (322) if at least one output protein pair (508) represented by one of the feature vector pairs is predicted by the trained machine-learning model to interact.

2. The method of claim 1, wherein the received indication is an indication that the at least one output protein pair is predicted to interact, the method further comprising applying a model interpretation software (216) for identifying a subset (512) of amino-acid residues in the plant protein sequence and pathogen protein sequence of the at least one output protein pair which had the highest impact on the prediction that the proteins of the at least one output protein pair will interact.

3. The method of claim 2, wherein applying the model interpretation software comprises computing relevance scores for the amino acid residues in the plant protein sequence and the pathogen protein sequence, the relevance scores indicating the contribution of the amino acid residues to the likelihood of the output protein pairs to interact, the method further comprising: using the relevance scores for identifying the sub-set of the amino acid residues.

4. The method of claim 3, further comprising visualizing the relevance scores as a heatmap overlaid on the plant protein sequences and / or pathogen protein sequences of the respective at least one output protein pair.65 KWS.224.01WO / KWS04605. The method of any one of claims 2-4, further comprising using experimental protein engineering for replacing one or more of the amino acids in the identified sub-set for increasing or decreasing the likelihood of interaction of the respective at least one output protein pair.

6. The method of any one of claims 2-5, further comprising: empirically testing the accuracy of the identification of the sub-set of specific amino acid residues having the highest impact on the prediction, and using the results of the testing for improving the trained machine-learning model.

7. The method of any one of claims 2-6, further comprising: applying the model interpretation software during the training of the machine-learning model after each training epoch to monitor evolution of residue relevance during model training, whereby in particular the model is retrained also based on feedback from the model interpretation software to enhance a prediction accuracy of the machine-learning-model.

8. The method of any preceding claims 2-7, wherein the application of the model interpretation software is integrated into an automated pipeline for continuous monitoring and updating of the machine-learning model as new protein interaction data becomes available.

9. The method of any preceding claims 2-8, further comprising: creating a plurality of plant protein sequence variants of the plant protein of the at least one output protein pair (508) by computationally permuting the sub-set of amino-acid residues identified by the model interpretation software in said plant protein sequence; and repeating steps a) to c) with the plant protein sequence variants for identifying at least one pair (516) of the plant proteins variants and pathogen proteins predicted by the trained machine-learning model to form a66 KWS.224.01WO / KWS0460 stronger or weaker PPI than the PPI predicted for the at least one output protein pair; and outputting the protein variant pair (516) identified to form a stronger or weaker PPI than the PPI predicted for the at least one output protein pair.

10. The method of any one of the previous claims, further comprising: receiving an initial set of protein sequences of the plant; transforming the initial set of plant protein sequences into a set of plant protein sequences having the same length; and using the set of plant protein sequences having the same length as the one or more plant protein sequences to be transformed into the plant protein feature vectors; receiving an initial set of protein sequences of the pathogen; analyzing the initial set of pathogen protein sequences for filtering out non-secreted proteins, thereby providing a sub-set of pathogen protein sequences; and using the set of pathogen protein sequences having the same length as the one or more pathogen protein sequences to be transformed into the pathogen protein feature vectors.

11. The method of any one of the previous claims, wherein the trained machinelearning model comprises a neural network configured to receive a pair of a protein feature vector and a pathogen protein feature vector as input.

12. The method of claim 11, wherein the neural network comprises a Siamese neural network (314) comprising a first (316) and a second (318) sub-network, wherein the Siamese neural network is configured to employ parameter sharing between67 KWS.224.01WO / KWS0460 the sub-networks, wherein the first sub-network is configured to process the protein feature vector and the second sub-network is configured to process the pathogen protein feature vector.

13. The method of any one of the previous claims 11-12, wherein the neural network comprises a network (320) having a transformer architecture, said network being in particular configured for receiving the output of the Siamese neural network, for computing, for each pair of plant-pathogen feature vectors, a likelihood of PPI formation, and for outputting a score being indicative of the likelihood.

14. The method of any one of the previous claims, wherein the feature vectors are vectors being indicative of chemical and / or electrical and / or structural properties of the amino acids constituting a protein.

15. The method of any one of the previous claims, wherein the feature vectors are sequences of feature sets, wherein the transformation of a respective one of the plant or pathogen protein sequences into a feature vector comprises: representing the protein sequence as a sentence consisting of a set of different characters representing protein-forming amino acids, the sentence comprising n tokens, wherein each token is a k-mer having a fixed size of k characters, the k-mers being derived by sliding a window having a length of k characters along the protein sequence, wherein in particular k is a number between 4 and 9, in particular between 5 and 8, in particular 6; computing the feature vector as a function of the sentence with the n tokens, wherein the features of the feature sets are derived from a respective one of the tokens of the protein sequence from which the feature vector was derived.

16. The method of claim 15, wherein the transformation of a respective one of the plant or pathogen protein sequences into a feature vector comprises:68 KWS.224.01WO / KWS0460 receiving a minimum window size and a maximum window size, wherein in particular the minimum window size is 1 and the maximum window size is a value smaller 12, in particular smaller 10; for each window size of all window sizes from the minimum to the maximum window size: perform the representation of the protein sequence and the computation of the feature vector in accordance with claim 15, wherein the sliding window length is equal to the window size; and use the feature vectors computed for the sliding windows of the different sizes ranging from the minimum window size to the maximum window size for performing the predicting by the machine-learning model.

17. The method of any one of the previous claims, wherein the feature vectors are sequences of feature sets, wherein the transformation of a respective one of the plant or pathogen protein sequences into a feature vector comprises: representing the protein sequence as a sequence of tokens, wherein each amino acids of the protein sequence is used as one of the tokens; computing the feature vector as a function of the sequence of tokens, wherein each of the tokens in the token sequence is transformed into a feature set comprising features of the amino acid represented by said token, and wherein the feature vector is a sequence of the feature sets.

18. The method of any one of the previous claims, further comprising using the at least one output protein pair (508) for at least one of: selectively breeding a plant comprising the plant protein of the at least one output protein pair predicted to interact or not to interact with the pathogen protein of the at least one output protein pair; selectively breeding a plant not comprising the plant protein of the at least one output protein pair predicted to interact or not to interact with the pathogen protein of the at least one output protein pair;69 KWS.224.01WO / KWS0460 identifying one or more individuals carrying a marker for use in marker- assisted breeding; genetically engineering a plant for generating a plant whose proteome comprises or does not comprise the plant protein comprised in the at least one output protein pair; using chemical mutagenesis methods for generating a plant whose proteome comprises or does not comprise the plant protein comprised in the at least one output protein pair; identifying candidate genes or candidate mutations adapted to increase pathogen resistance via the presence or absence of the plant protein comprised in the at least one output protein pair, and empirically testing the candidate genes or mutations or plants comprising the candidate genes or mutations; conducting a functional analysis of the plant protein; screening plant populations for plants whose genome results in the presence or absence of the plant protein comprised in the at least one output protein pair; predicting which varieties in a population might be most susceptible for a pathogen.

19. The method of any one of the previous claims 2-18, wherein the model interpretation software (216) is applied on the trained machine learning model and on the at least one output protein pair.

20. The method of any one of the previous claims 2-19, wherein the application of the model interpretation software comprises interrogating the machine-learning model, while it is processing or after it has processed the at least one output protein pair for predicting if the proteins of the pair interact, to trace back which70 KWS.224.01WO / KWS0460 amino acid residues or amino acid residue patterns had the strongest effect on the prediction of the machine-learning model.

21. The method of any one of the previous claims 2-20, wherein the model interpretation software is or comprises software separate from the machine learning model configured for predicting the PPI.

22. The method of any one of the previous claims 2-21, wherein the model interpretation software is or comprises an integral part of the machine learning model, wherein the integral part comprises attention layers.

23. The method of any one of the previous claims 2-22, wherein the trained machine-learning model is configured to compute, for each amino acid of the at least one output protein pair, a likelihood of protein-protein formation, and wherein the model interpretation software is configured to compute relevance scores that quantify the contribution of each amino acid residue to the predicted likelihood of protein-protein formation.

24. The method of any one of the previous claims 2-23, wherein the model interpretation software is selected from a group comprising: a machine learning model comprising long short-term memory - LSTM - layers; a software program or module implementing Layer-Wise Relevance Propagation - LRP; a software program or module implementing Local Interpretable Model- Agnostic Explanations - LIME; a software program or module implementing an attention mechanism, in particular one or more attention layers integral to a neural network structure comprising the machine learning model; a software program or module configured to make use of permutation importance; a software program or module configured to compute SHapley Additive exPlanations - SHAP - values for model interpretation and model description,71 KWS.224.01WO / KWS0460 wherein the SHAP values can in particular be KernelSHAP values or positional SHAP values - PoSHAP values; and a software program or module configured to perform Gradient-weighted Class Activation Mapping, Grad-CAM.

25. A computer system (200) configured for: a) transforming (102) one or more protein sequences (208) of a plant into a respective plant protein feature vector (302; 306) and transforming (104) one or more protein sequences (210) of a pathogen into a respective pathogen protein feature vector (304; 308); b) forming (106) a plurality of different feature vector pairs (313) respectively comprising one of the pathogen protein feature vectors and one of the plant protein feature vectors; c) providing (108) the vector pairs to a trained machine learning model (212), and in response receiving (110), for the pairs of feature vectors, an indication(322) if at least one output protein pair represented by one of the feature vector pairs is predicted by the trained machine-learning model to interact.