Bioinformatics method for determining therapeutic target regions

EP4591308A1Pending Publication Date: 2025-07-30CENT NAT DE LA RECH SCI (C N R S) +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023790046
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-20
Filing Date
2023-09-19
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Current methods for identifying therapeutic target regions for antiviral drugs are inefficient due to lack of precision in target identification, failure to account for biochemical interactions, and inclusion of unusable sequences, leading to high numbers of potential targets that require extensive and costly screening programs.

Method used

Integration of 3D structure and intramolecular chemical interaction analysis to validate and refine target residues, exclusion of low-quality sequences, and calculation of chemical interaction scores to select stable and relevant target regions, reducing the number of targets to be screened and improving the reliability of the method.

Benefits of technology

The enhanced method provides a more precise and reliable identification of therapeutic target regions, reducing the number of potential targets and increasing the likelihood of identifying effective drugs by focusing on stable and chemically relevant sites, thus streamlining the drug development process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) identifying (E1), in a set of previously aligned nucleotide and polypeptide sequences characteristic of said candidate protein, invariant residues and / or pairs of lethal synthetic residues referred to as target residues; b) identifying (E2) at least one candidate region consisting of at least one pair of target residues identified in step a), said at least one pair comprising target residues located at a determined distance in space and being exposed at the surface of the candidate protein, preferably in a pocket; c) determining (E3) the advantageous chemical interactions between each residue and / or between each pair of residues within the candidate protein, from the 2D and / or 3D structure of the candidate protein; said advantageous chemical interactions being hydrophobic bonds and / or hydrogen bonds and / or saline bridges and / or negative-repulsion and / or positive-repulsion bonds; d) selecting (E4) the residues linked by said advantageous chemical interactions, said residues being at a distance of at most 10 angstroms; e) selecting (E5) at least one therapeutic target region from among the candidate regions identified in step b), comprising the residues selected in step d).
Need to check novelty before this filing date? Find Prior Art

Description

[0001]TITLE: BIOINFORMATIC METHOD FOR DETERMINING THERAPEUTIC TARGET REGIONS TECHNICAL FIELD This application relates to a bioinformatics method for identifying reliable and sustainable therapeutic target regions, in order to optimize the search for new drugs, particularly antiviral drugs. STATE OF THE ART Despite the existence of treatments, RNA viruses still represent a serious public health problem. Indeed, their high mutation rate allows them to quickly acquire resistance to these treatments. To prevent the emergence of resistance, it is recommended to target invariant amino acids as a priority. Indeed, mutations in highly conserved positions lead to a deterioration or alteration of biological functions and could thus render the virus non-viable. However, due to their very small number,Invariant positions alone cannot constitute binding sites for a drug. To find other optimal binding sites accessible to a drug, Lao J. et al. propose to also identify pairs of covariant mutations called “synthetic lethals” (SLs) [Brouillet et al., Petitjean et al.]. SLs represent mutations that are not lethal but, when combined, render the virus non-viable. These SLs have already been studied in the search for anticancer drugs [Kuiken HJ, and Beijersbergen RL] and anti-HIV agents. Lao et al. propose a series of computational steps to identify the best target residues that, once mutated or blocked by a drug, could substantially affect the biological function of the targeted pathogen. The method proposed in Lao et al., however, has disadvantages. First,The targets identified by this method are not described in sufficient detail, which penalizes the user in their development program. For example, this method does not reveal whether the target proposed by the software has a large number of invariant residues, a specific volume, or is unlikely to mutate in the future. In addition, it does not take into account the fact that the initial batch of sequences may be poorly exploitable or, on the contrary, very reliable and therefore highly predictive. Finally, this method lacks crucial information to establish the relevance of the identified targets: the nature of the chemical bonds that may exist between the residues of the candidate protein, and more particularly in the target region. This biochemical information would strengthen the validity of the targets identified by the genetic approach of Lao et al. and select the most relevant ones. Generally speaking,the number of therapeutic target regions identified using a genetic screen by Lao et al. is too high. Since research organizations must then set up long and costly programs for screening libraries of molecules (antibodies, small molecules, etc.), it is preferable to reduce the number of target regions of interest as much as possible, by selecting only those that are likely to be long-lasting and stable, in particular by taking into account the biochemical links existing within the protein of interest. DISCLOSURE OF THE INVENTION According to a first aspect, the invention aims to improve the method of Lao et al. To this end, the present inventors propose to integrate into this method steps based on the 3D structure and the intramolecular chemical interactions of the target protein. These steps make it possible to confirm or validate the residues identified by the method of Lao et al.,according to the physicochemical quality of their interactions and their sensitivity to environmental parameters (pH, temperature, etc.). Thanks to these steps, only pairs of residues with advantageous physicochemical characteristics and located at a distance such that the chemical bonds between residues potentially have an influence will be selected. In addition, the present inventors propose to apply filters to the method of Lao et al. in order to exclude from the analysis nucleotide or peptide sequences of insufficient quality and thus avoid burdening the system by working on unusable or poorly indexed sequences. Thus, the present inventors have developed additional steps allowing more precise selection of the sequences to be tested and / or used as a priority. Finally, a method of this type must be able to generate results quickly, whether on proteins from different microorganisms,or on proteins from different variants of the same microorganism. It is therefore important to have a rapid and reliable method for processing several proteins in parallel, in particular allowing the 3D structures of proteins from known variants to be taken into account. All of these improvements make it possible to improve the reliability and interest of the method described in Lao et al., which was theoretical but difficult to use (because it was insufficiently described and poorly documented). The method of the present invention is more efficient and more reliable than that described in Lao et al., insofar as it has been scientifically enriched and supplemented by new criteria allowing the detected targets to be described in detail, both statistically and biochemically. Thanks to these improvements, the method of the invention becomes essential for identifying important target regions on pathogenic organisms,with a view to enabling researchers to design the drugs of tomorrow. To this end, a computer-implemented method is proposed for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) Identification, in a set of nucleotide and polypeptide sequences characteristic of said candidate protein, previously aligned, of invariant residues and / or pairs of synthetic lethal residues called target residues; b) Identification of at least one candidate region consisting of at least one pair of target residues identified in step a), said at least one pair comprising target residues located at a determined distance in space and being exposed to the surface of the candidate protein,preferably in a pocket; c) Determination of the advantageous chemical interactions between each residue and / or between each pair of residues within the candidate protein, from the 2D and / or 3D structure of the candidate protein; said advantageous chemical interactions being hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or bonds with negative repulsions and / or positive repulsions; d) Selection of the residues linked by said advantageous chemical interactions, said residues being at a distance of at most 10 Angstroms; e) Selection of at least one therapeutic target region from among the candidate regions identified in step b), comprising the residues selected in step d). The invention, according to the first aspect, is advantageously completed by the following characteristics,taken alone or in any of their technically possible combinations: - the step of identifying advantageous chemical interactions consists of a determination of the chemical interactivity of each residue within the protein and / or a determination of the chemical interactivity of each pair of residues spaced by a determined distance; - the distance determined in step b) is between 2 and 8 Angstroms, preferably 5 Angstroms; - the method comprises an evaluation of the chemical interactivity of the determined bonds comprising a calculation of at least one score from: a percentage of hydrophobic bonds and / or a percentage of hydrogen bonds and / or,a percentage of salt bridges and / or a percentage of negatively repulsive bonds and / or a percentage of positively repulsive bonds and / or a percentage of chemical bonds; - the target region is a pocket identified from a set of aligned 3D structures of the candidate protein and in which the chemical interactions are determined on a set of aligned 3D structures of the candidate protein; - the target pocket comprises at least five target residues chosen from invariant residues or synthetic lethal covariant residues, and in which said pocket has a volume between 60 ^̇, ^ and 500 ^̇ ^; - the method comprises an evaluation of the quality of the spatial location of the target residues, said evaluation being characterized by the calculation of the proportion of pairs of residues located less than 5 Angstroms from each other and / or the proportion of pairs of residues located less than 10 Angstroms from each other; - the method comprises a step of selecting the reference polypeptide sequence of the candidate protein, then a step of selecting the polypeptide sequences to be tested having lengths identical to the reference sequence, and / or having a number of mutations less than or equal to three standard deviations of the average number of mutations per sequence, and / or having an invariant residue at its N or C end; - the method comprises an evaluation of the quality of the selected initial sequences,said evaluation comprising the calculation of at least one score measuring the heterogeneity of the length of the initial sequences and / or the number of sequences having an aberrant number of poorly defined amino acids and / or the number of sequences having an aberrant number of missing residues. According to a second aspect, the invention aims to identify a therapeutic target region solely on the basis of the chemical interactions which exist between the residues of the region. In this respect, a computer-implemented method is provided for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) Determination of the advantageous chemical interactions between each residue and / or between each pair of residues within the candidate protein,from the 2D and / or 3D structure of the candidate protein; said advantageous chemical interactions are hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or negative repulsions and / or positive repulsions; b) Selection of the residues linked by said advantageous chemical interactions, said residues being at a distance of at most 10 Angstroms, said residues forming a therapeutic target region. DESCRIPTION OF THE FIGURES Other characteristics, aims and advantages of the invention will emerge from the description which follows, which is purely illustrative and non-limiting,and which should be read in conjunction with Figure 1 which illustrates a diagram of the steps of a computer-implemented method for determining at least one therapeutic target region on the surface of a candidate protein of a pathogenic organism. Figure 2 shows a diagram of the chemical interactions existing between the pairs of amino acids within the HA protein of the influenza virus (HB = hydrogen bond, SB = salt bridge, PR = positive repulsion, HI = hydrophobic interaction). The highlighted bonds are those corresponding to pairs already defined by the technique described in Lao et al and therefore confirmed by the chemical method of the invention,connecting pairs or groups of amino acids whose distance is suitable (highlighted in gray); these pairs will therefore be selected within the framework of the method of the invention. Figure 3 highlights the amino acids of interest selected in Figure 2 within the amino acids identified after implementation of steps E1 and E2 of the method of the invention (for more details, see Figure 3 of Lao et al.) DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a computer-implemented method for determining at least one therapeutic target region on the surface of a candidate protein of a pathogenic organism. Such a method may for example be implemented by a processing unit such as one or more processors or any other equivalent means. By "target pathogenic organism" is meant here any type of organism capable of causing diseases in a "host" (such as a human being,a plant or animal). This organism is preferably a microorganism such as a virus, a bacterium, a parasite, a fungus, a protozoan (amoeba, sporozoan or flagellates), etc. Some of these organisms have only been sequenced in the last few decades, such as: ^ Rotaviruses, Caliciviruses (Norwalk and hepatitis E), Ebola, chikungunya, SARS Coronavirus; ^ the bacteria Legionella, Campylobacter, Helicobacter, Mycobacterium and Escherichia coli strain O157:H7; ^ Sporozoa (parasitic protozoa) of the orders Coccidia (Cryptosporidium, Cyclospora, Toxoplasma) or Microsporidia (Enterocytozoon, Encephalitozoon, Nosema). It can also be a harmful macro-organism, such as a worm (belonging for example to the group of helminths, flatworms such as trematodes or cestodes, or nemathelminths such as the nematodes Ascaris, Toxocara, or Trichuris) or an insect. By extension,it is also included under the term "target pathogenic organism" tumor cells or cells infected by a pathogenic microorganism such as a virus, a bacterium, etc. Indeed, these cells often express on their surface or in their cytoplasm (or even in their nuclei) proteins involved in maintaining the proliferation signals of these cells, resistance to cell death, escape from the immune system, angiogenesis, activating invasion and metastases, replicative immortality, escape from growth factor suppressors, reprogramming of energy metabolism, which amplifies the disease (cancer or infection). It is therefore entirely logical to use the method of the invention to identify target regions suitable for therapeutic research, on proteins which are also involved in these disorders. Pathogenic organisms are essentially made up of proteins,some of which are necessary for their development, infectivity and / or pathogenicity. For each known pathogen, many proteins of this type have been identified and constitute a preferred target for researchers in the pharmaceutical industry. Indeed, invalidating the function and / or masking such proteins often makes it possible to stop the development, spread and / or deleterious effects of pathogens on human, plant or animal health. In the context of this application, these proteins will be called "candidate proteins", in that they have been previously identified as candidates potentially having an impact on the development, infectivity and / or pathogenicity of a pathogenic organism of interest. For a molecule to be selected as a "medicine", it is necessary that it substantially influences, in the short, medium or long term, the function of at least one of these candidate proteins. To do this,it must first be able to come into contact with the candidate protein (if possible, thanks to several contact zones). Thanks to 3D structures, it is now possible to determine which zones are located on the surface of the protein, and therefore which zones can be possible “contact zones” between a drug molecule and a protein of interest. For each protein, however, these contact zones are too numerous to be all screened in research programs. To facilitate their work and accelerate the identification of effective drugs, researchers need to know more precisely which “target regions” on the surface of candidate proteins have the most promising chemical and biological properties to be targeted by a drug. In the context of this application,a "therapeutic target region" is therefore called a privileged contact zone between a drug and a candidate protein, said zone having been selected in such a way that the interaction between these two elements (the drug on the one hand and the protein on the other hand) can be chemically strong, stable and significantly influence the biological function of the protein. Brief description of the steps of the method of the invention It is first considered that a candidate protein, which is known to act on the development, infectivity and / or pathogenicity of the pathogenic organism being studied, has been identified. To implement the method of the invention, this candidate protein must be known and well described in the literature. In particular, polypeptide sequences and 3D sequences must have already been characterized in the art,and be easily accessible. Nucleotide sequences coding for these polypeptide sequences must also be known. All these sequences are generally provided in official and openly accessible databases. The sequences of these databases, hereinafter referred to as “BDD1”, are generally anonymized, freely accessible and uploaded by different international research units. These are, for example, the Genbank, NCBI, INSDC, EMBL, HIVDB, LANL, FLUDB, GISAID, etc. databases, well known to those skilled in the art. In a step prior to the method of the invention, a polypeptide sequence and a reference nucleotide sequence must be selected (step E0). These reference sequences can be chosen from a phylogenetic tree (in this case, the reference sequence is at the root of the tree representing the sequences studied) or, if the tree does not exist, by recalculating an ancestral sequence,reconstructed by one of the following three methods: parsimony, maximum likelihood or Bayesian method. This reference sequence can also be a consensus sequence from the batch of sequences studied, but in this case it is only used to be compared to the other sequences, because a calculated consensus sequence can be a sequence that has never existed, so it is not necessarily functional. Consequently,The other known / listed nucleotide and polypeptide sequences for this candidate protein are identified in the “BDD1” databases. These additional sequences will hereinafter be referred to as “initial sequences”. A protein must perform a number of functions that can only be achieved if it adopts a certain structure in space and has the chemical radicals at the appropriate location. This is called the structure-function link. Thus, certain mutations change the structure of the protein and cause it to lose one or more functions. If these functions are essential, it follows that the virus becomes non-replicative and can therefore no longer develop. These mutations are mainly of three types: invariant positions, synthetic lethal pairs, and compensatory mutation pairs. A single mutated invariant position renders the protein non-functional, when to achieve the same goal,two positions must be mutated in the case of synthetic lethal pairs. Finally, a compensatory mutation pair is defined as follows: a first mutation makes the protein non-functional but a second could appear and restore the function of the protein studied. In these three cases, a strong constraint is imposed on the protein. In the context of the present invention, these initial sequences are first processed to select the invariant residues and / or the pairs of synthetic lethal residues (step E1). Once these invariant residues and / or pairs of synthetic lethal residues are known, at least one candidate region is identified, on the basis of other structural criteria (step E2). Steps E1 and E2 have been described by Lao et al. In the context of the present application,a "candidate region from E2" is a region containing at least two or three amino acids of interest which have been selected to be invariant residues and / or synthetic lethal (SL) covariant residues not derived from a common ancestor, close to each other, exposed on the surface of the candidate protein and possibly in a pocket, thanks to steps E1 and E2 of the present application, i.e. according to the method described in Lao et al. In a second step, the method of the invention provides for determining the chemical bonds involved globally between the residues of the protein, and in particular between the residues located in the candidate region (steps E3, E4, E5 in Figure 1). Thanks to this chemical interactivity information, at least one candidate region with the greatest chance of being an effective therapeutic target is identified. The method of the invention thus makes it possible to select,among the candidate regions obtained by following the indications of Lao et al., the therapeutic target region(s) most likely to allow the identification of effective drugs. As will be detailed below, one or more scores can be advantageously calculated to verify that the information from each step of the method of the invention is relevant for the rest of the process and in relation to the expected results. These scores make it possible to inform researchers about the quality and therefore the reliability of the results obtained. For the sake of readability of the description, the detailed expressions of the scores are given at the end of the description. Detailed description of the steps of the invention Step E1 consists of identifying, in a set of nucleotide and polypeptide sequences, characteristic of said candidate protein, and previously cleaned and aligned,invariant residues and / or pairs of synthetic lethal covariant residues, which will be called, in the context of the invention, “target residues”. Databases for storing the sequences of pathogenic organisms may contain erroneous sequences, which, if not eliminated, could generate false results. This is why, in the method of the invention, the initial sequences must first be “cleaned” (step E11). This “cleaning” consists of filtering the initial sequences, for example in the following manner: ^ Initial sequences having a length different from the sequence chosen as reference can be excluded (if the batch of initial sequences contains enough sequences); ^ If the batch of initial sequences does not contain many sequences and if the alignment seems possible (because the sequences do not contain hypervariable regions, and few poorly described regions),an alignment can be carried out to homogenize the lengths. In practice, from sequences of variable lengths, an alignment of a single length is obtained, which is greater than the length of the sequence with the greatest length. ^ The initial sequences having a number of mutations greater than three standard deviations of the average number of mutations per sequence are excluded. Furthermore, the N-terminal and C-terminal ends of the initial sequences are identified and the sequences are oriented in the same direction, so that their ends can be superimposable (for example, all the sequences can be oriented from the N-terminal end to the C-terminal end). This cleaning step E11 also makes it possible to identify, among all the initial sequences, those having heterogeneous lengths and / or to discard sequences which have unacceptable anomalies. Such anomalies are, for example,an aberrant number of mutations, an aberrant number of poorly defined amino acids, an aberrant number of missing amino acids, etc. In this regard, one or more scores or score functions can advantageously be calculated (step E11') to evaluate whether the initial sequences (before cleaning) do not present too many anomalies and / or whether their length is sufficiently homogeneous (scores S1, S2, S3, S4 detailed later). If one or more of these scores is not acceptable (value close to 0 and not 1), the method of the invention can be interrupted, because this means that the set of initial sequences before cleaning is not sufficiently robust to be exploitable. A score function noted f(S1QS) can also be calculated at this stage,in order to assess whether the set of initial sequences before cleaning is sufficiently complete and robust to effectively predict the existence of a therapeutic target region within the selected candidate protein. This score function f(S1QS) reflects the impact of the sequence data on the quality of the result obtained at the end of the process. It is also recommended to interrupt the process when the number of mutations belonging to the batches of initial sequences retained after “cleaning” is too low (typically, a number of mutations not allowing statistically correct results to be obtained, for example which does not allow the calculation of ^, ^(covariants with less than 5 representatives)). In this case, it is preferable to select a new set of sequences from the BDD1 database, to enrich it by downloading new sequences, or to change the candidate protein. Conversely, we consider that we have a good set of initial sequences when at least 1000 sequences of acceptable quality have been identified, these sequences carrying enough mutations. By “acceptable quality”, we mean here initial sequences presenting a number of anomalies less than three standard deviations of the average number of these anomalies per sequence. In this regard, the S7 score can be calculated, to measure the total number of initial sequences not containing an aberrant number of mutations. Such an S7 score depends on the S2, S3 and S4 scores presented above. Then, an alignment of the sequences obtained after cleaning is advantageously carried out (step E12).Sequence alignment can be implemented in two ways: i) either the sequences all have the same length: in this case the alignment is generated by aligning all the sequences on their N-terminal end, ii) or they do not all have the same length: in this case it is difficult to use multiple alignment methods, which are generally used for a smaller number of sequences. In this case, it is possible to select a representative sample of the population of sequences studied and to use the HMMER suite (http: / / hmmer.org) which allows to generate a profile on which all the sequences can then be aligned one by one (using the hmmbuild and hmmcalibrate functions). This method allows to align hundreds of thousands of sequences in a time of the order of a minute.The quality of the sequence alignment can be assessed (step E12') by calculating one or more scores (detailed later) relating to the impact of gaps in the alignment (score S8), the redundancy of the sequence set (score S10), the impact of hypervariable regions (score S11), the impact of deletions and insertions (score S12), the impact of post-translational modifications (score S13), and / or the impact of the existence of different subtypes (score S14). The scores S11, S12, and S13 are defined from data present in the literature. One or more score functions can be calculated at this stage to assess whether the set of aligned sequences after cleaning is sufficiently complete and robust to effectively predict the existence of a therapeutic target region within the selected candidate protein. The score function f(S3) detailed later reflects the quality of the alignment and the possible prediction.Some terms of this function describe the precision with which the initial sequences are described and have an impact on the statistical results (f(S3QS) and other terms of this function show the heterogeneity of the batch of sequences studied and have an impact on the description of the target itself f(S3SC). More precisely, the score function f(S3QS) translates the impact of the alignment on the statistical prediction of the target region and the score function f(S3SC) translates the impact of the alignment on the prediction of the target. If the predicted quality of the alignment is low, the process can be interrupted to be restarted from new initial sequences. At the end of this alignment step, the user has a set of cleaned and aligned sequences called “sequences to be tested”. In these sequences to be tested, the invariant residues are then identified (step E13).“Invariant residues” are by definition amino acids that do not change (or almost do not change) their position within the protein, in all the sequences to be tested studied. Since sequencing methods are not 100% reliable, an average error rate of 0.3% can be applied (Cheng C. et al, 2022). Thus, in the context of the present invention, a residue is defined as “invariant” if it is present at the same position on at least 99.7% of the sequences to be tested. The quality of the selection of invariant residues can advantageously be evaluated by calculating the mutational richness of the sequences (step E13'). In particular, the S18 score can be calculated (see below). Thanks to this score, it is possible to verify that the sequences on which the invariant residues have been detected are sufficiently heterogeneous.Indeed, to be able to state that a residue is invariant for functional reasons (mutations appeared at this position, but were not selected because they were lethal) it is necessary to be able to show that mutations appeared elsewhere on the sequence. Once the number of invariant residues is known, it is advantageous to determine the percentage of these residues by calculating the S19 score. This score makes it possible to evaluate the impact of invariant positions on the final result. The targets most likely to be stable in the long term (therefore being unable to mutate without modifying the replicative activity of the virus, thus preventing the appearance of mutations that could make the mutants resistant to treatment) are those that are the most invariant, therefore made up of the largest number of invariant residues. It is also possible to calculate the score function ^. ( ^^^^^^^^^^^ )which allows to evaluate the degree of invariance of the batch of sequences studied, and gives an idea of ​​its long-term mutational incapacity. This degree of invariance depends on several variables such as: the quality of the alignment, the total number of mutations, the number of invariants, the number of synthetic lethals (which represent a two-residue invariance) and the number of mutations which appeared for functional reasons and not due to the presence of a common ancestor. As specified in Figure 1, other residues of interest can also be selected in the method of the invention. These are residues which are not invariant, but which belong to pairs of covariants of interest (step E14). Several statistical tests make it possible to define the covariation of a pair of variables in a list of variables. The statistical test of the ^ ^is one of them. It makes it possible to determine, according to a threshold, whether the variables taken two by two are independent of each other. In the context of the present invention, it is notably possible to use the ^ ^ de Noivirt defined rather in the form (where A and B are the specific amino acids found at positions i and j), which allows to take into account each residue among 20 and not only the mutated or unmutated state of the original residue. Once this defined, it is preferable to readjust the results to reject false positives due to the multiplicity of tests performed. Thus, the p-values ​​can be readjusted using a method known as the false discovery rate. Residues are considered “dependent” or “covariant” if their p-value is below 0.05 (Noivirt O. et al.). The impact of the number of covariant positions as well as the impact of the number of pairs of covariant positions on the final result can be assessed by calculating the S20 and S21 scores respectively. The pairs of covariant residues identified by this calculation must then be filtered (steps E15 and E16). First, it is preferable to eliminate pairs of covariant residues that share common ancestors (step E16).This can be achieved in particular by studying the DNA sequences coding the initial sequences: after aligning these DNA sequences, the mutated codons causing non-synonymous mutations are identified and selected. Indeed, non-synonymous mutations induce the appearance of amino acids that are different on a physicochemical level; while, on the contrary, synonymous mutations code for the same residue. A linkage disequilibrium coefficient D', called Lewontin, can then be calculated with all the recoded DNA sequence data (differentiating synonymous positions from non-synonymous ones). Using this coefficient, couples sharing the same ancestor, whose covariation is therefore not the consequence of functional interdependencies, can be identified.Thanks to these steps, it is possible to determine whether the covariation of the two residues identified as “covariants” at step E14 results from the coevolution of these two residues or if it is due to the fact that they are phylogenetically linked to a common ancestor. In the latter case, the pairs of residues are not conserved in the method of the invention; this is then a “false positive” called “ancestral linkage disequilibrium” (Lao et al.; Petitjean et al.). At this stage, it is possible to calculate S24 and S25 scores which accurately reflect the impact of global non-synonymous mutability and the impact of synonymous mutability. Finally, it is advantageous to identify whether, among the pairs of covariants selected at the end of step E14, some are capable of inducing the impossibility for the virus to replicate (step E16). These particular pairs of covariants are known as “lethal synthetics.”In fact, it is important to distinguish between the two types of covariant mutations that exist, and which have completely opposite consequences: compensatory mutations (CM) and synthetic lethal mutations (SL) (Lao et al.; Petitjean et al.). To do this, it is possible, for example, to calculate a so-called dissimilarity coefficient ξ, which allows the test of ^ to be signed. ^ (^^, ^^). This sign allows us to differentiate between CM and SL. CM has a ξ which is positive when NobsA,i,B,j ≥ NexA,i,B,j with A and B two residues located respectively at positions i and j. Thus ξA,i,B,j = + while SLs have a ξ which is negative when NobsA,i,B,j ≥ NexA,i,B,j . Thus ξA,i,B,j = -^ ^(^^, ^^) with Nobs the number of pairs of residues A and B observed at positions i and j, Nex the number of pairs of residues A and B expected at positions i and j (Petitjean et al.). At this stage, it is possible to calculate the S22 score which reflects the impact of the number of SLs and their strength on the result obtained at the end of the process. Similarly, the S23 score can also be calculated, as it reflects the impact of the number of CMs and their strength on the result. The strength of a pair of covariants is defined by its ^ ^ (^^, ^^). Indeed, the more the ^ ^(^^, ^^) is high the higher the number of observed pairs is compared to the number of pairs expected if there were no covariation. S22 and S23 therefore define the strength of the covariation for this specific pair. At the end of these different steps, invariant residues and pairs of synthetic lethals are selected as being part of at least one “candidate region” of the candidate protein studied. It is here possible to calculate the S9 score which reflects the impact of the variance at the 5' and 3' ends of the sequences. Indeed, in the step which consists of aligning the DNA sequences to identify the covariant residues which share a common ancestor, one of the two techniques used consists of aligning all the sequences on their 5' end. However, the variance of this end reduces the chance of correctly aligning the sequences. It is then preferable to align on the 3' end, which must in turn be little variant.The candidate region that will be selected as the best therapeutic target must also satisfy a number of other advantageous conditions. In particular, the amino acids that compose it may be subject to constraints due to the 3D structure of the candidate protein. Conversely, it may be advantageous to determine whether the candidate region contains a “pocket” that would allow a drug to lodge in a stable and strong manner. To evaluate these aspects, the method of the invention contains a step E2 that evaluates the quality of the candidate regions obtained in step E1 with respect to the position of the amino acids that compose it, relative to the 3D structure of the candidate protein. Step E2 therefore consists of identifying, among the regions and residues identified in step E1, at least one pair of target residues located at a close distance in space, and being exposed on the surface of the candidate protein, preferably in a pocket.To be able to bind a small molecule, a therapeutic target must be composed of residues that are close to each other in space. Also, in the context of the present invention, the target residues of the candidate region will preferably be separated by a maximum of 10 Angstroms, preferably a maximum of 5 Angstroms. It is advantageous here to quantify by calculating different scores the target residues that are less than 5 Angstroms and / or 10 Angstroms apart, in order to evaluate whether they are sufficient in number to constitute a future target capable of binding a potential small molecule (drug). The scores S28, S29, S30, S31 can in particular be calculated (step E21'). To obtain this information, the method of the invention advantageously requires access to the known three-dimensional structures identified for the candidate protein.These structures are known and described in dedicated 3D structure databases, which will be called “BDD2” here. As for nucleotide and polypeptide sequences, these 3D sequences must often, prior to the method of the invention, have been processed (cleaning step E6, alignment step E7). The 3D structures are in fact often in pdb format. However, this format is not applied in the same way by the entire scientific community (format of the file itself, names of subunits or numbering of residues, etc.). Thus, a cleaning (step E6) of the structures is preferably implemented, making it possible to define a single and generalized format for all the pdb files used (standardization of the file format, numbering of residues and atoms, names of subunits and their respective positions, among others). In addition, dozens or even hundreds of structures of the same protein are available.Just as with sequences, the 3D structures identified for the same protein are preferably aligned. These are structural alignments that allow us to know how different the structures identified for this protein are. The alignment is carried out using a reference model. We can then calculate the average difference between all these structures, which is called "RMSD". A three-dimensional structure is defined by the position in space of each of the atoms that constitute it. If many of these positions are missing (missing data), the 3D structure is wobbly in the sense that only some of these atoms have a specific place in the structure.If a lot of data is missing for each of the structures studied (knowing that the missing data are not located at the same position in the protein space), it becomes difficult to make a structural alignment and even to compare the residues located at the same position. The fact that a protein has several subunits further complicates its structure. Moreover, in this case, we speak of a quaternary structure and no longer only a tertiary structure. In this case, it is necessary not only to define the position of each atom of each subunit, but also the position of each subunit relative to each other. Under these conditions, the quality of the 3D structures can therefore be evaluated to verify that they will effectively predict the existence of a target region. Several scores are thus calculated (step E7').In this respect, we can calculate the S5 score, which quantifies the number of 3D structures with an aberrant number of missing data, as well as the S6 score, which evaluates the impact of the existence (if any) of different subunits on the final result. The S15, S16, and S17 scores also allow us to evaluate the alignment of the 3D structures. S15 is, for structures, the equivalent of Sseq for sequences. We look at whether the number of known 3D structures for the candidate protein is high or not. If there are fewer than 200 known 3D structures, this number is insufficient to have a consistent average (this figure can be reduced, however, as 3D structures become more and more reliable). S16 gives an idea of ​​the variation in structure resolutions. Indeed, 3D structures are determined using different techniques. The two most used are X-ray diffraction and electron microscopy.A resolution threshold is defined for each of these structures, depending on the technique used. S16 evaluates this variation in resolution. S17 shows the structural heterogeneity of the batch of structures studied. Each of the atoms in the structure is aligned with the corresponding atom in the next structure. Thus, for each atom, we can associate a number of positions equal to the number of structures studied. An average value of this position in space is calculated as well as the deviation from the average. If we perform this calculation for all the atoms in the structure, it becomes possible to calculate an average deviation that gives an idea of ​​the heterogeneity of the structures, spatially speaking. The TM described in this score is a variant of the RMSD that allows it to be normalized so that it is not dependent on the total number of positions in the structure. We can also calculate score functions evaluating the quality of the structural alignment via the functions. ^^^ ^^^^, ^^^ ^^^ ^. The function ^ ( ^ ^ ) reflects the quality of alignment and possible prediction. Some terms in this function describe the accuracy with which 3D structures are described and have an impact on statistical results ^ ^ ^ ^^^ ^ and other terms of this function show the heterogeneity of the batch of sequences studied and have an impact on the description of the target itself More precisely, the score function ^ ^ ^ ^^^ ^ translates the impact of alignment on the statistical prediction of the target region and the score function translates the impact of alignment on target prediction. From the 3D coordinates of all residue atoms in space, pairs of residues that are close in space, i.e. less than 5 or 10 Angstroms apart, are selected (step E21). Furthermore, the accessibility of target residues and / or their exposure on the protein surface must be taken into account. Indeed, an effective therapeutic target must not be buried in the 3D structure of the protein, otherwise the drug will not be able to reach it. Thus, from the three-dimensional structure of the candidate protein, the residues that are buried in the protein are distinguished from those that are exposed on its surface and therefore accessible (to do this, it is possible to use the ASA program for example).In the present method, only candidate regions containing at least two, preferably at least three, accessible residues are selected (step E22). Here again, scores reflecting accessibility can be calculated to assess the possibility of obtaining sufficient candidate region(s) (step E22'). The score S32 evaluates the percentage of accessible positions and the score S33 evaluates the percentage of accessible residues. Finally, it may be advantageous to assess whether the previously selected candidate region has a 3D structure comparable to a “pocket” (step E23). To do this, it is possible, for example, to use structure prediction software, for example the Fpocket software (Le Guilloux et al.), in which, to take into account the existence of small and large pockets, it is preferable to reduce the minimum and maximum radii of the alpha spheres to 2.5 Å and 4 Å respectively. The Fpocket software (Le Guilloux et al., 2009) allows to determine all the pockets, whose cardinal is ^. ^^ ^ ^^ ^ ^^ ℎ ^ ^ ^ on the surface of a protein (providing a three-dimensional structure as input). The volume of a pocket that can accommodate a small drug molecule preferably meets the following constraints: 60 Å ^ < ^^^ℎ^ < 500 Å ^ ( ^^^^^^^^ℎ^^^^^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ) The percentage of pockets meeting this criterion for the protein studied can be calculated ^ ^^^^^ = ^ ^ ^ ^ ^^ ^ ^ ℎ ^^ ^ ^^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ÷ ^ ^^ ^ ^^ ^ ^ ℎ ^^ ^ ^. We can also determine the sum of the cumulative volumes of the pockets 60-500: which is ^^^^^^^^^^^^^^(step E21'). At the end of this step, the networks of spatially close SL invariant or covariant residues, which are exposed on the surface of the candidate protein and preferably in a pocket having a volume between 60 Å3 and 500 Å3, will finally be retained. Preferably, networks of at least five target residues will be retained. In a complementary manner, the SL invariant or covariant residues can be represented in the form of graphs (step E8). Indeed, from the detected networks, groups of interdependent residues are formed by means of mathematical graphs. In these graphs are integrated the invariant residues and the pairs of synthetic lethal residues which form invariance groups.The nodes of the graph are the residues, the edges define the link between the residues (synthetic lethal or invariant, in this case they are linked to all) and only exists if the two residues are positioned less than 10 Angstroms apart and on the surface of the candidate protein. Advantageously, this type of graph can be visualized using the free software Graphviz. These graphs are a way to evaluate the quality of the identified regions. Step E3 consists of determining whether there are advantageous chemical interactions between the residues and / or between the pairs of residues within the candidate protein, based on the 3D structure of the candidate protein. This step E3 can alternatively be limited to the analysis of the advantageous chemical interactions existing between the residues and / or between the pairs of residues selected within the candidate region obtained in step E2, based on the 3D structure of this candidate region.Genetics allows functionally detecting the impact of microscopic physicochemical changes occurring at the residue level. Thus, being invariant (or being part of an invariance group) is the consequence of the physicochemical quality of one or more residues (selection pressure requires the maintenance of this / these residues at this position). Conversely, if the physicochemistry of a residue is changed (for example by a mutation or by the external environment), the function and / or structure of the region may be affected. Based on this observation, it is recommended to take into account the physicochemical links existing between the residues of the proteins studied, to strengthen or invalidate the results obtained previously, with the aim of identifying the most relevant target regions.Thus, it is necessary to determine, for each amino acid of the candidate protein, the chemical bonds (for example hydrogen, hydrophobic, ionic and repulsive) in which this amino acid could potentially participate (step E3 of Figure 1). In particular, it is necessary to identify, from the three-dimensional structure of the protein, the pairs of residues whose amino acids are sufficiently close so that the identified chemical bonds can effectively influence the function and / or stability of these residues (see step E4 of Figure 1). The additional steps E3 and E4 therefore make it possible to highlight pairs of amino acids having a chemical interaction influencing their function, or likely to evolve during modifications of the protein's environment (pH, temperature, ionic strength, etc.).Protein stability is indeed often pH-dependent, and varies depending on the subtype and / or origin of the host organism. For example, the hemagglutinin (HA) of the human influenza virus is more stable than the hemagglutinin of the avian virus (Galloway SE et al.). Moreover, mutations can lead to changes in the chemical interaction network, and stabilize or destabilize this protein (Byrd-Leotis L. et al.). The intraprotein physicochemical network is made up of several types of interactions, such as hydrophobic, hydrogen, as well as salt bridges, negative repulsive and positive repulsive interactions, among others (Dyson HJ et al.; Hubbard RE & Kamran Haider M.; Sticke DF et al.; Barlow DJ & Thornton J. M; Harrison JS et al.). These interactions allow the protein to have a very particular shape: this is why they are important.If they were not there, the structure of the protein and surely its function would be modified. This is why these regions rich in these chemical bonds are important to determine, in the context of the method of the invention. Unlike hydrophobic and hydrogen interactions for which pH sensitivity is negligible (or insufficiently documented), electrostatic interactions are strongly affected by pH variations (Harrison, J.S. et al.; Pahari S. et al.). Histidine residues (pKa ≈ 6.4) are biological pH sensors because they are partially charged at neutral pH and positively charged at acidic pH (Pahari S. et al.; Kampmann T. et al.). Arginine and lysine (pKa ≈ 13.8 and 10.7) are more basic and invariably protonated under physiological conditions (Pahari S. et al.; Fitch, C. A. et al.). However, large fluctuations in pKa can occur depending on the microenvironment (Pahari S. et al.; Harris TK& Turner GJ; Harms MJ et al.; Di Russo NV et al.; Baumgart M. et al.). Therefore, the pKa of the negatively charged carboxyl group of aspartate and glutamate (pKa ≈ 3.4 and 4.1) approaches the pH values ​​reached during endosomal maturation (Pahari S. et al.; Mellman I., et al.). Therefore, the breaking of salt bridges (or at least their weakening in the case where the hydrogen bond remains) can occur if pH < pKacarboxyl group (Meuzelaar H. et al.). Finally, cation-cation and anion-anion repulsions induce significant destabilizations (Harrison JS et al.). When the pH decreases, negative repulsions can also be disrupted and potentially form hydrogen bonds. The method described here therefore contains a step of characterizing the chemical interactions existing between all the atoms of the protein, or at least those present in the candidate region selected previously.This step aims in particular to identify the following interactions (Hubbard RE & Kamran Haider M.; Barlow DJ & Thornton J. M; Harrison JS et al.; Donald JE et al.; Freitas RF de & Schapira M.; Onofrio A. et al.): - Hydrophobic bonds between non-polar atoms of the type (CB, CG, CE, CD1, CD2, CE2, CE3, CZ2, CZ3, CH2, CE1, CZ, CG1, CG2, CD, CH2) belonging to the following hydrophobic amino acids (ALA, MET, TRP, PHE, TYR, VAL, LEU, ILE, PRO). Atoms presenting this type of bonds will be selected when their distance is less than 4.2 Å. - Hydrogen bonds consisting of a proton donor atom (OG, OG1, NE2, ND2, ND1, NE2, NZ, NE, NH1, NH2, OH, NE1) belonging to a proton donor amino acid (SER, THR, GLN, ASN, HIS, LYS, ARG, TYR, TRP) and a proton acceptor atom (OG, OG1, OE1, OE2, OD1, OD2, ND1, NE2, OH) belonging to a proton acceptor amino acid (SER, THR, GLU, ASP, GLN, ASN, HIS, TYR).Atoms with this type of bond will be selected when their distance is less than 3.5 Å (Hubbard RE & Kamran Haider M.; Sticke DF et al.). These hydrogen bonds can consist of interactions between two side chains or between a side chain and the main chain of the protein. The proton donors (N) of the main chain or the proton acceptors of this same chain can belong to any amino acid of the main chain. - Electrostatic bonds, of two kinds: o Attractive bonds (salt bridges or ionic interaction) when a cation (NE, NH1, NH2, NZ, NE2, ND1) of a positively charged residue (ARG, LYS, HIS), in particular when it is at a distance of less than 4 Å from an anion (OD1, OD2, OE1, OE2) carried by a negatively charged residue (ASP, GLU).o Repulsive bonds consisting of two atoms with identical charges (- / -) or (+ / +), especially when they are separated by a distance of less than 5 Å (Barlow DJ & Thornton JM; Harrison JS et al.). All positively and negatively charged atoms can be taken into account in the calculation, as long as the charges are equally distributed between the ionizable groups, by stabilizing the resonance between the charges. Different scores can be calculated to evaluate the chemical interactivity of each amino acid within the candidate region or within the protein. In particular, it is possible to determine the percentage of hydrophobic bonds and / or the percentage of hydrogen bonds and / or the percentage of salt bridges and / or the percentage of negative repulsions and / or the percentage of positive repulsions and / or the percentage of chemical bonds (step E3').In practice, the method of the invention therefore advantageously contains a step of calculating the distance between each atom of the protein, for example from a pdb, SwissProt, uniprot or Modbase file. Preferably, the distances between atoms belonging to the same position are not calculated unless they belong to different protomers. Following this calculation, the atoms separated by a distance of at most 5 Å are selected. This step is preferably carried out before listing the chemical interactions between the atoms. Thus, only the chemical interactions linking nearby atoms are to be taken into consideration. It is possible, however, to carry out the two steps in reverse order. From the number. ^ ^ ^^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^representing the total number of pairs of atoms located within 5 angstroms, several scores can be calculated. In particular, it is possible to calculate the percentage of hydrophobic bonds and / or the percentage of hydrogen bonds and / or the percentage of salt bridges and / or the percentage of negative repulsions and / or the percentage of positive repulsions and / or the percentage of chemical bonds (step E3'). It is also possible to calculate a score function ^ ( ^ ^^^^^^ )which gives an image of the overall chemical interactions of the protein studied is calculated (step E3'). Thanks to all these steps, the residues linked by said advantageous chemical interactions, and being at a distance such that these bonds influence the function of these residues, are selected. This is step E4 in Figure 1. At this stage, it is possible to represent the different chemical bonds in the form of graphs (step E8'). Indeed, from the previously detected bonds, mathematical graphs are formed. The nodes of the graph are the positions / residues, the edges define the chemical link that exists between the positions. In a complementary manner, this type of graph can be visualized using the free software Graphviz. These graphs are an additional means of evaluating the identified bonds.Step E5 identifies the best therapeutic target region(s) for generating a potential drug by cross-referencing the results obtained at the end of step E2 with those obtained at the end of step E4. The quality of the selected targets can advantageously be evaluated by calculating several score functions. For example, it is possible to calculate the invariance of each target region by the function ^. ( ^ ^^^^ ) . The invariance of each target region ensures that it will be stable over time. In addition, the function ^(^ ^ ) allows you to obtain a score between 0 and 1 for each target S X.This score allows to evaluate how effective the identified target is in retaining the drug. In other words, these two scoring functions allow to evaluate the effectiveness of the target region. The target regions can advantageously be represented in the form of graphs (step E8''). Indeed, from the previously detected pairs, groups of interdependent residues using mathematical graphs are formed. In these graphs are integrated the invariant residues and the pairs of synthetic lethal residues that form invariance groups, and which are linked by the advantageous chemical bonds. The nodes of the graph are the selected positions / residues, the edges define the link that exists between the positions (synthetic lethal, invariant, or advantageous chemical interaction) and only exists if the two residues at this position are positioned less than 10 Angstroms and on the surface of the candidate protein.In addition, this type of graph can be visualized using the free software Graphviz. These graphs are an additional way to evaluate the quality of the identified regions. The complexity of the graphs can be calculated (step E8''') by the S26 score as well as the number of connected graphs by the S27 score. Such complexity is useful: if the graph is complex, this reflects a significant number of targets. If it is very complex, there may be intersections between non-empty targets. If there are many sun subgraphs, this means that a residue is linked to many others and that the majority of the function falls on it. Definitions of the S1 scores: ^. ^ : Heterogeneity of sequence length. Let us consider ^ the mean and ^ the standard deviation of the amino acid lengths of the sequences in the sample studied. ^^ ^ > ^ ^^^^^ ^ ^ = 0 A score ^ ^ = 1 indicates absolute heterogeneity while ^ ^= 0 indicates no heterogeneity. S2: ^ ^^^^^ : Sequences with an aberrant number of mutations. Let ^ be the mean and ^ the standard deviation of the number of mutations calculated per sequence of the sample studied, in comparison with a reference sequence. The sample studied consists of the polypeptide sequences downloaded before the cleaning steps and the reference sequence introduced in step E0 previously described. Knowing that the number of aberrant mutations is defined as ^^^ ^^ >= ^ + 3 × ^ and that ^^^ ^^ ^^^^^ corresponds to the number of sequences having a mutation number greater than or equal to ^^^ ^^ . We can then calculate: ^ ^^^^^ = 1 − (^^^ ^^ ^^^^^ ÷ ^^^ ^^ ) with ^^^ ^^ the total number of sequences. The closer S2 is to 1, the greater the number of aberrant mutations. S3: ^ ^^^^^: Sequences with an aberrant number of poorly defined amino acids (AA) (noted NYP). In this respect, it is noted that during the sequencing step of nucleotide sequences, it is sometimes impossible to define exactly the exact nitrogenous base (A, T, G, C) found at this location although it is possible to affirm that a nucleotide does indeed exist at this position. This position is noted N. When sequencing cannot define the exact nucleotide, but it is possible to affirm that it is a purine, then a P is added, and if it is a pyrimidine, then a Y is added. This results, in the majority of cases, in an impossibility to define the amino acid which will be found at this position, which on the protein alignment will be noted as a "gap". It will therefore no longer be possible to differentiate between a difficulty in determining the nitrogenous base during sequencing and the fact that sequencing has not been done.Thus, from the nucleotide sequences, we can calculate the number of poorly defined amino acids. Let ^ be the mean and ^ the standard deviation of the number of mutations calculated per sequence of the sample studied, in comparison with a reference sequence. Knowing that the number of aberrant NYP is defined as ^^^. ^^ >= ^ + 3 × ^ and that ^^^ ^^ ^^^^^ corresponds to the number of sequences having a number of NYP greater than or equal to ^^^ ^^ . We can then calculate: ^ ^^^^^ = 1 − (^^^ ^^ ^^^^^ ÷ ^^^ ^^ ) The closer S3 is to 1, the greater the number of aberrant amino acids. S4: ^ ^^^^^ : Sequences with an aberrant number of gaps. Let ^ be the mean and ^ the standard deviation of the number of mutations calculated per sequence of the sample studied, in comparison with a reference sequence. Knowing that the number of aberrant gaps is defined as ^^^ ^^ >= ^ + 3 × ^ and that ^^^ ^^ ^^^^^corresponds to the number of sequences having a gap number greater than or equal to ^^^ ^^ . We can then calculate: ^ ^^^^^ = 1 − (^^^ ^^ ^^^^^ ÷ ^^^ ^^ ) The closer the S4 score is to 1, the more aberrant the number of gaps is. F(S1QS): Impact of sequence data on the quality of statistical prediction. This function varies from 0 to 1. ^^^ ^^^ ^ allows you to assess whether the input data is sufficient and robust in the face of the prediction to be made: S5: ^ ^^^^^ : 3D structure having an aberrant number of missing data. Consider, ^ ^^ ^ ^^ ^ ^as the total number of files studied and ^ the mean and ^ the standard deviation of the number of amino acids per file of the sample studied, in comparison to a reference file. This reference file can be chosen in two ways. Either this file is defined as by the scientific community, or because it represents the root of a phylogenetic tree containing all or the majority of the proteins studied Knowing that the number of missing amino acids (AA) aberrant is defined as: ^^ ^^ = ^ + 3 × ^ And that ^ ^^ ^ ^^ ^ ^ ^^^^ corresponds to the number of files with a number of missing AAs greater than or equal to ^^^ ^^ . We can then calculate: We note that the choice of a particular reference structure has little impact since its total number of residues will be very close to another described structure not having an aberrant number of uninformed positions. S6: ^ ^^^^^^^^: impact of the existence of different subunits on the final result. If the protein studied has several subunits (each noted as subunit), then they will be listed, regardless of the number of sequences in the alignment. If it only has one subunit, then this score will be equal to 1. S7: ^ ^^^^ ^ : Importance of the number of sequences (theoretical threshold at 1000). ^^^ ^^ is the total number of sequences in the sample excluding sequences with an aberrant number of mutations ∀ ^^^ ^^ ∈ ℕ, ^^ (^^^ ^^ − ^^^^ ^^ ^^^^^ ∪ ^^^ ^^ ^^^^^ ∪ ^^^ ^^ ^^^^^ ^) > 1000, ^^ (^^^ ^^ − ^^^^ ^^ ^^^^^ ∪ ^^^ ^^ ^^^^^ ∪ ^^^ ^^ ^^^^^ ^) < 1000, S8: ^ ^^^^^^: evaluates the impact of gaps in the alignment. Consider ^ the mean and ^ the standard deviation of the number of gaps calculated per sequence of the sample studied after alignment, in comparison to the same sequence before alignment. ^ ^^ ^ < ^ ^^^^^ ^ ^^^^^^ = 1 − ^ ^^ ^ > ^ ^^^^^ ^ ^^^^^^ = 0 S9 : ^ ^^^^ and ^ ^^^^ : impact of 5' and 3' end variance. 5 amino acids (AA) from the 5' end and 5 AA from the 3' end are studied. Invariant residues (inv), synthetic lethal (SL) and compensatory mutation (CM) pairs are listed at the ends. to ^^^^^ ^^^ ^ ^^^^ ^^^^ > 1 ^^^^ ^^ ^^^ ^ ^^^^ = 1 S10 : ^ ^^^: Redundancy of the sequence batch. Consider ^ the mean and ^ the standard deviation of the histogram of sequence redundancy. The polypeptide sequences corresponding to the candidate protein may be very heterogeneous or not very heterogeneous in terms of the mutations they carry. An extreme situation may be that out of X sequences, they all have the same AA sequence. We would then assume that this batch is very homogeneous, that the redundancy is absolute, and therefore would have a score equal to 0. Conversely, when the polypeptide sequences are very different taken two by two, then there is very little redundancy, and the score tends towards 1. ^ ^^ ^ < ^ ^^^^^ ^ ^^^ = ^ ^^ ^ > ^ ^^^^^ ^ ^^^ = 1 If S10 = 1 absolute heterogeneity, if S10= 0 no heterogeneity S11: ^ ^^^ : impact of hypervariable regions. Hypervariable regions (hyp) are listed from the bibliography and counted, regardless of the number of sequences in the alignment. ^^^^ = 1 ÷ ^ ℎ^^^^^^ S12 : ^ ^^^^^ : impact of deletions and insertions. Insertions-deletions (indels) are listed from the bibliography and counted, regardless of the number of sequences in the alignment. S13: ^ ^^^^^^^ : impact of post-translational modifications. The different types of post-translational modifications (postrad) are listed from the bibliography and counted, regardless of the number of sequences in the alignment. S14: ^ ^^^^^^^ : impact of the existence of different subtypes. The different subtypes (each noted as "subtype") belonging to the batch of sequences studied are listed, regardless of the number of sequences in the alignment. Indeed, some microorganisms mutate so quickly that the evolution over time of these different variants allows the appearance of variant subtypes. ^ ^^^^^^^ = 1 ÷ ^ ^^^^^^^^^^^^^^^^^ ^(^ ^): quality of alignment of primary sequences + ^ ^^^ + (^ ^^^ + ^ ^^^^^^^ + ^ ^^^^^^^ ) ÷ 3 × 2) ÷ 17 ^^^ ^^^ ^ : quality of the statistical prediction of the alignment ^^^ ^^^ ^: Impact of alignment on target prediction S18: ^ ^^^ : Mutational richness of the sequences. Let us consider ^ the mean and ^ the standard deviation of the number of mutations calculated per sequence of the sample of sequences having a number of mutations < ^^^ ^^ , compared to a reference sequence. ^^ ^ > ^ ^^^^^ ^ ^^^ = 1 If S18 = 1, high heterogeneity, if S18= 0 no heterogeneity. S19: ^ ^^^: impact of invariant positions. We define the percentage of invariant residues (Target Quality QC). Invariant residues are identified as follows: the same amino acid is found at a given position in at least 99.7% of the sequences. The sum of these positions gives ^^^^ ^ ^^^ = ^^^^ ÷ ^ ^^^ ^^ ^ ^^^ ^^^ ^^ ^^^^^^^^^ ^^ ^′^^^^^^^^^^ ^ ( ^^^^^^^^^^^ ) : degree of invariance This score function allows to evaluate the true degree of invariance of the batch of sequences studied, thus giving an image of its essentiality. Indeed, the invariant positions are such that they cannot be selected in the event of mutation since they are essential to the replicability of the organism from which the candidate protein comes. S20: ^ ^^^^^: impact of the number of covariant positions. We define the percentage of covariant residues as follows: sum of the positions found in a pair of covariants (^^^^^ ^^ ^^ ) having a ^ ^ defined according to the Noirvit protocol (Noivirt, et al., 2005). ^ ^^^^^ = ^^^^^ ^^ ^^ ÷ ^ ^^^ ^^ ^ ^^^ ^^^ ^^ ^^^^^^^^^ ^^ ^ ^ ^^^^^^^^^^ S21: ^ ^^^^^^^^ : impact of the number of pairs of covariant positions. We define the percentage of pairs of covariant residues as follows: sum of the pairs of residues having a ^ ^ defined according to the Noirvit protocol ^^^^^ ^^ ^^^ ^^ ^^ S24: ^ ^^^^ : impact of global non-synonymous mutability. It is calculated as follows: ^ ^ ^^^^ ^^ ^^ ^ ^^^é^^^^ ^^^^^^é ^^^^ Lao et al. where A denotes a residue of the studied sequence which is non-synonymous with the chosen reference sequence and where S denotes a residue of the studied sequence which is synonymous with the chosen reference sequence. S25: ^ ^^^^ : impact of synonymous mutability. It is calculated as follows: ^ ^ ^^^^ ^^ ^^ ^ ^^^ é^^^^ ^^^^^^^^^^ Lao et al. S22: ^ ^^ : impact of the number of SLs and their strength in the overall result. This score is the sum of the negative dissimilarity coefficients ^ therefore of the pairs of SL residues reported to the sum of all the ^ ^ (SL plus CM) S23: ^ ^^ : impact of the number of CMs and their strength in the overall result. This score is the sum of the dissimilarity coefficients ^ positive therefore of the pairs of CM residues reported to the sum of all the ^ ^ (SL plus CM) S27: ^ ^^^^^^^^^: Graph complexity. We specify that a graph G is defined by a pair (S,A) with S a finite set of vertices, and A a finite set of pairs of vertices (si, sj) in S². A pair is therefore a pair of vertices connected by an edge. It is a question of studying the invariant residues and SL involved in several pair relations (close in space and located on the surface of the protein, being either invariant or involved in a synthetic lethal relation). From the graph of links, we define this score as being the average of the number of edges (a) per node (^)̿ S28: ^ ^^^^^^ : Number of connected graphs. We specify that a graph is connected if each pair of vertices is connected by an edge. S28: ^ ^^^^^: residues close in space (5 angstroms). This score corresponds to the proportion of residue pairs located less than 5 angstroms from each other (^^). From a reference pdb file, the calculation of the distance between the 2 ^^ of a residue pair is carried out for all residue pairs. The number of pairs separated by less than 5 angstroms is evaluated and noted ^ ^ ^^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ . S29: ^ ^^^^^^ : residues close in space (10 angstroms). This score corresponds to the proportion of residue pairs located less than 10 angstroms from each other (^^). From a reference pdb file, the calculation of the distance between the 2 ^^ of a residue pair is carried out for all residue pairs. The number of pairs separated by less than 10 angstroms is noted ^ ^^ ^ ^^ ^ ^^ ^^ ^ ^ ^ ^ ^ ^ . S30: ^^^^^^^^^^^: invariants and SL close in space (5 angstroms) From a reference pdb file, the calculation of the distance between the 2 ^^ of a pair of residues is carried out for all pairs of residues. The number of pairs being separated by less than 5 angstroms is evaluated and noted ^ ^ ^ ^ ^^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ . S31: ^^^^^^^^^^^^: invariant and SL close in space (10 angstroms) From a reference pdb file, the calculation of the distance between the 2 ^^ of a pair of residues is carried out for all pairs of residues. The number of pairs separated by less than 10 angstroms is noted ^ ^^ ^ ^^^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ ^ ^ . S32: ^ ^^^ : Percentage of accessible residues. Where ^^^^^ represents the number of accessible residues and ^ ^ ^ ^ ^^ ^ ^ the number of buried residues can be determined by ASA software (Alland et al., 2005) S33: ^ ^^^^^^^^ : Percentage of accessible invariant and SL residues. Or ^^^^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ represents the number of accessible residues and ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^^ ^ ^ ^ ^ ^ ^ the number of buried residues. They are determined by the ASA software (Alland et al., 2005) ^ ( ^ ^^^^^^ ) : chemical reactivity With ^ ^^^^^^^^^ the percentage of chemical bonds, ^ ( ^ ^^^^ ) : Invariance score for each target We define two sets A and B: - Set A groups the connected or very weakly connected graphs (each of these sub-graphs represents a subset of set A). Let us recall the definition of an edge connecting two residues: the two residues are less than 10 angstroms apart, they are present on the surface of the molecule and are either invariant or are part of a SL pair. - Set B groups the targets defined by the F-pocket software (each target defines a subset of set B). It is then a question of linking a subset of set A to a subset of set B. Once these subsets are linked two by two (one from set A to one from set B), we define the intersection of their element (^^^^^ ^ ) as well as their union (^^^^^ ^ ). more ^^^ ^^^ ^ ^ ^ ^ ^ ÷ ( ^^^ ^^^^^^ ^ ^ ^ ) is close to 1, the more total the invariance of the target x is and therefore the less it will be subject to variability in the future. Thus such a target will drastically reduce the appearance of new drug-resistant variants. Indeed, for these variants to exist, positions of the target would have to be able to mutate. However, such a mutated target would no longer allow the replicability of the variant and would therefore call into question its existence. ^(^ ^ ): score between 0 and 1, for each target ÷ (^^^^^ ^ × ( ^^^^^ ^ − 1 )) ÷ 2))) ÷ 4 This scoring function allows us to give a score to each of the targets determined by our software and defines the "druggability" of this target. "Druggability" means being able to effectively bind a small molecule that would have therapeutic capabilities and therefore represent a future drug. To this definition is added that this target would leave little or no possibility of therapeutic escape. Thus a small molecule, a future drug, must be able to chemically bind to a group of residues (^ ^^^^^^^^^ ) located in a concave space (called a pocket), the volume of which can be determined (^^^ ^ ), and having the greatest possible invariance (^ ( ^ ^^^^ ) . ^(^ ^ ): quality of three-dimensional alignment: ^(^ ^ ) = ^^ ^^^^^ × 8 + ^ ^^^^^^^^^ × 5 + ^^^ ^^^ ^: quality of statistical prediction of structural alignment: ^^^^^^ ^ = on target prediction: EXAMPLE The method of the invention was implemented on the hemagglutinin (HA) protein of the influenza virus. This implementation highlighted several advantageous chemical bonds which made it possible to select 7 pairs / groups of amino acids linked by significant chemical and genetic bonds, the distance of which is acceptable (Figure 2, see the grayed-out amino acids). These particularly advantageous pairs / groups were then taken into account in the analysis of the therapeutic target regions of the HA protein identified during the implementation of steps (E1) and (E2) of the method of the invention. As shown in Figure 3 (using the information from Figure 3 of Lao et al.), taking into account these 7 preferential amino acid groups restricts the number of therapeutic target regions within HA from 6 to 3.These 3 therapeutic targets will therefore be the subject of pharmaceutical drug screening with an increased probability of identifying reliable and efficient active ingredients. The method of the invention therefore makes it possible to refine the method previously proposed in Lao et al. by adding a step (E3) which requires determining all the chemical interactions present between each residue and / or between each pair of residues within the candidate protein separated by a distance of at most 5 Angstroms (see Figure 2). Thanks to this additional step, only the pairs / groups of residues possessing advantageous physicochemical characteristics and located at a distance such that the chemical bonds between residues potentially have an influence (E4), will be selected. Thus, the “chemical interactivity” information proposed in the present invention reinforces the validity of the targets identified by the genetic approach of Lao et al.by selecting the most relevant ones, so as to significantly reduce the number of therapeutic targets to be used in molecular library screening programs and increase the probability of identifying more effective drugs more quickly. BIBLIOGRAPHICAL REFERENCES Alland et al., 2005 Alland et al.: Alland C, Moreews F, Boens D, Carpentier M, Chiusa S, Lonquety M, Renault N, Wong Y, Cantalloube H, Chomilier J, et al. 2005. RPBS: a web resource for structural bioinformatics. Nucleic Acids Res 33:W44-49 Barlow, DJ & Thornton, JM Ion-pairs in proteins. Journal of Molecular Biology 168, 867–885 (1983) Baumgart, M. et al. Design of buried charged networks in artificial proteins. Nat Commun 12, 1895 (2021). Brouillet S., Valere T., Ollivier E., Marsan L., Vanet A., Co-lethality studied as an asset against viral drug escape: the HIV protease case. Biol Direct. 2010 Jun 17;5:40. Byrd-Leotis, L., Galloway, SE, Agbogu, E. & Steinhauer, DAInfluenza hemagglutinin (HA) stem region mutations that stabilize or destabilize the structure of multiple HA subtypes. J Virol 89, 4504–4516 (2015). Cheng C, Xiao P. Evaluation of the correctable decoding sequencing as a new powerful strategy for DNA sequencing. Life Sci Alliance. 2022 Apr 14;5(8):e202101294. Childers, M. C., Towse, C.-L. & Daggett, V. The effect of chirality and steric hindrance on intrinsic backbone conformational propensities: tools for protein design. Protein Eng Des Sel 29, 271–280 (2016). Di Russo, N. V., Estrin, D. A., Martí, M. A. & Roitberg, A. E. pH-Dependent conformational changes in proteins and their effect on experimental pK(a)s: the case of Nitrophorin 4. PLoS Comput Biol 8, e1002761 (2012). Donald, J. E., Kulp, D. W. & DeGrado, W. F. Salt bridges: geometrically specific, designable interactions. Proteins 79, 898–915 (2011). Dyson, H. J., Wright, P. E. & Scheraga, H. A.The role of hydrophobic interactions in initiation and propagation of protein folding. Proceedings of the National Academy of Sciences 103, 13057–13061 (2006). Fitch, C. A., Platzer, G., Okon, M., Garcia-Moreno E, B. & McIntosh, L. P. Arginine: Its pKa value revisited. Protein Sci 24, 752–761 (2015). Freitas, R. F. de & Schapira, M. A systematic analysis of atomic protein–ligand interactions in the PDB. Med. Chem. Commun. 8, 1970–1981 (2017). Galloway, S. E., Reed, M. L., Russell, C. J. & Steinhauer, D. A. Influenza HA subtypes demonstrate divergent phenotypes for cleavage activation and pH of fusion: implications for host range and adaptation. PLoS Pathog 9, e1003151 (2013). Harms, M. J. et al. The pKa Values of Acidic and Basic Residues Buried at the Same Internal Location in a Protein Are Governed by Different Factors. Journal of Molecular Biology 389, 34–47 (2009). Harris, T. K. & Turner, G. J. Structural Basis of Perturbed pKa Values of Catalytic Groups in Enzyme Active Sites.IUBMB Life 53, 85–98 (2002). Harrison, J. S. et al. Role of Electrostatic Repulsion in Controlling pH-Dependent Conformational Changes of Viral Fusion Proteins. Structure 21, 1085–1096 (2013). Hubbard, R. E. & Kamran Haider, M. Hydrogen Bonds in Proteins: Role and Strength. in eLS (John Wiley & Sons, Ltd, 2010). Kampmann, T., Mueller, D. S., Mark, A. E., Young, P. R. & Kobe, B. The Role of Histidine Residues in Low-pH-Mediated Viral Membrane Fusion. Structure 14, 1481– 1487 (2006). Kuiken H.J.,, Beijersbergen R.L., Exploration of synthetic lethal interactions as cancer drug targets. Future Oncol. 2010 Nov;6(11):1789-802 Lao J. and Vanet A. A New Strategy to Reduce Influenza Escape: Detecting Therapeutic Targets Constituted of Invariance Groups. Viruses. 2017 Mar 2;9(3):38. Mellman, I., Fuchs, R. & Helenius, A. Acidification of the endocytic and exocytic pathways. Annu. Rev. Biochem. 55, 663–700 (1986). Meuzelaar, H., Vreede, J. & Woutersen, S.Influence of Glu / Arg, Asp / Arg, and Glu / Lys Salt Bridges on α -Helical Stability and Folding Kinetics. Biophysical Journal 110, 2328– 2341 (2016). Le Guilloux, V.; Schmidtke, P.; Tuffery, P. Fpocket: An open source platform for ligand pocket detection. BMC Bioinform. 2009, 10, 168 Noivirt, O.; Eisenstein,M.; Horovitz, A. Detection and reduction of evolutionary noise in correlated mutation analysis. Protein Eng. Des. Sel. 2005, 18, 247–253 Onofrio, A. et al. Distance-dependent hydrophobic–hydrophobic contacts in protein folding simulations. Phys. Chem. Chem. Phys. 16, 18907–18917 (2014). Pahari, S., Sun, L. & Alexov, E. PKAD: a database of experimentally measured pKa values of ionizable groups in proteins. Database 2019, baz024 (2019). Petitjean M., Badel A., Veitia R.A., Vanet A., Synthetic lethals in HIV: ways to avoid drug resistance : Running title: Preventing HIV resistance. Biol Direct. 2015 Apr 17;10:17 Sticke, D. F., Presta, L. G., Dill, K. A. & Rose, G. D.Hydrogen bonding in globular proteins. Journal of Molecular Biology 226, 1143–1159 (1992).

Claims

CLAIMS 1. A computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) Identification (E1), in a set of previously aligned nucleotide and polypeptide sequences characteristic of said candidate protein, of invariant residues and / or pairs of synthetic lethal residues called target residues; b) Identification (E2) of at least one candidate region consisting of at least one pair of target residues identified in step a), said at least one pair comprising target residues located at a determined distance in space and being exposed on the surface of the candidate protein, preferably in a pocket; c) Determination (E3) of the advantageous chemical interactions between each residue and / or between each pair of residues within the candidate protein,from the 2D and / or 3D structure of the candidate protein; said advantageous chemical interactions being hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or bonds with negative repulsions and / or with positive repulsions; d) Selection (E4) of the residues linked by said advantageous chemical interactions, said residues being at a distance of at most 10 Angstroms; e) Selection (E5) of at least one therapeutic target region from among the candidate regions identified in step b), comprising the residues selected in step d).

2. Method according to claim 1, wherein the step of identifying (E3) the advantageous chemical interactions consists of a determination of the chemical interactivity of all the residues within the protein and / or a determination of the chemical interactivity of all the pairs of residues within the protein,said residues or pairs of residues being spaced by a determined distance of at most 10 Angstroms.

3. Method according to any one of claims 1 to 2, wherein the distance between the invariant residues or between the pairs of synthetic lethals determined in step b) is between 2 and 8 Angstroms, preferably 5 Angstroms.

4. Method according to any one of claims 1 to 3, comprising an evaluation (E3') of the chemical interactivity of the determined bonds (E3) comprising, a calculation of at least one score from: a percentage of hydrophobic bonds and / or a percentage of hydrogen bonds and / or, a percentage of salt bridges and / or a percentage of bonds with negative repulsions and / or a percentage of bonds with positive repulsions and / or a percentage of chemical bonds.

5. Method according to any one of the preceding claims, wherein the target region is a pocket identified from a set of aligned 3D structures of the candidate protein and wherein the chemical interactions are determined on a set of aligned 3D structures of the candidate protein.

6. Method according to claim 5, wherein the target pocket comprises at least five target residues chosen from invariant residues or synthetic lethal covariant residues, and wherein said pocket has a volume between 60 ^̇ ^ and 500 ^̇ ^7. Method according to any one of the preceding claims comprising an evaluation of the quality of the spatial location of the target residues, said evaluation being characterized by the calculation of the proportion of pairs of residues located less than 5 Angstroms from each other and / or the proportion of pairs of residues located less than 10 Angstroms from each other.

8. Method according to any one of the preceding claims, comprising a step of selecting (E0) the reference polypeptide sequence of the candidate protein, then a step of selecting the polypeptide sequences to be tested having lengths identical to the reference sequence, and / or having a number of mutations less than or equal to three standard deviations of the average number of mutations per sequence, and / or having an invariant residue at its N or C end. 9.Method according to claim 8, comprising an evaluation (E11') of the quality of the selected polypeptide sequences, said evaluation comprising the calculation of at least one score measuring the heterogeneity of the length of the selected polypeptide sequences and / or the number of selected polypeptide sequences having an aberrant number of poorly defined amino acids and / or the number of sequences having an aberrant number of missing residues compared to the reference polypeptide sequence of the candidate protein.

10. A computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) Determining (E3) the advantageous chemical interactions between each residue and / or between each pair of residues within the candidate protein, from the 2D and / or 3D structure of the candidate protein; said advantageous chemical interactions being hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or bonds with negative repulsions and / or positive repulsions; and b) Selecting the residues (E4) linked by said advantageous chemical interactions, said residues being at a distance of at most 10 Angstroms, said residues forming a therapeutic target region.