Bioinformatics methods for determining therapeutic target regions

By integrating 3D structure and chemical interactions, the method enhances the precision of therapeutic target region identification, reducing the number of potential targets and improving the efficiency of drug development by focusing on stable and chemically strong regions.

JP2025536116APending Publication Date: 2025-10-31CENT NAT DE LA RECH SCI (C N R S) +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517544
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-20
Filing Date
2023-09-19
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods for identifying therapeutic target regions for antiviral drugs, such as those proposed by Lao et al., lack precision, fail to account for biochemical interactions, and result in large numbers of potential targets, leading to inefficient and costly screening programs.

Method used

A computer-implemented method that incorporates 3D structure and intramolecular chemical interactions to validate residues, filters out low-quality sequences, and selects residues based on favorable chemical bonds and spatial proximity, reducing the number of potential targets to those that are stable and chemically strong.

Benefits of technology

The method provides a more reliable and efficient identification of therapeutic target regions, ensuring that the selected targets are statistically and biochemically valid, enabling faster and more effective drug design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536116000001_ABST
    Figure 2025536116000001_ABST
Patent Text Reader

Abstract

The present invention provides a computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, the method comprising: a) identifying pairs of invariant and / or lethal synthetic residues, called target residues, in a set of pre-aligned nucleotide and polypeptide sequences characteristic of a candidate protein (E1); b) identifying at least one candidate region (E2) consisting of at least one pair of target residues identified in step a), wherein at least one pair is located at a determined distance in space and comprises target residues exposed on the surface of the candidate protein, preferably in a pocket; c) determining (E3) from the 2D and / or 3D structure of the candidate protein favorable chemical interactions between each residue and / or each pair of residues in the candidate protein, wherein the favorable chemical interactions are hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or negative repulsive bonds and / or positive repulsive bonds; d) selecting residues that are linked by favorable chemical interactions, the residues being at a distance of at most 10 Angstroms (E4); e) selecting from among the candidate regions identified in step b) at least one therapeutic target region comprising the residue selected in step d) (E5).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to bioinformatics methods for identifying reliable and durable therapeutic target regions to optimize the search for new drugs, particularly antiviral drugs. [Background technology]

[0002] Despite the existence of treatments, RNA viruses remain a serious public health problem. Their high mutation rate allows them to rapidly acquire resistance to these treatments. To prevent the emergence of resistance, it is recommended to preferentially target invariant amino acids. In fact, mutations at highly conserved positions can lead to deterioration or alteration of biological functions, rendering the virus inviable. However, due to their extremely low number, invariant positions alone cannot constitute a drug binding site.

[0003] To find other optimal drug-accessible binding sites, Lao J. et al. also proposed identifying pairs of co-mutations called "synthetic lethals" (SL). [Brouillet et al., Petitjean et al.]. SLs represent mutations that are not lethal but render the virus nonviable when combined. These SLs have already been investigated in the search for anti-cancer drugs [Kuiken HJ and Beijersbergen RL] and anti-HIV drugs. Lao et al. propose a series of computational steps to identify the best target residues that, when mutated or blocked by drugs, can substantially affect the biological function of the targeted pathogen.

[0004] However, the method proposed by Lao et al. has drawbacks. First, the targets identified by this method are not described with sufficient precision, which puts users at a disadvantage in their development programs. For example, this method does not reveal whether the targets proposed by the software have a large number of invariant residues, a specific volume, or a low likelihood of future mutations. Furthermore, it does not take into account the fact that the initial batch of sequences may not be very useful or, conversely, may be very reliable and therefore highly predictive. Finally, this method lacks information about the nature of chemical bonds that may exist between residues of candidate proteins, more specifically in the target region, which is important information for establishing the relevance of the identified targets. This biochemical information would strengthen the validity of the targets identified by Lao et al.'s genetic approach and allow the selection of the most relevant ones.

[0005] Typically, the number of therapeutic target regions identified by genetic screening as described by Lao et al. is too large. Therefore, research institutions must then set up long and expensive screening programs for libraries of molecules (antibodies, small molecules, etc.). Therefore, it is preferable to reduce the number of target regions of interest as much as possible and select only target regions that are likely to be long-lasting and stable, especially by taking into account existing biochemical links within the protein of interest. Summary of the Invention

[0006] According to a first aspect, the present invention aims to improve the method of Lao et al. To achieve this, the inventors propose to incorporate into this method steps based on the 3D structure and intramolecular chemical interactions of the target protein. These steps make it possible to confirm or validate the residues identified by the method of Lao et al., depending on the physicochemical properties of their interactions and their sensitivity to environmental parameters (pH, temperature, etc.). These steps allow the identification of residues with favorable physicochemical features and potential chemical bonds between residues. Only pairs of residues located at such a distance as to have a significant effect are selected.

[0007] In addition, we propose applying a filter to the Lao et al. method to exclude nucleotide or peptide sequences of insufficient quality from the analysis, thus avoiding burdening the system by working with unusable or poorly indexed sequences. Therefore, we developed a further step for more precise selection of sequences to be tested and / or initially used.

[0008] Finally, this type of method must be able to rapidly generate results, whether for proteins from different microorganisms or for proteins from different mutants of the same microorganism. It is therefore important to be able to have a fast and safe method for processing several proteins in parallel, which allows taking into account the 3D structure of proteins, especially from known mutants.

[0009] All of these developments make it possible to improve the reliability and interest of the method described in Lao et al., which was theoretical but difficult to use (because it was poorly explained and poorly documented). The method of the present invention is more effective and reliable than the method described in Lao et al. in that it has been scientifically enhanced and supplemented by new criteria that allow the detected targets to be described in detail both statistically and biochemically. These improvements make the method of the present invention essential for identifying important target regions in pathogenic organisms, allowing researchers to design future drugs.

[0010] To this end, a computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, the method comprising: a) identifying pairs of invariant and / or synthetic lethal residues, referred to as target residues, in a set of pre-aligned nucleotide and polypeptide sequences characteristic of candidate proteins; b) identifying at least one candidate region consisting of at least one pair of target residues identified in step a), wherein at least one pair is located at a determined distance in space and comprises target residues exposed on the surface of the candidate protein, preferably within a pocket; c) determining from the 2D and / or 3D structure of the candidate protein favorable chemical interactions between each residue and / or each pair of residues in the candidate protein, wherein the favorable chemical interactions are hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or negative repulsive bonds and / or positive repulsive bonds; d) selecting residues that are linked by favorable chemical interactions, the residues being at most 10 angstroms apart; e) selecting from among the candidate regions identified in step b) at least one therapeutic target region that comprises the residue selected in step d).

[0011] According to a first aspect, the present invention is advantageously accomplished by incorporating the following features, taken alone or in any of the technically possible combinations: - the step of identifying favorable chemical interactions consists of determining the chemical interactions of each residue in the protein and / or determining the chemical interactions of each pair of residues separated by the determined distance; - the distance determined in step b) is between 2 and 8 angstroms, preferably 5 angstroms; The method may include calculating the percentage of hydrophobic bonds and / or the percentage of hydrogen bonds and / or the percentage of salt bridges and / or the percentage of negative repulsive bonds and / or the percentage of positive repulsive bonds. and / or evaluating the determined chemical interactions of the bonds, including calculating at least one score from the percentage of chemical bonds; - the target region is a pocket identified from a set of aligned 3D structures of the candidate protein, and chemical interactions are determined for the set of aligned 3D structures of the candidate protein; - the target pocket comprises at least five target residues selected from synthetic lethal invariant residues or covariant residues, and the pocket

[0012]

number

[0013] According to a second aspect, the present invention aims to identify therapeutic target regions based solely on the chemical interactions that exist between residues in the region. To this end, a computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism is provided, the method comprising: a) determining, from the 2D and / or 3D structure of the candidate protein, favorable chemical interactions between each residue and / or each pair of residues in the candidate protein, wherein the favorable chemical interactions are hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or negative repulsive bonds and / or positive repulsive bonds; b) selecting residues that are linked by favorable chemical interactions, the residues being at most 10 angstroms apart, and the residues forming a therapeutic target region. [Brief explanation of the drawings]

[0014] [Figure 1] Further features, objects and advantages of the present invention will become apparent from the following description, which is purely illustrative and non-limiting, and should be read in conjunction with Figure 1, which shows a diagram of the steps of a computer-implemented method for determining at least one therapeutic target region on the surface of a candidate protein of a pathogenic organism. [Figure 2-1] 1 shows a diagram of the chemical interactions that exist between pairs of amino acids in the HA protein of influenza virus (HB = hydrogen bond, SB = salt bridge, PR = positive repulsion, HI = hydrophobic interaction). The highlighted bonds are those that have already been defined by the techniques described in Lao et al. and therefore correspond to pairs identified by the chemical methods of the present invention, connecting pairs or groups of amino acids with a favorable distance (highlighted in grey), and therefore these pairs will be selected as part of the methods of the present invention. [Figure 2-2] 1 shows a diagram of the chemical interactions that exist between pairs of amino acids in the HA protein of influenza virus (HB = hydrogen bond, SB = salt bridge, PR = positive repulsion, HI = hydrophobic interaction). The highlighted bonds are those that have already been defined by the techniques described in Lao et al. and therefore correspond to pairs identified by the chemical methods of the present invention, connecting pairs or groups of amino acids with a favorable distance (highlighted in grey), and therefore these pairs will be selected as part of the methods of the present invention. [Figure 3-1] Selected amino acids of interest are highlighted in Figure 2 among the amino acids identified after performing steps E1 and E2 of the method of the invention (see Figure 3 of Lao et al. for further details). [Figure 3-2]Selected amino acids of interest are highlighted in Figure 2 among the amino acids identified after performing steps E1 and E2 of the method of the invention (see Figure 3 of Lao et al. for further details). [Figure 3-3] Selected amino acids of interest are highlighted in Figure 2 among the amino acids identified after performing steps E1 and E2 of the method of the invention (see Figure 3 of Lao et al. for further details). DETAILED DESCRIPTION OF THE INVENTION

[0015] The present invention relates to a computer-implemented method for determining at least one therapeutic target region on the surface of a candidate protein of a pathogenic organism. Such a method may be implemented by a processing unit, such as, for example, one or more processors or any other equivalent means.

[0016] The term "target pathogenic organism" is used herein to refer to any type of organism that can cause disease in a "host" (such as a human, plant, or animal). The organism is preferably a microorganism, such as a virus, bacterium, parasite, fungus, protozoan (amoeba, sporozoa, or flagellate). Some of these organisms have been sequenced in the last few decades, e.g., · Rotavirus, Calicivirus (Norwalk and Hepatitis E), Ebola, Chikungunya, SARS coronavirus; Legionella, Campylobacter, Helicobacter, Mycobacterium, and Escherichia coli O157:H7 strains; Sporozoa (parasitic protozoa) such as Coccidiophores (Cryptosporidium, Cyclospora, Toxoplasma) or Microsporidiophores (Enterocytozoon, Encephalitozoon, Nosema).

[0017] It may also be a harmful macroorganism, such as, for example, a worm (belonging for example to the group of flatworms, such as helminths, flukes or cestodes, or nematodes, such as Ascaris, Toxocara or Trichuris) or an insect.

[0018] In a broad sense, the term "target pathogenic organism" also includes tumor cells or cells infected with pathogenic microorganisms such as viruses, bacteria, etc. Indeed, these cells often express proteins on their surface or in their cytoplasm (or even in their nucleus) that are involved in maintaining growth signals in these cells, resisting cell death, evading the immune system, activating angiogenesis, invasion and metastasis, replicative immortality, evading growth factor inhibitor genes, reprogramming energy metabolism, and thereby amplifying the disease (cancer or infectious disease). Therefore, it is entirely reasonable to use the methods of the present invention to identify target regions suitable for therapeutic studies in proteins involved in these disorders as well.

[0019] Pathogenic organisms are essentially composed of proteins, some of which are necessary for their development, infectivity, and / or pathogenicity. For every known pathogen, a large number of proteins of this type have been identified and are a major target for researchers in the pharmaceutical industry. Indeed, disabling the function of such proteins and / or masking such proteins often makes it possible to stop the development, proliferation, and / or harmful effects of the pathogen on human, plant, or animal health. In the context of the present application, these proteins are used to refer to proteins that are essential for the development, infection, and / or pathogenicity of the pathogenic organism of interest. They will be referred to as "candidate proteins" in that they have already been identified as candidates for potential effects on infectivity and / or pathogenicity.

[0020] For a molecule to be selected as a "drug," it must have a substantial short-, medium-, or long-term effect on the function of at least one of these candidate proteins. To do this, it must first be able to contact the candidate protein (preferably through several contact zones). 3D structures now allow us to determine which zones are localized on the protein surface and, therefore, which zones could be potential "contact zones" between a drug molecule and the protein of interest. However, for each protein, these contact zones are too numerous to screen in all research programs. To facilitate their study and accelerate the identification of effective drugs, researchers need to know more precisely which "target regions" on the surface of candidate proteins have the most promising chemical and biological properties to be targeted by the drug. Therefore, in this application, a "therapeutic target region" is defined as a zone of favorable contact between a drug and a candidate protein, selected so that the interaction between these two elements (the drug on the one hand and the protein on the other) is chemically strong and stable and can significantly affect the biological function of the protein.

[0021] Brief Description of the Method Steps of the Invention It is assumed that candidate proteins known to affect the development, infectivity, and / or virulence of the pathogenic organism under study have already been identified. To perform the methods of the present invention, the candidate proteins must be known and well-described in the literature. Notably, the polypeptide and 3D sequences must already be characterized and readily available in the art. The nucleotide sequences encoding these polypeptide sequences must also be known. All these sequences are generally provided in public, openly accessible databases. The sequences in these databases, hereafter referred to as "BDD1," are generally anonymized, freely accessible, and uploaded by various international research organizations. Examples include Genbank, NCBI, INSDC, EMBL, HIVDB, LANL, FLUDB, GISAID, etc., all of which are well known to those skilled in the art.

[0022] In the previous step of the method of the present invention, a polypeptide sequence and a reference nucleotide sequence must be selected (step E0). These reference sequences can be selected from a phylogenetic tree (in this case, the reference sequence is at the root of the phylogenetic tree representing the sequence being studied), or, if no phylogenetic tree exists, by recalculating an ancestral sequence reconstructed by one of three methods: parsimony, maximum likelihood, or Bayesian. This reference sequence can also be a consensus sequence from the batch of sequences being studied; however, in this case, the calculated consensus sequence may be a sequence that never existed and is therefore not necessarily functional, and is therefore only useful for comparison with other sequences.

[0023] Subsequently, other known / recorded nucleotide and polypeptide sequences for this candidate protein are identified in the "BDD1" database. These additional sequences are hereafter referred to as "initial sequences."

[0024] Proteins must adopt a specific structure in space and perform a specific number of functions that can only be achieved if they have the right chemical groups in the right position. This is called the structure-function link. Therefore, certain mutations change the structure of the protein and therefore cause it to lose one or more functions. If these functions are essential, the virus becomes non-replicating and therefore can no longer develop. These mutations are mainly of three types: invariant positions, synthetic lethal pairs, and compensatory pairs. A single mutated invariant position can cause the protein to In the case of synthetic lethal pairings, two positions must be mutated to achieve the same goal, rendering the protein nonfunctional. Finally, paired compensatory mutations are defined as those in which the first mutation renders the protein nonfunctional, but the second mutation appears and restores the function of the studied protein. In these three cases, a high level of stress is imposed on the protein.

[0025] In the present invention, these initial sequences are first processed to select pairs of invariant and / or synthetic lethal residues (Step E1). Once these pairs of invariant and / or synthetic lethal residues are known, at least one candidate region is identified based on other structural criteria (Step E2). Steps E1 and E2 were described by Lao et al.

[0026] In the context of the present application, an "E2-derived candidate region" is a region comprising at least two or three amino acids of interest selected by steps E1 and E2 of the present application, i.e., according to the method described in Lao et al., to be invariant and / or synthetic lethal (SL)-covariant residues that are not derived from a common ancestor, are close to each other, and are exposed on the surface and possibly in a pocket of the candidate protein.

[0027] In the second step, the method of the present invention provides for determining the overall chemical bonds involved between protein residues, particularly those located in the candidate region (steps E3, E4, and E5 in FIG. 1). This chemical interaction information identifies at least one candidate region that is most likely to be an effective therapeutic target. Thus, the method of the present invention makes it possible to select, from the candidate regions obtained according to the guidelines of Lao et al., a therapeutic target region that is most likely to enable the identification of an effective drug.

[0028] As detailed below, one or more scores may be advantageously calculated to check that the information obtained from each step of the method of the present invention is relevant with respect to the rest of the method and the expected results. These scores provide researchers with information about the quality and reliability of the results obtained. For ease of reading the description, detailed score expressions are provided at the end of the description.

[0029] Detailed Description of the Steps of the Invention Step E1 consists of identifying pairs of invariant and / or synthetic lethal covariant residues, which will be called "target residues" in the context of the present invention, in a set of pre-cleaned and aligned nucleotide and polypeptide sequences characteristic of the candidate protein.

[0030] Databases used to store sequences of pathogenic organisms may contain erroneous sequences that, if not eliminated, can produce erroneous results.

[0031] This is why, in the method of the invention, the initial sequence must first be "cleaned" (step E11). This "cleaning" consists, for example, in filtering the initial sequence as follows: Initial sequences with a different length than the sequence selected as reference can be filtered out (if the batch of initial sequences contains enough sequences). If the batch of initial sequences does not contain many sequences and if alignment seems possible (because the sequences do not contain any hypervariable regions and do not contain any poorly described regions), an alignment can be performed to homogenize the length, in effect obtaining a single length alignment from sequences of varying lengths that is longer than the length of the sequence with the longest length. Initial sequences with a number of mutations greater than three standard deviations from the mean number of mutations per sequence is excluded.

[0032] In addition, the N- and C-terminal ends of the initial sequences can be marked and the sequences oriented in the same direction so that the ends overlap (e.g., all sequences can be oriented from N-terminal end to C-terminal end).

[0033] This cleaning step E11 also makes it possible to identify which of the set of initial sequences have non-uniform lengths and / or to exclude sequences with unacceptable anomalies. Such anomalies can be, for example, an abnormal number of mutations, an abnormal number of ambiguous amino acids, an abnormal number of missing amino acids, etc. In this respect, one or more scores or score functions can be advantageously calculated (step E11') to assess whether the initial sequences (before cleaning) do not contain too many anomalies and / or whether their lengths are sufficiently uniform (scores S1, S2, S3, S4, detailed below). If one or more of these scores are not acceptable (a value close to 0 and not 1), the method of the present invention can be interrupted, as this means that the set of initial sequences before cleaning is not robust enough to be used.

[0034] To assess whether the initial sequence set before cleaning is sufficiently complete and robust to effectively predict the presence of therapeutic target regions within the selected candidate proteins, a scoring function, denoted f(S1QS), can also be calculated at this stage. This scoring function f(S1QS) reflects the impact of the sequence data on the quality of the results obtained at the end of the method.

[0035] Also, if the number of mutations belonging to the batch of initial sequences retained after "cleaning" is too small (typically a number of mutations that does not allow statistically correct results to be obtained, e.g., χ 2 If the number of mutations does not allow the calculation of σ (covariates with fewer than five representatives), it is recommended to stop the method. In this case, it is preferable to select a new set of sequences from the BDD1 database and enrich it by downloading new sequences or modifying the candidate protein. Conversely, a good set of initial sequences is considered to exist when at least 1,000 sequences of acceptable quality have been identified, and these sequences carry sufficient mutations.

[0036] As used herein, "acceptable quality" refers to initial sequences having a number of aberrations that is less than three standard deviations from the average number of these aberrations per sequence. In this regard, a score S7 can be calculated to measure the total number of initial sequences that do not contain an abnormal number of mutations. Such a score S7 depends on the scores S2, S3, and S4 presented above.

[0037] The sequences obtained after cleaning are then advantageously aligned (step E12). Sequence alignment can be implemented in two ways. i) the sequences are all the same length: in this case, the alignment is generated by aligning all the sequences at their N-termini, or ii) They do not all have the same length: it is difficult to use multiple alignment methods, which are generally used for smaller numbers of sequences. In this case, a representative sample of the population of studied sequences is selected and a profile is generated using the HMMER suite (http: / / hmmer.org), allowing all sequences to then be aligned one by one (using the hmmbuild and hmmcalibrate functions). This method allows hundreds of thousands of sequences to be aligned in just one minute.

[0038] The quality of the sequence alignment is evaluated by the effect of gaps in the alignment (score S8), redundancy of the sequence batch (score S10), the effect of hypervariable regions (score S11), deletions and The impact of insertions (score S12), post-translational modifications (score S13), and / or the presence of different subtypes (score S14) can be assessed (step E12') by calculating one or more scores (described in more detail below). Scores S11, S12, and S13 are defined based on data available in the literature.

[0039] One or more scoring functions can be calculated at this stage to assess whether the set of aligned sequences after cleaning is sufficiently complete and robust to effectively predict the presence of therapeutic target regions within the selected candidate protein. The scoring function f(S3), detailed below, reflects the quality of the alignment and possible predictions. Some terms of this function describe the accuracy with which the initial sequence is described and have an impact on the statistical results (f(S3QS)), while other terms of this function indicate the heterogeneity of the batch of sequences studied and have an impact on the description of the target itself (f(S3SC)). More precisely, the scoring function f(S3QS) reflects the impact of alignment on the statistical prediction of the target region, and the scoring function f(S3SC) reflects the impact of alignment on the prediction of the target.

[0040] If the predicted quality of the alignment is low, the method can be interrupted in order to restart from a new initial sequence.

[0041] At the end of this alignment step, the user is presented with a set of cleaned, aligned sequences known as "test sequences."

[0042] Invariant residues are then identified in these test sequences (step E13). By definition, an "invariant residue" is an amino acid that does not change (or changes very little) its position in the protein in all of the test sequences studied. Because sequencing methods cannot be 100% reliable, an average error rate of 0.3% can be applied (Cheng C. et al., 2022). Therefore, in the context of the present invention, a residue is defined as "invariant" if it is present in the same position in at least 99.7% of the test sequences.

[0043] The quality of the selection of invariant residues can advantageously be evaluated by calculating the mutational abundance of the sequence (step E13'). In particular, a score S18 can be calculated (see below). This score makes it possible to check that the sequences in which the invariant residues are detected are sufficiently heterogeneous. Indeed, to be able to claim that a residue is invariant for functional reasons (mutations occurred at this position but were not selected because they were lethal), it is necessary to be able to show that mutations occurred elsewhere in the sequence.

[0044] Once the number of invariant residues is known, it is advantageous to determine the percentage of these residues by calculating the score S19. This score is used to evaluate the influence of the invariant positions on the final result. The targets most likely to be stable over time (and therefore unable to mutate without altering the replication activity of the virus, thus preventing the emergence of mutations that may make the mutant resistant to treatment) are those that are most invariant and therefore composed of the largest number of invariant residues.

[0045] The score function f(S invariance It is also possible to calculate the σ σ , which assesses the degree of invariance of a batch of studied sequences and gives an idea of ​​its long-term immutability. This degree of invariance depends on several variables, such as the quality of the alignment, the total number of mutations, the number of immutations, the number of synthetic lethals (representing invariance by two residues), and the number of mutations that arise for functional reasons and are not due to the presence of a common ancestor.

[0046] Other residues of interest can also be selected in the method of the present invention, as specified in FIG. 1. These are residues that are not invariant but belong to a covariant pair of interest (step E1 4).

[0047] Several statistical tests can be used to define the covariation of pairs of variables in a list of variables. 2 One such statistical test is the χ test, which is used to determine whether paired variables are independent of each other according to a threshold. 2 It is possible to use 2 Noivirt is defined as (Ai,Bj), where A and B are the specific amino acids found at positions i and j, which takes into account each of the 20 residues, as well as the mutated or non-mutated state of the original residue. 2 Once (Ai, Bj) are defined, it is preferable to rescale the results to reject false positives resulting from multiple tests performed. Therefore, the p-values ​​can be rescaled using a method known as the "false discovery rate." Residues are considered "dependent" or "covarying" if their p-values ​​are less than 0.05 (Noivirt O. et al.).

[0048] The impact of the number of covariant positions and the number of pairs of covariant positions on the final result can be assessed by calculating the scores S20 and S21, respectively.

[0049] The pairs of covariant residues identified by this calculation must then be filtered (steps E15 and E16).

[0050] As a first step, it is preferable to exclude pairs of covariant residues that share a common ancestor (step E16). This can be achieved, in particular, by studying the DNA sequences encoding the initial sequences. After aligning these DNA sequences, mutant codons that cause nonsynonymous mutations are identified and selected. In fact, nonsynonymous mutations result in the appearance of physicochemically distinct amino acids, while synonymous mutations encode the same residue. The so-called Lewontin coefficient of linkage disequilibrium D' can then be calculated using the set of recorded DNA sequence data (distinguishing between synonymous and nonsynonymous positions). This coefficient can be used to identify pairs that share the same ancestor and, therefore, whose covariation is not the result of functional interdependence. These steps make it possible to determine whether the covariation of two residues identified as "covariant" in step E14 is the result of coevolution of these two residues or whether it is due to the fact that they are phylogenetically linked to a common ancestor. In the latter case, the residue pair will not be conserved in the method of the present invention, resulting in a "false positive" known as "ancestral linkage disequilibrium" (Lao et al., Petitjean et al.).

[0051] At this stage, scores S24 and S25 can be calculated that accurately reflect the impact of the overall nonsynonymous and synonymous variation.

[0052] Finally, advantageously, it is identified whether any of the covariant pairs selected at the end of step E14 are capable of inducing the virus to be unable to replicate (step E16). These particular covariant pairs are known as "synthetic lethal". In fact, two types exist, which have completely opposite consequences: compensatory mutations. It is important to distinguish between covariant mutations (CM) and synthetic lethal (SL) mutations (Lao et al., Petitjean et al.). To do this, it is possible, for example, to calculate the so-called dissimilarity coefficient ξ, χ 2(Ai,Bj) allows a code to be assigned. This code distinguishes between CM and SL. A CM has ξ that is positive if NobsA,i,B,j ≥ NexA,i,B,j, and A and B are two residues located at positions i and j, respectively. Thus, ξA,i,B,j = +χ 2 (Ai,Bj), while SL has ξ that is negative when NobsA,i,B,j ≥ NexA,i,B,j. Therefore, ξA,i,B,j = -χ 2 (Ai,Bj), where Nobs is the number of pairs of residues A and B observed at positions i and j, and Nex is the number of pairs of residues A and B observed at positions i and j. is the number of pairs of residues A and B predicted at positions i and j (Petitjean et al.).

[0053] At this stage, a score S22 can be calculated, which reflects the impact of the number of SLs and their strength on the result obtained at the end of the method. Similarly, a score S23 can also be calculated to reflect the impact of the number of CMs and their strength on the result. The strength of a covariant pair is determined by its χ 2 (Ai,Bj). In fact, χ 2 The higher (Ai, Bj), the greater the number of pairs observed relative to the number of pairs expected in the absence of covariation. Thus, S22 and S23 define the strength of covariation for this particular pair.

[0054] At the end of these various steps, pairs of invariant residues and synthetic lethal factors are selected as part of at least one "candidate region" of the candidate protein to be studied.

[0055] Here, it is possible to calculate a score S9 that reflects the influence of variance at the 5' and 3' ends of the sequence. Indeed, in the step of aligning DNA sequences to identify covariant residues that share a common ancestor, one of the two techniques used is to align all sequences at their 5' ends. However, variance at this end reduces the chance of accurately aligning the sequences. In this case, it may be preferable to align on the 3' end, which should then be less variable.

[0056] The candidate region selected as the best therapeutic target must also satisfy a certain number of other favorable conditions. Notably, the amino acids that make up the candidate region may be stressed by the 3D structure of the candidate protein. Conversely, it may be advantageous to determine whether the candidate region contains a "pocket" that allows the drug to remain in a stable and potent manner. To evaluate these aspects, the method of the present invention includes a step E2 of evaluating the quality of the candidate region obtained in step E1 with respect to the positions of the amino acids that make up the candidate region with respect to the 3D structure of the candidate protein.

[0057] Step E2 therefore consists of identifying at least one target pair of residues that are located in close spatial distance from the regions and residues identified in step E1 and that are exposed on the surface of the candidate protein, preferably within a pocket.

[0058] To be able to bind a small molecule, a therapeutic target must consist of residues that are spatially close to each other, and in the context of the present invention, target residues in a candidate region will preferably be at most 10 angstroms apart, preferably at most 5 angstroms apart.

[0059] Advantageously, herein, target residues that are less than 5 Å and / or less than 10 Å apart are quantified by calculating a difference score to assess whether they are in sufficient number to constitute a future target that can be bound to a potential small molecule (drug). Scores S28, S29, S30, S31 may in particular be calculated (step E21′).

[0060] To obtain this information, the method of the present invention advantageously requires access to known and recorded three-dimensional structures of the candidate proteins. These structures are known and described in a dedicated 3D structural database, referred to herein as "BDD2". As in the case of nucleotide and polypeptide sequences, these 3D sequences will often have to be processed (cleaning step E6, alignment step E7) prior to the method of the present invention.

[0061] In practice, 3D structures are often in pdb format. However, this format is not uniformly applied by the entire scientific community (the format of the file itself, the names of the subunits, or the numbering of the residues, etc.). Therefore, the structure is preferably cleaned (step E6) to define a single and generalized format for all pdb files used (among other things, standardization of the file format, the numbering of the residues and atoms, the names of the subunits, and their respective positions).

[0062] Furthermore, tens or even hundreds of structures of the same protein are available. In the same way as sequences, the 3D structures recorded for the same protein are preferably aligned. These are structural alignments that reveal how different the structures recorded for this protein are. The alignment is based on a reference model. Then, the average deviation between all these structures is calculated, which is called "RMSD".

[0063] A three-dimensional structure is defined by the position in space of each of its constituent atoms. If many of these positions are missing (missing data), the 3D structure is flawed in the sense that only some of these atoms have a fixed position in the structure. If much data is missing for each of the studied structures (keeping in mind that the missing data are not at the same position in the protein space), it becomes difficult to perform structural alignment and even compare residues at the same position. The fact that a protein has several subunits further complicates its structure. Furthermore, in this case, we are talking about a quaternary structure, not just a tertiary structure. In this case, not only must the position of each atom in each subunit be defined, but also the position of each subunit relative to each other.

[0064] Therefore, under these conditions, the quality of the 3D structures can be evaluated to ensure that they are effective in predicting the presence of the target region. Therefore, several scores are calculated (step E7'). At this point, it is possible to calculate score S5, which quantifies the number of 3D structures with an abnormal number of missing data, and score S6, which evaluates the impact of the presence (if any) of different subunits on the final result.

[0065] Scores S15, S16, and S17 are also used to evaluate the alignment of 3D structures. For structure, S15 is the equivalent of Sseq for sequence. We examine whether the number of known 3D structures for the candidate protein is large. If there are fewer than 200 known 3D structures, this number is insufficient for a consistent average (however, this number can be reduced as 3D structures become more and more reliable).

[0066] S16 gives an idea of ​​the variation in structural resolution. In fact, 3D structures are determined using a variety of techniques. The two most widely used are X-ray diffraction and electron microscopy. A resolution threshold is defined for each of these structures depending on the technique used. S16 evaluates this variation in resolution. S17 shows the structural heterogeneity of the batch of structures studied. Each atom in a structure is aligned with its corresponding atom in the next structure. It is therefore possible to associate for each atom a number of positions equal to the number of structures studied.

[0067] The average value for this position in space is calculated along with the deviation from the mean. If this calculation is performed for all atoms in the structure, it is possible to calculate the average deviation, which gives an idea of ​​the heterogeneity of the structure, spatially. The TM described by this score is a variant of RMSD that normalizes it so that it is independent of the total number of positions in the structure.

[0068] Calculate the score function and use the functions f(S4), f(S 4QS ), f(S 4QC It is also possible to evaluate the quality of the structural alignment via the function f(S4). The function f(S4) reflects the quality of the alignment and the possible predictions. Several terms of this function describe the accuracy with which the 3D structure is described, and the statistical results f(S 4QS ), and the other terms of this function describe the heterogeneity of the batch of sequences studied and account for the target itself f(S 4QC ) more precisely, the score function f(S 4QS ) reflects the influence of alignment on the statistical prediction of the target region, and the score function f(S 4QC ) reflects the influence of alignment on target prediction.

[0069] From the 3D coordinates of all atoms of the residues in space, pairs of residues that are spatially close, ie, less than 5 or 10 Angstroms apart, are selected (step E21).

[0070] Furthermore, the accessibility of target residues and / or their exposure on the surface of the protein must be taken into consideration. A valid therapeutic target must not be buried in the 3D structure of the protein, otherwise drugs will not be able to reach it. Therefore, based on the three-dimensional structure of the candidate protein, residues that are buried in the protein are distinguished from those that are exposed on its surface and therefore accessible (for example, the ASA program can be used to do this). In the method of the present invention, only candidate regions containing at least two, preferably at least three accessible residues are selected (step E22).

[0071] Here again, an accessibility score can be calculated to assess the likelihood of obtaining a sufficient candidate region (step E22'): score S32 assesses the percentage of accessible positions, and score S33 assesses the percentage of accessible residues.

[0072] Finally, it may be advantageous to evaluate whether the previously selected candidate regions have a 3D structure resembling a "pocket" (step E23). To do this, it is possible to use, for example, structural prediction software such as Fpocket (Le Guilloux et al.), preferably reducing the minimum and maximum radii of the alpha sphere to 2.5 Å and 4 Å, respectively, to take into account the presence of small and large pockets.

[0073] The Fpocket software (Le Guilloux et al., 2009) identifies the radix-numbered nucleotides on the surface of a protein.

[0074]

number

[0075] The volume of the pocket that can accommodate a small drug molecule preferably satisfies the following constraints:

[0076]

number

[0077] Percentage of pockets that meet this criterion for the studied protein

[0078]

number

[0079]

number

[0080] At the end of this step, an array of spatially close SL invariant or covariant residues will be selected that is exposed on the surface of the candidate protein and preferably lies within a pocket with a volume between 60 Å and 500 Å. An array of at least five target residues is preferred.

[0081] In a complementary way, the SL invariant or covariant residues can be represented in the form of a graph (step E8). Indeed, from the detected network, groups of interdependent residues are formed using mathematical graphs. In these graphs, pairs of invariant and synthetic lethal residues are combined to form invariant groups. The nodes of the graph are residues, and the edges define bonds between residues (synthetic lethal or invariant residues, in this case they are all connected), which exist only if two residues are located within 10 Å of each other and on the surface of the candidate protein. Ideally, this type of graph can be visualized by the free Graphviz software. These graphs are a means of assessing the quality of the identified regions.

[0082] Step E3 consists of determining whether any favorable chemical interactions exist between residues and / or residue pairs in the candidate protein based on the 3D structure of the candidate protein. Alternatively, this step E3 can be limited to analyzing the favorable chemical interactions that exist between selected residues and / or residue pairs in the candidate region obtained in step E2 based on the 3D structure of this candidate region.

[0083] Genetics allows us to functionally detect the impact of microscopic physicochemical changes occurring at residues. Thus, the fact that a residue is invariant (or part of an invariant group) is the result of the physicochemical quality of one or more residues (selective pressures force the maintenance of this or these residues at this position). Conversely, if the physicochemistry of a residue changes (e.g., due to mutation or the external environment), the function and / or structure of the region may be affected. Based on this observation, it is recommended to take into account the physicochemical bonds existing between the residues of the protein being studied in order to reinforce or invalidate previously obtained results, with the aim of identifying the most relevant target regions.

[0084] Thus, for each amino acid in the candidate protein, the chemical bonds (e.g., hydrogen, hydrophobic, ionic, and repulsive) that this amino acid may potentially participate in are determined (step E3 in Figure 1). In particular, the three-dimensional structure of the protein is used to identify residue pairs in which the amino acids are sufficiently close so that the identified chemical bonds can effectively affect the function and / or stability of these residues (see step E4 in Figure 1).

[0085] Thus, the additional steps E3 and E4 allow us to identify amino acid pairs with chemical interactions that affect their function or that may change when the protein's environment (pH, temperature, ionic strength, etc.) is altered. Indeed, protein stability is often pH-dependent and varies based on the subtype and / or origin of the host organism. For example, the hemagglutinin (HA) of human influenza viruses is more stable than the hemagglutinin of avian viruses (Galloway, 2004). Furthermore, mutations lead to changes in the chemical reaction network, which in turn alters the function of this protein. The quality can be stabilized or destabilized (Byrd-Leotis L. et al.).

[0086] The intra-protein physicochemical network is composed of several types of interactions, such as hydrophobic and hydrogen interactions, as well as salt bridges, negatively repulsive interactions, and positively repulsive interactions, among others (Dyson HJ et al., Hubbard RE and Kamran Haider M., Sticke DF et al., Barlow DJ and Thornton JM, Harrison JS et al.). These interactions ensure that proteins have a very specific shape, which is why they are important. If they are not there, the structure of the protein, and certainly its function, will be altered. This is why it is important to determine these regions that are rich in these chemical bonds in the context of the method of the present invention.

[0087] Unlike hydrophobic and hydrogen interactions, whose pH sensitivity is negligible (or not well documented), electrostatic interactions are strongly affected by pH fluctuations (Harrison, JS et al., Pahari S. et al.). Histidine residues (

[0088]

number

[0089]

number

[0090]

number

[0091] Thus, the methods described herein involve characterizing the chemical interactions that exist between all atoms of a protein, or at least in preselected candidate regions, with the aim of identifying, in particular, the following interactions (Hubbard RE and Kamran Haider M., Barlow DJ and Thornton JM, Harrison JS et al., Donald JE et al., Freitas RFde and Schapira M., Onofrio A. et al.): Hydrophobic bonds between non-polar atoms of the following types (CB, CG, CE, CD1, CD2, CE2, CE3, CZ2, CZ3, CH2, CE1, CZ, CG1, CG2, CD, CH2) belonging to the following hydrophobic amino acids (ALA, MET, TRP, PHE, TYR, VAL, LEU, ILE, PRO). Atoms with this type of bond will be selected if their distance is less than 4.2 Å. Hydrogen bonds are formed between proton-donating atoms (OG, OG1, NE2, ND2, ND1, NE2, NZ, NE, NH1, NH2, OH, NE1) belonging to proton-donating amino acids (SER, THR, GLN, ASN, HIS, LYS, ARG, TYR, TRP) and proton-accepting atoms (OG, OG1, OE1, OE2, OD1, OD2, ND1, NE2, OH) belonging to proton-accepting amino acids (SER, THR, GLUN, ASP, GLN, ASN, HIS, TYR). Atoms with this type of bond are selected if their distance is less than 3.5 Å (Hubbard RE, Kamran Haider M., Sticke DF et al.). These hydrogen bonds can consist of interactions between two side chains or between a side chain and the main chain of a protein. The main chain proton donor (N) or the proton acceptor of this same chain can belong to any amino acid in the main chain. -There are two types of electrostatic bonds: Attractive bonds (salt bridges or ionic interactions) between cations (NE, NH1, NH2, NZ, NE2, ND1) of positively charged residues (ARG, LYS, HIS), especially when they are at a distance of less than 4 Å from anions (OD1, OD2, OE1, OE2) carried by negatively charged residues (ASP, GLU). Repulsive bonding between two atoms with identical charges (- / -) or (+ / +), especially when they are separated by a distance of less than 5 Å (Barlow DJ, Thornton JM, Harrison JS, et al.).

[0092] By stabilizing the resonance between the charges, all positively and negatively charged atoms can be considered in the calculations, as long as the charges are evenly distributed among the ionizable groups.

[0093] To assess the chemical interactions of each amino acid in the candidate region or in the protein, different scores can be calculated, in particular the percentage of hydrophobic bonds, and / or the percentage of hydrogen bonds, and / or the percentage of salt bridges, and / or the percentage of negative repulsions, and / or the percentage of positive repulsions, and / or the percentage of chemical bonds can be determined (step E3').

[0094] Therefore, in practice, the method of the present invention advantageously includes a step of calculating the distance between each atom of the protein, for example from a pdb, SwissProt, uniprot, or Modbase file. Preferably, distances between atoms belonging to the same position are not calculated unless they belong to different protomers. Following this calculation, atoms separated by a maximum distance of 5 Å are selected. This step is preferably performed before recording the chemical interactions between atoms. Therefore, only chemical interactions connecting nearby atoms must be considered. However, it is also possible to perform the two steps in the reverse order.

[0095] A number representing the total number of atomic pairs within 5 angstroms

[0096]

number

[0097] The score function f(S cochem ), which gives a picture of the overall chemical interactions of the protein studied (step E3').

[0098] All these steps select residues that are connected by favorable chemical interactions and that are at such a distance that these connections affect the function of these residues. This is step E4 in Figure 1.

[0099] At this stage, it is possible to represent the various chemical bonds in graph form (step E8'). In fact, a mathematical graph is formed from the previously detected bonds. The nodes of the graph are positions / residues, and the edges define the chemical bonds between the positions. Additionally, this type of graph can be visualized by the free Graphviz software. These graphs are an additional way to evaluate the identified bonds.

[0100] Step E5 identifies the best therapeutic target regions for generating potential drugs by cross-referencing the results obtained after step E2 with the results obtained after step E4.

[0101] The quality of the selected targets can advantageously be evaluated by calculating some score function. For example, the constancy of each target region can be calculated using the function f(S invX) can be calculated by the function f(S X ) is each S X For each target, a score between 0 and 1 is provided. This score is used to evaluate how effective the identified target is at preserving the drug. In other words, these two scoring functions can be used to evaluate the effectiveness of the target region.

[0102] The target region can advantageously be represented in the form of a graph (step E8''). Indeed, from the previously detected pairs, groups of interdependent residues are formed by means of mathematical graphs. These graphs integrate invariant residues and pairs of synthetic lethal residues that form the invariant group and are connected by favorable chemical bonds. The nodes of the graph are the selected positions / residues, and the edges define the bonds between the positions (synthetic lethal, invariant, or favorable chemical interactions) and only exist if both residues at that position are located on the surface of the candidate protein within 10 Angstroms. Additionally, this type of graph can be visualized by the free Graphviz software. These graphs are an additional means of assessing the quality of the identified regions.

[0103] The complexity of the graph can be calculated by score S26 (step E8'"), and the number of connected graphs can be calculated by score S27. Such complexity is useful: if a graph is complex, this means that there are many targets. If it is very complex, there may be intersections between targets that are not empty. If there are many subgraphs, this means that one residue is connected to many other residues and that a large part of the function falls into that.

[0104] Score Definitions S1:S L :Sequence length heterogeneity.

[0105] Let μ be the mean amino acid length of the sequences in the sample studied and σ the standard deviation.

[0106]

number

[0107] Score S L = 1 indicates absolute heterogeneity, while S L = 0 indicates no heterogeneity.

[0108] S2:

[0109]

number

[0110] Let μ be the average number of mutations calculated for each sequence in the studied sample compared to the reference sequence, and σ be the standard deviation. The studied sample consists of the polypeptide sequences downloaded before the cleaning step and the reference sequence introduced in step E0 described above.

[0111] The number of abnormal mutations is mut ab ≥ μ + 3 × σ,

[0112]

number

[0113]

number

[0114]

number

[0115] The closer S2 is to 1, the greater the number of abnormal mutations.

[0116] S3:

[0117]

number

[0118] Let μ be the mean number of mutations calculated for each sequence in the sample studied compared to the reference sequence, and σ the standard deviation.

[0119] The number of abnormal NYPs is NYP ab ≥ μ + 3 × σ,

[0120]

number

[0121]

number

[0122] The closer S3 is to 1, the greater the number of abnormal amino acids.

[0123] S4:

[0124]

number

[0125] Let μ be the mean number of mutations calculated for each sequence in the sample studied compared to the reference sequence, and σ the standard deviation.

[0126] An abnormal number of gaps ab ≥ μ + 3 × σ,

[0127]

number

[0128]

number

[0129] The closer the score S4 is to 1, the more abnormal the number of gaps is.

[0130] F(S1QS): Impact of sequence data on statistical prediction quality.

[0131] This function varies from 0 to 1. f(S 1QS ) is the input data used to make the prediction It is used to assess whether a system is adequate and robust against

[0132]

number

[0133] S5:

[0134]

number

[0135]

number

[0136] It has been found that the number of missing abnormal amino acids (AA) is defined as follows: AA ab =μ+3×σ Also,

[0137]

number

[0138]

number

[0139] It is noted that the choice of a particular reference structure has little effect, as its residue count will be very close to that of another described structure that does not have an unusual number of unspecified positions.

[0140] S6:S subtunit The effect of the presence of different subunits on the final result.

[0141] If the studied protein has several subunits (each designated as a subunit), they will be listed regardless of the number of sequences in the alignment. If it has only one subunit, this score will be equal to 1.

[0142]

number

[0143] S7:

[0144]

number

[0145]

number

[0146]

number

[0147] S8:

[0148]

number

[0149] Let μ be the mean number of calculated gaps per sequence of the sample studied after alignment compared to the same sequences before alignment, and σ the standard deviation.

[0150]

number

[0151] S9:

[0152]

number

[0153] The 5'-end and 3'-end 5 amino acids (AA) are studied. Invariant residues (inv), synthetic lethal (SL) pairs, and compensatory mutation (CM) pairs are listed at the end.

[0154]

number

[0155] S10:S red :Array batch redundancy.

[0156] Let μ be the mean of the sequence redundancy histogram and σ be the standard deviation. Polypeptide sequences corresponding to candidate proteins can be highly heterogeneous with respect to the mutations they carry, or only slightly heterogeneous. An extreme situation could be that the X sequences all have the same AA sequence. This batch is then assumed to be very homogeneous, with redundancy absolute and therefore a score equal to 0. Conversely, if the polypeptide sequences are very different when viewed pairwise, the redundancy will be very small and the score will tend to approach 1.

[0157]

number

[0158] If S10=1, there is absolute heterogeneity, if S10=0, there is no heterogeneity.

[0159] S11:S hyp :Influence of hypervariable regions.

[0160] Hypervariable regions (hyp) are listed and counted from the references regardless of the number of sequences in the alignment.

[0161]

number

[0162] S12:S indel :Effects of deletions and insertions.

[0163] Insertions-deletions (indels) are listed and counted from the references regardless of the number of sequences in the alignment.

[0164]

number

[0165] S13:S postrad :Influence of post-translational modifications.

[0166] The different types of post-translational modifications (postrads) are listed and counted from the references, regardless of the number of sequences in the alignment.

[0167]

number

[0168] S14:S subtype :The impact of the presence of different subtypes.

[0169] The different subtypes (each designated "subtype") belonging to the batch of sequences studied are enumerated, regardless of the number of sequences in the alignment. Indeed, some microorganisms mutate so rapidly that the evolution of these different variants over time leads to the emergence of variant subtypes.

[0170]

number

[0171] f(S3): Primary sequence alignment quality

[0172]

number

[0173] f(S 3QS ): the quality of statistical predictions of alignments

[0174]

number

[0175] f(S 3QC ): The effect of alignment on target prediction

[0176]

number

[0177] S18:S mut :The abundance of sequence variation.

[0178] μ is the number of mutations compared to the reference sequence <mut ab Consider the mean of the number of mutations calculated for each sequence in the sample of sequences with σ as the standard deviation.

[0179]

number

[0180] If S18=1, there is high inhomogeneity, if S18=0, there is no inhomogeneity.

[0181] S19:S inv : Invariant position effects.

[0182] A percentage of invariant residues (Target Quality QC) is defined. Invariant residues are identified as those in which the same amino acid is found at a given position in at least 99.7% of the sequences. The sum of these positions is:

[0183]

number

[0184]

number

[0185] f(S invariance ): Invariance

[0186] This scoring function is used to assess the true degree of invariance of the batch of sequences studied, thus giving an image of its essentiality. In fact, invariant positions are those positions that cannot be selected in the event of mutation, because they are essential for the replicative capacity of the organism from which the candidate protein originates.

[0187]

number

[0188] S20:S covar : The effect of the number of covariant positions.

[0189] The percentage of covariant residues is the pairwise covariant

[0190]

number

[0191]

number

[0192] S21:S prcovar : The effect of the number of covariant position pairs.

[0193] The percentage of covariant residue pairs is χ 2 (Noirvit protocol

[0194]

number

[0195]

number

[0196] S24:

[0197]

number

[0198] This is calculated as follows:

[0199]

number

[0200] D' AiAj and D' SiSj is shown in Lao et al. A indicates a residue in the studied sequence that is nonsynonymous with the selected reference sequence, and S indicates a residue in the studied sequence that is synonymous with the selected reference sequence.

[0201] S25:

[0202]

number

[0203] This is calculated as follows:

[0204]

number

[0205] S22:S SL : The influence of the number of SLs and their intensity on the overall results.

[0206] This score is the sum of the negative dissimilarity coefficients ξ and therefore all γ 2 The sum of SL residue pairs relative to the sum of (SL+CM).

[0207]

number

[0208] S23:S CM : The effect of number and intensity of CMs on the overall results.

[0209] This score is the sum of the positive dissimilarity coefficients ξ and therefore all γ 2 The sum of CM residue pairs over the sum of (SL+CM).

[0210]

number

[0211] S27:S MEdge :Graph complexity.

[0212] A graph G is defined by the pair (S,A), where S is a finite set of vertices and A is a set of vertices in S. 2 A pair of vertices in (s i ,s j ) is a finite set of vertices. Hence, a pair is a pair of vertices joined by an edge.

[0213] The aim is to study the invariant and SL residues involved in several pairwise relationships (spatially close, located on the protein surface, either invariant or involved in synthetic lethal relationships). From the graph of binding, this score is calculated by the node

[0214]

number

[0215]

number

[0216] S28:S connect : The number of connected graphs. A graph is connected if every pair of vertices is connected by an edge.

[0217] S28:S Near5 : spatially close residues (5 Å).

[0218] This score corresponds to the percentage of pairs of residues located within 5 angstroms of each other (Cα). Starting from the reference pdb file, for every residue pair the distance between the 2 Cα of the residue pair is calculated. The number of pairs separated by less than 5 angstroms is evaluated,

[0219]

number

[0220]

number

[0221] S29:S Near10 : spatially close residues (10 Å).

[0222] This score corresponds to the percentage of residue pairs that are located within 10 Angstroms of each other (Cα). Starting from the reference pdb file, for every residue pair the distance between the 2 Cα of the residue pair is calculated. The number of pairs separated by less than 10 Angstroms is calculated as

[0223]

number

[0224]

number

[0225] S30:S NearInvSL5 : Spatially close (5 Å) invariant residues and SL.

[0226] Starting from the reference pdb file, the distance between the 2Cα of the residue pair is calculated for all residue pairs. The number of pairs separated by less than 5 Å is evaluated.

[0227]

number

[0228]

number

[0229] S31:S NearInvSL10 : Spatially close (10 Å) invariant residues and SL.

[0230] Starting from the reference pdb file, the distance between the 2Cα of the residue pair is calculated for all residue pairs. The number of pairs separated by less than 10 Angstroms is calculated.

[0231]

number

[0232]

number

[0233] S32:S acc : Percentage of accessible residues.

[0234]

number

[0235]

number

[0236]

number

[0237] S33:

[0238]

number

[0239]

number

[0240]

number

[0241]

number

[0242]

number

[0243]

number

[0244] Two sets A and B are defined. Set A groups together connected or very loosely connected graphs (each of these subgraphs represents a subset of Set A). Recall the definition of an edge joining two residues: the two residues are less than 10 Angstroms apart, they are on the surface of the molecule, and they are either invariant or form part of an SL pair. Set B groups together targets defined by the Fpocket software (each target defines a subset of Set B).

[0245] The next step is to combine subsets of set A with subsets of set B. When these subsets are combined pairwise (one from set A, one from set B), the intersection of their elements is x ), and their union is defined (union x ).

[0246]

number

[0247]

number

[0248] f(S X ): Score 0 to 1 for each target

[0249]

number

[0250] This scoring function assigns a score to each target determined by our software and defines the "druggability" of this target. "Druggability" means that it can efficiently bind to a small molecule with therapeutic potential, and therefore represents a future drug. In addition to this definition, this target leaves little or no possibility of therapeutic escape. Therefore, a small molecule that is a future drug will have a group of residues (S) located within a concave space (called a pocket). cochemTot ) and its volume (vol x ) can be determined and the maximum possible invariance (f(S invX )

[0251] f(S4): 3D alignment quality:

[0252]

number

[0253] f(S 4QS ): Quality of statistical prediction of structural alignment:

[0254]

number

[0255] f(S 4QC ): Impact of 3D alignment on target prediction:

[0256]

number

[0257] The methods of the present invention have been performed on the influenza virus hemagglutinin (HA) protein.

[0258] This exercise highlighted several advantageous chemical bonds that allowed the selection of seven amino acid pairs / groups linked with acceptable distances between them by significant chemical and genetic bonds (see Figure 2, shaded amino acids).

[0259] These particularly advantageous pairs / groups are then subjected to steps (E1) and (E2) of the method of the invention. These seven prioritized amino acid groups were considered in the analysis of therapeutic target regions of the HA protein identified during the study. As shown in Figure 3 (using information from Figure 3 in Lao et al.), the inclusion of these seven prioritized amino acid groups reduces the number of therapeutic target regions within HA from six to three. These three therapeutic targets are therefore the focus of drug screening with a high likelihood of identifying reliable and effective active ingredients.

[0260] Therefore, the method of the present invention improves on the method previously proposed in Lao et al. by adding a step (E3) that requires the determination of all chemical interactions that exist between each residue and / or each pair of residues in the candidate protein separated by a distance of up to 5 Å (see FIG. 2). This additional step results in the selection of only pairs / groups of residues (E4) that have favorable physicochemical characteristics and are located at such a distance that chemical bonding between the residues is potentially influential.

[0261] Thus, the "chemical interaction" information proposed in this invention enhances the effectiveness of targets identified by the genetic approach of Lao et al. by selecting the most relevant ones, thus significantly reducing the number of therapeutic targets to be used in molecular library screening programs and enhancing the likelihood of identifying more effective drugs faster.

[0262] References Alland et al.,2005 Alland et al.:Alland C,Moreews F,Boens D,Carpentier M,Chiusa S,Lonquety M,Renault N,Wong Y,Cantalloube H,Chomilier J,et al.2005.RPBS:a web resource for structural bioinformatics.Nucleic Acids Res 33:W44-49 Barlow,D.J.&Thornton,J.M.Ion-pairs in proteins.Journal of Molecular Biology168,867-885(1983) Baumgart,M.et al.Design of buried charged networks in artificial proteins.Nat Commun12,1895(2021). Brouillet S.,Valere T.,Ollivier E.,Marsan L.,Vanet A.,Co-lethality studied as an asset against viral drug escape:the HIV protease case.Biol Direct.2010 Jun 17;5:40. Byrd-Leotis,L.,Galloway,S.E.,Agbogu,E.&Steinhauer,D.A.Influenza hemagglutinin(HA)stem region mutations that stabilize or destabilize the structure of multiple HA subtypes.J Virol89,4504-4516(2015). Cheng C,Xiao P.Evaluation of the correctable decoding sequencing as a new powerful strategy for DNA sequencing.Life Sci Alliance.2022 Apr 14;5(8):e202101294. Childers,M.C.,Towse,C.-L.&Daggett,V.The effect of chirality and steric hindrance on intrinsic backbone conformational propensities:tools for protein design.Protein Eng Des Sel29,271-280(2016). Di Russo,N.V.,Estrin,D.A.,Marti,M.A.&Roitberg,A.E.pH-Dependent conformational changes in proteins and their effect on experimental pK(a)s:the case of Nitrophorin 4.PLoS Comput Biol8,e1002761(2012). Donald,J.E.,Kulp,D.W.&DeGrado,W.F.Salt bridges:geometrically specific,designable interactions.Proteins79,898-915(2011). Dyson,H.J.,Wright,P.E.&Scheraga,H.A.The role of hydrophobic interactions in initiation and propagation of protein folding.Proceedings of the National Academy of Sciences103,13057-13061(2006). Fitch,C.A.,Platzer,G.,Okon,M.,Garcia-Moreno E,B.&McIntosh,L.P.Arginine:Its pKa value revisited.Protein Sci24,752-761(2015). Freitas,R.F.de&Schapira,M.A systematic analysis of atomic protein-ligand interactions in the PDB.Med.Chem.Commun.8,1970-1981 (2017). Galloway,S.E.,Reed,M.L.,Russell,C.J.&Steinhauer,D.A.Influenza HA subtypes demonstrate divergent phenotypes for cleavage activation and pH of fusion:implications for host range and adaptation.PLoS Pathog9,e1003151(2013). Harms,M.J.et al.The pKa Values of Acidic and Basic Residues Buried at the Same Internal Location in a Protein Are Governed by Different Factors.Journal of Molecular Biology389,34-47(2009). Harris,T.K.&Turner,G.J.Structural Basis of Perturbed pKa Values of Catalytic Groups in Enzyme Active Sites.IUBMB Life53,85-98(2002). Harrison,J.S.et al.Role of Electrostatic Repulsion in Controlling pH-Dependent Conformational Changes of Viral Fusion Proteins.Structure21,1085-1096(2013). Hubbard,R.E.&Kamran Haider,M.Hydrogen Bonds in Proteins:Role and Strength in ELS(John Wiley&Sons,Ltd,2010). Kampmann,T.,Mueller,D.S.,Mark,A.E.,Young,P.R.&Kobe,B.The Role of Histidine Residues in Low-pH-Mediated Viral Membrane Fusion.Structure14,1481-1487(2006). Kuiken H.J.,Beijersbergen R.L.,Exploration of synthetic lethal interactions as cancer drug targets.Future Oncol.2010 No v;6(11):1789-802 Lao J.and Vanet A.A New Strategy to Reduce Influenza Escape:Detecting Therapeutic Targets Constituted of Invariance Groups.Viruses.2017 Mar 2;9(3):38. Mellman,I.,Fuchs,R.&Helenius,A.Acidification of the endocytic and exocytic pathways.Annu.Rev.Biochem.55,663-700(1986). Meuzelaar,H.,Vreede,J.&Woutersen,S.Influence of Glu / Arg,Asp / Arg,and Glu / Lys Salt Bridges on α-Helical Stability and Folding Kinetics.Biophysical Journal110,2328-2341(2016). Le Guilloux,V.;Schmidtke,P.;Tuffery,P.Fpocket:An open source platform for ligand pocket detection.BMC Bioinform.2009,10,168 Noivirt,O.;Eisenstein,M.Horovitz,A.Detection and reduction of evolutionary noise in correlated mutation analysis.Protein Eng.Des.Sel.2005,18,247-253 Onofrio,A.et al.Distance-dependent hydrophobic-hydrophobic contacts in protein folding simulations.Phys.Chem.Chem.Phys.16,18907-18917 (2014). Pahari,S.,Sun,L.&Alexov,E.PKAD:a database of experimentally measured pKa values of ionizable groups in proteins.Database 2019,baz024(2019). Petitjean M.,Badel A.,Veitia R.A.,Vanet A.,Synthetic lethals in HIV:ways to avoid drug resistance:Running title:Preventing HIV resistance.Biol Direct.2015 Apr 17;10:17 Sticke,D.F.,Presta,L.G.,Dill,K.A.&Rose,G.D.Hydrogen bonding in globular proteins.Journal of Molecular Biology226,1143-1159(1992).

Claims

1. 1. A computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising: a) identifying pairs of invariant and / or synthetic lethal residues, referred to as target residues, in a set of pre-aligned nucleotide and polypeptide sequences characteristic of said candidate protein (E1); b) identifying at least one candidate region consisting of at least one pair of target residues identified in step a), said at least one pair being located at a determined distance in space and comprising target residues exposed on the surface of said candidate protein, preferably in a pocket (E2); c) determining from the 2D and / or 3D structure of said candidate protein favorable chemical interactions between each residue and / or each pair of residues in said candidate protein, said favorable chemical interactions being hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or negative repulsive bonds and / or positive repulsive bonds (E3); d) selecting residues linked by said favorable chemical interactions, said residues being at a distance of at most 10 Angstroms (E4); e) selecting from among the candidate regions identified in step b) at least one therapeutic target region comprising the residue selected in step d).

2. 2. The method of claim 1, wherein the step (E3) of identifying favorable chemical interactions consists of determining the chemical interactions of all residues in the protein and / or determining the chemical interactions of all pairs of residues in the protein, the residues or pairs of residues being separated by a determined distance of at most 10 Angstroms.

3. The method according to claim 1 or 2, wherein the distance between the invariant residues or synthetic lethal pairs determined in step b) is between 2 and 8 Angstroms, preferably 5 Angstroms.

4. The method of any one of claims 1 to 3, further comprising an evaluation (E3') of the chemical interactions of the determined bonds (E3), wherein the evaluation comprises calculating at least one score from the percentage of hydrophobic bonds and / or the percentage of hydrogen bonds and / or the percentage of salt bridges and / or the percentage of negative repulsive bonds and / or the percentage of positive repulsive bonds and / or the percentage of chemical bonds.

5. 5. The method of claim 1, wherein the target region is a pocket identified from a set of aligned 3D structures of the candidate protein, and the chemical interactions are determined for the set of aligned 3D structures of the candidate protein.

6. a target pocket comprising at least five target residues selected from synthetic lethal invariant residues or covariant residues, said pocket comprising: [Equation 1] 6. The method of claim 5, wherein the volume of the mixture is

7. 7. The method according to any one of claims 1 to 6, comprising an assessment of the quality of the spatial location of target residues, characterized in that said assessment is characterized by calculating the percentage of residue pairs located within 5 Angstroms of each other and / or the percentage of residue pairs located within 10 Angstroms of each other.

8. 8. The method of any one of claims 1 to 7, comprising the step (E0) of selecting a reference polypeptide sequence for the candidate protein, followed by the step of selecting a polypeptide test sequence that has the same length as the reference sequence, and / or has a number of mutations that is no more than three standard deviations of the average number of mutations per sequence, and / or has invariant residues at its N- or C-terminus.

9. 9. The method of claim 8, further comprising evaluating (E11') the quality of the selected polypeptide sequences, wherein said evaluating comprises calculating at least one score measuring the heterogeneity in length of the selected polypeptide sequences compared to the reference polypeptide sequences of the candidate protein, and / or the number of selected polypeptide sequences having an unusual number of ambiguous amino acids, and / or the number of sequences having an unusual number of deleted residues.

10. 1. A computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising: a) determining from the 2D and / or 3D structure of said candidate protein favorable chemical interactions between each residue and / or each pair of residues in said candidate protein, said favorable chemical interactions being hydrophobic bonds and / or hydrogen bonds and / or salt bridges and / or negative repulsive bonds and / or positive repulsive bonds (E3); b) selecting residues linked by said favorable chemical interaction, said residues being at most 10 angstroms apart, said residues forming a therapeutic target region (E4).