Computer-implemented method for treatment recommendation
A computer-implemented method optimizes ligand configurations for personalized cancer therapy by simulating protein-ligand interactions, addressing the challenge of selecting effective cancer treatments for individual tumors.
Patent Information
- Application Number
- PCT/EP2025/068138
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-03
- Filing Date
- 2025-06-26
- Publication Date
- 2026-01-08
AI Technical Summary
Selecting the most appropriate cancer therapy for a patient is challenging due to the complexity of cancer, its genetic mutations, and the variability between individual tumors, despite advances in cancer genomics and molecular profiling.
A computer-implemented method that optimizes ligand configurations based on mutated protein structures, generates protein-ligand complex ensembles, and scores them for treatment recommendations, using molecular docking and dynamics simulations to identify effective drug candidates.
This method provides accurate and efficient treatment recommendations by simulating protein-ligand interactions, allowing for personalized cancer therapy and reducing the risk of treatment resistance.
Smart Images

Figure EP2025068138_08012026_PF_FP_ABST
Abstract
Description
[0001] Computer-implemented method for treatment recommendation
[0002] Field of Invention
[0003] The present invention relates to the field of computer-implemented methods for treatment recommendation, and more particularly to techniques for optimizing ligand configurations based on mutated protein structures and generating protein-ligand complex ensembles for scoring and recommendation.
[0004] Background
[0005] Cancer is a complex and multifaceted disease that occurs due to the uncontrolled growth and spread of abnormal cells in the body. It can affect any part of the body and can lead to various health problems or even death if not detected and treated early.
[0006] Cancer cells are prone to mutations because they have lost their ability to regulate their growth and division. Normally, cells in the human body grow and divide in a controlled manner, and any abnormalities or mistakes in their DNA are quickly repaired. However, in cancer cells, the DNA repair mechanism is impaired, which allows mutations to accumulate over time.
[0007] Moreover, cancer cells can also acquire new mutations as they continue to grow and spread, which can make them more aggressive and resistant to treatment. During cancer treatment, some cancer cells may be killed, but others may survive and continue to evolve, developing new mutations that make them resistant to the treatment. As a result, the cancer may progress despite treatment, requiring new or different treatments to be used.
[0008] This is why cancer treatment often involves a combination of therapies, such as chemotherapy, radiation therapy, and targeted therapies, to attack the cancer cells from different angles and minimize the risk of acquired resistance. Selecting the most appropriate therapy for a patient's cancer can be challenging, as there are many different types of cancer, each with unique characteristics and genetic mutations. Additionally, even within a specific type of cancer, there can be considerable variation between individual tumors.
[0009] Clinicians typically rely on a combination of factors to guide their selection of therapy, including the patient's medical history, the type and stage of the cancer, the results of diagnostic tests and imaging studies, and the patient's overall health and preferences. In recent years, advances in cancer genomics and molecular profiling have provided new tools for identifying specific genetic mutations or other biomarkers that may be driving a patient's cancer. This information can help guide the selection of targeted therapies that are more likely to be effective.
[0010] Despite these advances, however, there is still much that is not known about the effectiveness of different cancer treatments, particularly in the context of individual patients with specific genetic mutations or other biomarkers.
[0011] Summary
[0012] It is therefore an objective of the present invention to provide a computer implemented method allowing for the testing of the efficacy of drugs with digital tools in tailor-made models of the patient’s cancer-driving proteins.
[0013] According to the first aspect there is provided a computer-implemented method of treatment recommendation. A computer-implemented method may be understood as a method that involves the use of a computer or processor to carry out at least some steps of the method. The method comprises providing a first digital resource representative for a mutated protein structure, abbreviated as MPS, related to a target protein type. A mutated protein structure refers to changes in the three-dimensional shape of a protein that occur due to changes in its amino acid sequence, for example due to a mutation. When a protein's structure is altered, its function may also be affected, which can lead to the development of diseases such as cancer. The term mutated protein structure refers to a protein structure that has been altered from its normal or wild-type state due to genetic mutations or other factors. The MPS is related to a target protein type. The target protein type may be any protein that is involved in cancer or other diseases. EGFR and TSRAF are examples of target protein types that could be used in the computer-implemented method. EGFR (epidermal growth factor receptor) is a transmembrane protein that is involved in regulating cell growth and division. Mutations in EGFR are commonly found in various types of cancer, including lung cancer, and drugs that target EGFR are used in cancer treatment. TSRAF (transducin-like enhancer of split-related protein) is a transcription factor that is involved in regulating gene expression. Mutations in TSRAF have been linked to various types of cancer, including leukemia and breast cancer.
[0014] The method further comprises providing a second digital resource representative for an initial molecular configuration of at least one ligand. The initial molecular configuration of at least one ligand refers to the specific arrangement of atoms and bonds within a molecule that is capable of binding to a receptor site on a target protein. Ligands are molecules that bind to specific receptor sites on proteins, and the initial molecular configuration of the ligand refers to the starting point of a molecular configuration of the ligand.
[0015] Next, the initial molecular configuration of the ligand is optimized into an optimized molecular configuration of at least one ligand based on the provided MPS. The term optimized molecular configuration refers to a configuration or a plurality of configurations of the ligand that are predicted to have the most favorable interaction with the MPS. In other words, the optimized molecular configuration of the ligand is a molecular configuration of the ligand that has the most favorable interactions with the mutated protein structure.
[0016] The method further comprises sampling a plurality of alternate conformations of the initial molecular configuration of at least one ligand and generating a third digital resource representative of a plurality of protein-ligand complex ensembles. The third digital resource comprises at least a first protein-ligand complex ensemble representative of the MPS and the optimized molecular configuration of the at least one ligand, and a plurality of second protein-ligand complex ensembles representative of the MPS and respective sampled alternate conformations of the at least one ligand. This is done to explore the conformational space of the ligand and identify alternative configurations that may also interact favorably the with MPS. Once the alternate conformations are sampled, the method generates a third digital resource that is representative of a plurality of protein-ligand complex ensembles. This digital resource contains, for example, information about the various conformations of the ligand and their interactions with the MPS, as well as the energy landscape of the ligand-MPS complex
[0017] Next, after generating the first and plurality of second protein-ligand complex ensembles, the method involves scoring these ensembles based on a scoring parameter that represents the energy of interaction between the ligand and the MPS. The energy of interaction is a factor in determining the strength of binding between the ligand and the MPS and is used as a measure of the ligand efficacy as a drug candidate.
[0018] Finally, the scored first and second protein-ligand complex ensembles and the respective recommendation categories are output.
[0019] By scoring the protein-ligand complex ensembles based on a scoring parameter, and outputting the scored first and second protein-ligand complex ensembles, the method provides the advantage that it allows for the generation of a more comprehensive set of protein-ligand complex ensembles, which can lead to more accurate predictions of the interaction between the ligand and the MPS. Additionally, by scoring the protein-ligand complex ensembles based on a scoring parameter representative for at least an energy of interaction, the method can identify ligands that are more likely to have a favorable interaction with the MPS, which can lead to more effective treatment recommendations. Another advantage is that the method allows for the automatic outputting of treatment recommendations based on the results of the method, which can save time and improve the efficiency of the treatment recommendation process. Additionally, by providing recommendation categories, the method can help clinicians to quickly identify the most effective treatment options for their patients. Yet another advantage is that it allows for the testing of the efficacy of drugs with digital tools in tailor-made models of the patient's cancer-driving proteins. By optimizing the molecular configuration of the ligand based on the provided MPS, the method can generate more accurate predictions of the interaction between the ligand and the MPS, which can lead to more effective treatment recommendations.
[0020] The computer-implemented method may further comprise classifying the scored first and second protein-ligand complex ensembles into one or more treatment recommendation categories after the step of scoring the first and the plurality of second protein-ligand complex ensembles, and outputting the classified first and second protein-ligand complex ensembles and the respective recommendation categories. Classification of the protein-ligand complex ensembles into treatment recommendation categories can help to identify the most effective treatment options for individual patients based on the predicted interaction between the ligand and the mutated protein structure. These recommendation categories can be based on various factors, such as the strength of the interaction, the specificity of the ligand for the mutated protein structure, or the predicted effectiveness of the ligand in inhibiting the activity of the mutated protein structure. The computer- implemented method may also comprise classifying the scored first and second protein-ligand complex ensembles by protein-type and ligand complex. In the method described, the scoring is done per each protein-ligand complex, which can be as many as the number of protein template structures for each protein type. The classification, however, is done per pair of protein-type and ligand, based on the overall suitability of the ligand for treatment. It is noted that the protein-type ligand complex refers to a set of protein-ligand complexes that involve the same type of protein structure which is different from a protein-ligand complex. For example, the protein-type may refer to a specific type of receptor protein, such as the EGFR protein, and the ligand molecules may include various tyrosine kinase inhibitors that target the EGFR protein. By outputting the classified protein-ligand complex ensembles and recommendation categories, the method can provide clinicians with a clear and concise report of the most effective treatment options for their patients based on the predicted interaction between the ligand and the mutated protein structure. This can help to guide treatment decisions and improve the overall effectiveness of cancer treatment.
[0021] The computer-implemented method may further comprise providing the first digital resource representative of the MPS by inputting a first input parameter representative for a reference amino acid sequence of a wild type target protein, such as EGFR, BRAF, HER2, TNF- alpha, p53, AKL, KRAS, myosin, etc., and inputting a second input parameter representative for a mutation in the wild type target protein. The term first input parameter representative for a reference amino acid sequence refers to a variable or input used in the computational method that represents a known amino acid sequence for a particular protein.
[0022] To generate the fourth digital resource representative for a mutated amino acid sequence, the method applies the mutation represented by the second input parameter to the reference amino acid sequence of the wild type target protein represented by the first input parameter. The first digital resource representative for the MPS is then generated based on the generated fourth digital resource representative for the mutated amino acid sequence. This allows the method to generate the MPS for a specific mutation in a wild type target protein, which can be used to optimize the molecular configuration of the ligand and generate protein-ligand complex ensembles that are tailored to the specific mutation in the target protein. By providing a more accurate representation of the mutated protein structure, this approach improves the accuracy of the treatment recommendations generated by the method and increases the likelihood of successful cancer treatment.
[0023] Generating the first digital resource representative for the MPS may further comprise a plurality of first model parameters representative for a plurality of protein template structures. Generating the first digital resource representative for the MPS refers to the process of creating a digital representation of the mutated protein structure that is used in the computer-implemented method. To generate the first digital resource representative for the MPS, the method identifies a protein template structure that corresponds to the wild type target protein, aligns the mutated amino acid sequence of the mutated target protein with the identified protein template structure, and computationally adjusts the positions and orientations of the particles in the digital representation of the mutated target protein sequence to maximize similarity to those of the identified protein template structure. The plurality of first model parameters are inputs of data that describe a set of protein template structures. The plurality of first model parameters representative for a plurality of protein template structures means that each of the first model parameters describes a different protein structure that can be used as a template for the mutated protein structure. The computer- implemented method may further comprise tracking first changes of the protein template structure corresponding to the wild type target protein to the generated MPS. Tracking first changes of the protein template structure involves monitoring for example the alignment of the mutated amino acid sequence with the identified protein template structure and adjusting the digital representation of the MPS accordingly to ensure that it accurately represents the mutated protein structure. This can help to improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment. By tracking changes to the protein template structure, the method can ensure that the digital representation of the MPS accurately reflects the mutated protein structure and can be used to optimize the molecular configuration of the ligand and generate protein-ligand complex ensembles that are tailored to the specific mutation in the target protein.
[0024] The computer-implemented method may further comprise outputting a three-dimensional representation of the mutated protein structure. Outputting a three-dimensional representation of the mutated protein structure can provide clinicians with a visual representation of the structural changes that have occurred due to the mutation. This can help to improve their understanding of the disease and guide the selection of appropriate treatment options. Additionally, a three- dimensional representation of the mutated protein structure can be used to optimize the molecular configuration of the ligand and generate protein-ligand complex ensembles that are tailored to the specific mutation in the target protein. By providing a more accurate representation of the mutated protein structure, this approach can improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment.
[0025] The computer-implemented method may include highlighting one or more mutations in the three-dimensional representation of the mutated protein structure. Highlighting the mutations in the three-dimensional representation of the mutated protein structure can provide clinicians with a clear visual representation of the specific changes that have occurred due to the mutation. This can help to improve their understanding of the disease and guide the selection of appropriate treatment options. By highlighting the mutations in the three-dimensional representation of the mutated protein structure, the method can also help to identify ligands that are more likely to have a favorable interaction with the mutated protein structure. This can improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment.
[0026] The computer-implemented method may comprise storing the plurality of first model parameters representative for the plurality of protein template structures on a database. Storing the first model parameters on a database can help to improve the efficiency and scalability of the computer-implemented method by allowing for rapid access to a large number of protein template structures. This can facilitate the selection of the most appropriate protein template structure for the mutated protein structure and improve the accuracy of the digital representation of the MPS. Additionally, storing the first model parameters on a database can allow for the method to be easily updated with new protein template structures as they become available. This can help to ensure that the method remains up-to-date and accurate over time and can improve the overall effectiveness of cancer treatment. The plurality of first model parameters representative for the plurality of protein template structures may comprise at least one of an atomic coordinate of particles in the protein structure, distance lengths between particles, bond lengths and / or angles between particles, torsion angles of amino acid side chains and backbone, solvent accessibility of amino acid residues, hydrogen bonding patterns between amino acid residues, secondary structure information such as alphahelices and beta-sheets, electrostatic potential maps of the protein structure, energy functions that describe the stability and / or interactions of the protein structure. These first model parameters describe various aspects of the protein structure that can be used as a reference to generate a digital representation of the mutated protein structure. By including these parameters in the plurality of first model parameters, the method can generate a more accurate digital representation of the MPS and improve the accuracy of the treatment recommendations generated by the method. For example, the electrostatic potential maps of the protein structure can be used to predict the electrostatic interactions between the ligand and the MPS, while the energy functions can be used to predict the stability of the protein-ligand complex ensembles. By including these parameters in the plurality of first model parameters, the method can generate more accurate predictions of the interaction between the ligand and the mutated protein structure, which can lead to more effective treatment recommendations.
[0027] The computer-implemented method may comprise the step of optimizing the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand using molecular docking. Molecular docking is a computational method that can be used to predict the binding modes and affinities of ligands to protein structures. In the context of the computer-implemented method described, molecular docking can be used to predict the optimal molecular configuration of the ligand that will have the most favorable interaction with the MPS.
[0028] By using molecular docking to optimize the molecular configuration of the ligand, the method can generate more accurate predictions of the interaction between the ligand and the mutated protein structure, which can lead to more effective treatment recommendations. Additionally, by optimizing the molecular configuration of the ligand, the method can identify ligands that are more likely to have a favorable interaction with the MPS, which can increase the likelihood of successful cancer treatment.
[0029] The molecular docking may comprise generating a digital set of possible ligand configurations and evaluating each ligand configuration of the digital set of possible ligand configuration for the ability to bind to or their strength of interaction with the target protein. By evaluating each ligand configuration in the digital set for the ability to bind to the target protein, the method can identify ligands that are more likely to have a favorable interaction with the MPS. This can improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment.
[0030] The computer-implemented method may comprise the step of evaluating each ligand configuration by scoring the energy of interaction between the ligand and the MPS. Scoring the energy of interaction between the ligand and the MPS can help to identify ligands that are more likely to have a favorable interaction with the MPS. By assigning a score to each ligand configuration based on its energy of interaction, the method can identify ligands that are more likely to be effective in inhibiting the activity of the mutated protein structure. Additionally, by using the energy of interaction as a scoring parameter, the method can identify ligands that are more specific to the mutated protein structure and less likely to interact with other proteins in the body. This can increase the effectiveness of cancer treatment while minimizing the risk of side effects.
[0031] The computer-implemented method may comprise the step of optimizing the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand using a molecular dynamics simulation. Molecular dynamics simulation is a computational method that can be used to simulate the movement and interactions of atoms and molecules over time. In the context of the computer-implemented method described in the patent application, molecular dynamics simulation can be used to predict the optimal molecular configuration of the ligand that will have the most favorable interaction with the MPS. By using molecular dynamics simulation to optimize the molecular configuration of the ligand, the method can generate more accurate predictions of the interaction between the ligand and the mutated protein structure, which can lead to more effective treatment recommendations. Additionally, by optimizing the molecular configuration of the ligand, the method can identify ligands that are more likely to have a favorable interaction with the MPS, which can increase the likelihood of successful cancer treatment.
[0032] In the molecular dynamics simulation step of the computer-implemented method, a digital set of possible ligand configurations may be generated and then a simulation is run to simulate the motion and interaction of particles in the ligand configurations. This is done by solving one or more equations of motion for each particle in the system. The simulation is used to evaluate each ligand configuration in the digital set for the strength of interaction with the target protein, which can help to identify ligands that are more likely to be effective in inhibiting the activity of the mutated protein structure. During the optimization of the initial molecular configuration of the ligand, the method may further comprise tracking second changes molecular of the initial configuration to the optimized molecular configuration of the ligand. This involves monitoring the molecular changes that occur during the optimization process and adjusting the molecular configuration of the ligand accordingly. By tracking these changes, the method can ensure that the optimized molecular configuration accurately reflects the most favorable interaction with the MPS. This can improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment.
[0033] The computer-implemented method may further comprise scoring the first and second changes by evaluating the changes in the molecular interactions between the optimized molecular configuration of the ligand and the MPS. This scoring process can be based on a variety of factors, such as the energy of interaction between the ligand and the MPS, the degree of similarity between the protein-ligand complex and scientific data, or other relevant information about the MPS and the ligand. By scoring the first and second changes in this way, the method can identify the most promising ligand-MPS complexes and provide treatment recommendations based on the predicted interaction between the ligand and the mutated protein structure.
[0034] During the step of optimizing the molecular configuration of the ligand, the first and second changes may be propagated to ensure that the optimization process accurately reflects the most favorable interaction with the MPS. Propagating the first and second changes involves ensuring that any changes made to the ligand configuration during the optimization process are consistent with the changes in the MPS. This can help to ensure that the optimized molecular configuration accurately reflects the most favorable interaction with the mutated protein structure and can improve the accuracy of the treatment recommendations generated by the method.
[0035] The third digital resource representative for the first protein-ligand complex ensemble may include one or more interaction parameters that define the interaction between the ligand and the MPS in the respective protein-ligand complex ensemble. These interaction parameters may include, for example, the energy of interaction between the ligand and the MPS, the specific amino acid residues in the MPS that are involved in the interaction, or the type of chemical bond between the ligand and the MPS. By including these interaction parameters in the third digital resource, the method can provide clinicians with a detailed understanding of the molecular interactions between the ligand and the mutated protein structure. This can help to guide treatment decisions and improve the overall effectiveness of cancer treatment. The third digital resource representative of the plurality of protein-ligand complex ensembles may comprise second model parameters that describe the interaction between the MPS and the ligand in the respective protein-ligand complex ensembles. These second model parameters may include, for example, empirical potentials that describe the energy of interaction between the ligand and the MPS. Empirical potentials are mathematical functions that are used to describe the energy of interaction between two molecules based on their distance and orientation. By using these empirical potentials to describe the interaction between the ligand and the MPS, the method can generate more accurate predictions of the interaction between the two molecules and identify ligands that are more likely to be effective in inhibiting the activity of the mutated protein structure. Examples of protein-ligand complex ensembles that may be modelled using empirical potentials include EGFR:afatinib and EGFR:gefitinib, which are protein-ligand complexes involving EGFR and two tyrosine kinase inhibitors that are used in cancer treatment.
[0036] The computer-implemented method may further comprise applying coarse-graining to the third digital resource representative of the plurality of protein-ligand complex ensembles. This involves grouping the particles or molecules of the molecular configuration of the protein and / or the ligand into one or more moieties, respectively. The second model parameters are then used to describe the interaction between the moieties of the protein and the ligand in the respective proteinligand complex ensembles. By using coarse-graining to simplify the molecular models, the method can reduce the computational complexity of the simulation and make it more efficient. This can facilitate the generation of a larger number of protein-ligand complex ensembles and improve the accuracy of the treatment recommendations generated by the method.
[0037] The computer-implemented method may further comprise outputting one or more treatment recommendations based on the scored first and second protein-ligand complex ensembles.
[0038] According to another aspect, there is provided a computer program product comprising a computer-executable program of instructions for performing, when executed on a computer, the steps of the method of any one of the method embodiments described above.
[0039] It will be understood by the skilled person that the features and advantages disclosed hereinabove with respect to embodiments of the method may also apply, mutatis mutandis, to embodiments of the computer program product.
[0040] According to yet another aspect of the present invention, there is provided a digital storage medium encoding a computer-executable program of instructions to perform, when executed on a computer, the steps of the method of any one of the method embodiments described above. It will be understood by the skilled person that the features and advantages disclosed hereinabove with respect to embodiments of the method may also apply, mutatis mutandis, to embodiments of the digital storage medium.
[0041] According to yet another aspect of the present invention, there is provided a device programmed to perform a method comprising the steps of any one of the methods of the method embodiments described above.
[0042] According to yet another aspect of the present invention, there is provided a method for downloading to a digital storage medium a computer-executable program of instructions to perform, when executed on a computer, the steps of the method of any one of the method embodiments described above.
[0043] It will be understood by the skilled person that the features and advantages disclosed hereinabove with respect to embodiments of the method may also apply, mutatis mutandis, to embodiments of the method for downloading.
[0044] Brief description of the figures
[0045] The accompanying drawings are used to illustrate presently preferred non-limiting exemplary embodiments of devices of the present invention. The above and other advantages of the features and objects of the present invention will become more apparent and the present invention will be better understood from the following detailed description when read in conjunction with the accompanying drawings, in which:
[0046] Figure 1 schematically illustrates a flowchart of an exemplary embodiment of a computer- implemented method of treatment recommendation;
[0047] Figure 2 schematically illustrates a flowchart of an exemplary embodiment of a computer- implemented method of treatment recommendation, e.g. a further development of the exemplary embodiment shown in Figure 1 ;
[0048] Figure 3, 4 and 5 schematically illustrates protein structures and ligands in unmutated, mutated and optimized configuration forms, respectively;
[0049] Figures 6 schematically illustrates a flowchart of a computer-implemented method to generate a first digital resource representative for the mutated protein structure based on a generated fourth digital resource representative for the mutated amino acid sequence.
[0050] Description of embodiments
[0051] According to a first aspect there is provided a computer-implemented method of treatment recommendation. A computer-implemented method may be understood as a method that involves the use of a computer or processor to carry out at least some steps of the method. Preferably, all the steps are executed by the computer or processor. The computer-implemented method comprises providing 100 a first digital resource representative for a mutated protein structure, abbreviated as MPS 1000’, shown in figure 4, related to a target protein type. A mutated protein structure refers to a protein structure that comprises at least of a set of atoms and their atomic coordinates, the MPS is derived from a protein template structure and includes changes in its chemical composition, described as mutations to the protein that occur due to changes in its amino acid sequence, for example due to a mutation. When a protein's structure is altered, its function may also be affected, as will be explained in relation to figures 3 and 4, which can lead to the development of diseases such as cancer. Figure 4 schematically illustrates a mutated protein structure 1000 where a part 1100 of the MPS 1000 has mutated. An unmutated protein structure is schematically shown in figure 3. In said figure 3, the part 1100 of the protein structure 1000 is the original part. Examples of protein structures that can be mutated include RCA1 and BRCA2 which are tumor suppressor genes that play a role in DNA repair. Mutations in these genes are associated with an increased risk of breast and ovarian cancer. KRAS is a proto-oncogene that regulates cell growth and division. Mutations in KRAS are commonly found in many types of cancer, including lung, colon, and pancreatic cancer. TP53 is another tumor suppressor gene that helps regulate cell growth and division. Mutations in TP53 are associated with an increased risk of many types of cancer, including lung, breast, and colon cancer. EGFR is a receptor protein that helps regulate cell growth and division. Mutations in EGFR are commonly found in non-small cell lung cancer and are associated with a poor prognosis. BCL2 is a protein that helps regulate cell death. Mutations in BCL2 are associated with an increased risk of lymphoma and other types of cancer. Other examples include BRAF, HER2, TNF-alpha, p53, Aik, Kras, MYOSIN, etc. It is noted that the present application is not limited to oncology but can be applied in other fields such as neurodegenerative diseases where mutations in proteins such as alpha-synuclein and tau are associated with neurodegenerative diseases such as Parkinson's and Alzheimer's. Mutations in proteins such as EDE receptor and apoB are associated with cardiovascular diseases such as atherosclerosis. Mutations in proteins such as CFTR and dystrophin are associated with genetic disorders such as cystic fibrosis and muscular dystrophy. Mutations in proteins such as hemagglutinin and neuraminidase are associated with the virulence of influenza viruses. In each of these fields, understanding the impact of mutated protein structures on protein function is critical for developing treatments and diagnostic tools.
[0052] A digital resource representative for a mutated protein structure refers to a collection of data, information, and other resources that describe the mutated protein structure and that are available electronically and can be accessed through digital means. This may include databases, computational tools, software applications, and other resources that are specifically designed to analyse, model, and visualize the three-dimensional structure of the mutated protein. Such digital resources can be used to better understand the impact of the mutation on the protein's structure and function, and to identify potential therapeutic drugs that can be used as potentially suitable treatments for diseases caused by the mutated protein.
[0053] It is noted that the first digital resource representative for an MPS 1000’ can also be generated from a wild type target protein 1000, shown in figure 3, as will be elaborated below, that generated first digital resource representative for the MPS 1000’ can then be provided to the method in the step of providing 100 the first digital resource representative for a mutated protein structure, MPS 1000’, related to a target protein type.
[0054] The term mutated protein structure, abbreviated as MPS 1000’ throughout this text, refers to a protein structure that has been altered from its normal or wild- type state, shown in figure 3, due to genetic mutations or other factors. The MPS 1000’is always related to a target protein type. The target protein type may be any protein that is involved in cancer or other diseases. EGFR and TSRAF are examples of target protein types that could be used in the computer-implemented method. EGFR (epidermal growth factor receptor) is a transmembrane protein that is involved in regulating cell growth and division. Mutations in EGFR are commonly found in various types of cancer, including lung cancer, and drugs that target EGFR are used in cancer treatment. TSRAF (transducin-like enhancer of split-related protein) is a transcription factor that is involved in regulating gene expression. Mutations in TSRAF have been linked to various types of cancer, including leukemia and breast cancer.
[0055] The method 100 further comprises providing 200 a second digital resource representative for an initial molecular configuration 1300 of at least one ligand, shown in figure 3. The initial molecular configuration 1300 of at least one ligand refers to the specific arrangement of atoms and bonds within a molecule that is capable of binding to a receptor site on a target protein. Ligands are molecules that bind to specific receptor sites on proteins, and the initial molecular configuration 1300 of the ligand refers to the starting point of a molecular configuration of the ligand. The initial molecular configuration 1300 of at least one ligand is shown in figure 3.
[0056] Next, the initial molecular configuration of the ligand is optimized 300 into an optimized molecular configuration of the at least one ligand based on the provided MPS 1000’. The term optimized molecular configuration refers to a configuration or a plurality of configurations of the ligand that are predicted to have the most favourable interaction with the MPS 1000’. A ligand 1300’ which has an improved molecular configuration with respect to the MPS 1000’ is shown in figure 5. In other words, the optimized molecular configuration of the ligand is a molecular configuration of the ligand that has the most favorable interactions with the mutated protein structure. It will be clear that more than one possible molecular configuration of the at least one ligand with highly favourable interaction with MPS can be determined, for example three ligands may be determined to have a highly favourable optimized molecular configuration. One approach that can be used to optimize the ligand's molecular configuration based on the provided MPS is by searching through available chemical space. This corresponds to the step of sampling 300 a plurality of alternate conformations of the initial molecular configuration of the at least one ligand. This involves searching through a database of chemical compounds and identifying those that are most likely to interact with the MPS. The database comprises a predefined set of starting structures for each ligand configuration that have been developed for presently available drug treatments for a disease, for example a predefined set of starting structure of ligand configuration that are available for lung oncology.
[0057] This can be done using computational methods such as virtual screening, which involves docking a library of compounds to the MPS and selecting those that show the highest binding affinity. Once a set of potential ligands has been identified, stereochemically acceptable conformations can be sampled using computational methods such as molecular dynamics simulations or Monte Carlo simulations.
[0058] After the ligand's molecular configuration has been optimized based on the provided MPS, the next step is to generate 400 a third digital resource representative of a plurality of proteinligand complex ensembles. This third digital resource represents a collection of data, information, and other resources that describe the first protein-ligand complex ensemble representative of the MPS and the optimized molecular configuration of the at least one ligand are available electronically and can be accessed through digital means. It includes information about the predicted interactions between the optimized ligand and the MPS, as well as information about the resulting protein-ligand complex structures. The third digital resource comprises at least a first protein-ligand complex ensemble representative of the MPS 1000’ and the optimized molecular configuration of the at least one ligand, and a plurality of second protein-ligand complex ensembles representative of the MPS 1000’ and respective sampled alternate conformations of the at least one ligand. This is done to explore the conformational space of the ligand and identify alternative configurations that may also interact favorably with the MPS 1000’. Once the alternate conformations are sampled, the method generates a third digital resource that is representative of a plurality of protein-ligand complex ensembles. This digital resource contains, for example, information about the various conformations of the ligand and their interactions with the MPS 1000’, as well as the energy landscape of the ligand-MPS 1000’ complex. It is noted that in the context of molecular modeling and simulation, the terms "conformation" and "configuration" are often used interchangeably to refer to the spatial arrangement of atoms and bonds within a molecule. Both terms describe the specific three-dimensional shape of a molecule, which can have a significant impact on its biological activity and interactions with other molecules. The terms "conformation" and "configuration" can be used interchangeably because they both refer to the spatial arrangement of a molecule's constituent atoms and bonds. At this point, the first protein-ligand complex ensemble representative of the MPS and the optimized molecular configuration of the at least one ligand, as well as the plurality of second protein-ligand complex ensembles representative of the MPS and respective sampled alternate conformations of the at least one ligand, can be used as digital resources for further analysis and development. These resources provide valuable information about the interaction between the optimized ligand and the MPS, as well as the predicted binding modes and interaction energies between the ligand and the MPS in different conformations. They can be used to identify potential drug candidates or therapeutic agents that can be recommended for the treatment of the disease. Additionally, the resources can be used to design and optimize the properties of the ligand to improve its binding affinity and selectivity for the MPS, and to gain insights into the molecular mechanisms underlying the interaction between the ligand and the MPS. As such, a method for generating protein-ligand complex ensemble representative of the MPS and the optimized molecular configuration of the at least one ligand, as well as the plurality of second protein-ligand complex ensembles representative of the MPS comprises:
[0059] - providing 100 a first digital resource representative for a mutated protein structure, MPS, related to a target protein type;
[0060] - providing 200 a second digital resource representative for an initial molecular configuration of at least one ligand;
[0061] - optimizing 300 the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand based on the provided MPS;
[0062] - sampling 300 a plurality of alternate conformations of the initial molecular configuration of the at least one ligand; and
[0063] - generating 400 a third digital resource representative of a plurality of protein-ligand complex ensembles, the third digital resource comprising at least a first protein-ligand complex ensemble representative of the MPS and the optimized molecular configuration of the at least one ligand, and a plurality of second protein-ligand complex ensembles representative of the MPS and respective sampled alternate conformations of the at least one ligand.
[0064] After generating 400 the first and plurality of second protein-ligand complex ensembles, the method can comprise scoring 500 these ensembles based on a scoring parameter that represents the energy of interaction between the ligand and the MPS 1000’. The energy of interaction is a factor in determining the strength of binding between the ligand and the MPS 1000’ and is used as a measure of the ligand's efficacy as a drug candidate. The scoring parameter may also take into account the degree of similarity between the protein-ligand complex and scientific data, which may include experimental data, structural information, or other relevant information about the MPS 1000’ and the ligand 1300, 1300’. The overall number that represents the energy and degree of similarity between the complex and scientific data is used to rank the ensembles and identify the most promising ligand-MPS 1000’ complexes which may potentially be considered as an effective treatment solution for the mutated disease. Referring to figures 3, 4 and 5 it will be clear that ligand 1300 shown in figures 3 and 4 will not effectively bind with the MPS 1100’
[0065] Finally, the scored first and second protein-ligand complex ensembles and the respective recommendation categories are output 700. By scoring the protein-ligand complex ensembles based on a scoring parameter, and outputting the scored first and second protein-ligand complex ensembles, the method provides the advantage that it allows for the generation of a more comprehensive set of protein-ligand complex ensembles, which can lead to more accurate predictions of the interaction between the ligand and the MPS 1000’. Additionally, by scoring the protein-ligand complex ensembles based on a scoring parameter representative for at least an energy of interaction, the method can identify ligands that are more likely to have a favorable interaction with the MPS 1000’, which can lead to more effective treatment recommendations. Another advantage is that the method allows for the automatic outputting of treatment recommendations based on the results of the method, which can save time and improve the efficiency of the treatment recommendation process. Additionally, by providing recommendation categories, the method can help clinicians to quickly identify the most effective treatment options for their patients. Yet another advantage is that it allows for the testing of the efficacy of drugs with digital tools in tailor-made models of the patient's cancer-driving proteins. By optimizing the molecular configuration of the ligand based on the provided MPS 1000’, the method can generate more accurate predictions of the interaction between the ligand and the MPS 1000’, which can lead to more effective treatment recommendations.
[0066] As shown in figure 2 the computer-implemented method may further comprise classifying 600 the scored first and second protein-ligand complex ensembles into one or more treatment recommendation categories after the step of scoring the first and the plurality of second proteinligand complex ensembles, and outputting 700 the classified first and second protein-ligand complex ensembles and the respective recommendation categories. Classification of the proteinligand complex ensembles into treatment recommendation categories can help to identify the most effective treatment options for individual patients based on the predicted interaction between the ligand and the mutated protein structure. These recommendation categories can be based on various factors, such as the strength of the interaction, the specificity of the ligand for the mutated protein structure, or the predicted effectiveness of the ligand in inhibiting the activity of the mutated protein structure. The recommendation categories can also be based on the selectivity and / or toxicity of the ligand, as well as the predicted impact of the ligand on the function of the mutated protein structure. By outputting the classified protein-ligand complex ensembles and recommendation categories, the method can provide clinicians with a report of the most effective treatment options for their patients based on the predicted interaction between the ligand and the mutated protein structure. This can help to guide treatment decisions and improve the overall effectiveness of cancer treatment. Examples of possible treatment recommendation categories that could be used to classify the scored protein-ligand complex ensembles include but are not limited to: High efficacy: This category could include protein-ligand complex ensembles that show the highest binding affinity and selectivity for the mutated protein structure, and are predicted to have the most favorable impact on the protein's function. These ligands are likely to be the most effective at treating diseases associated with the mutated protein structure. Moderate efficacy: This category could include protein-ligand complex ensembles that show moderate binding affinity and selectivity for the mutated protein structure, and are predicted to have a moderate impact on the protein's function. These ligands may be less effective than those in the high efficacy category, but still have potential as treatment options. Low efficacy: This category could include protein-ligand complex ensembles that show low binding affinity and selectivity for the mutated protein structure, and are predicted to have a minimal impact on the protein's function. These ligands are less likely to be effective as treatment options but may still have potential as starting points for further optimization. High toxicity: This category could include protein-ligand complex ensembles that show high toxicity or adverse effects, such as off-target binding or cellular toxicity. These ligands may be effective at treating the mutated protein structure but are likely to cause significant harm to healthy cells or tissues. Unknown efficacy: This category could include protein-ligand complex ensembles for which there is insufficient data to classify them into one of the other categories. These ligands may require further study or experimental validation before they can be considered as potential treatment options.
[0067] The computer-implemented method may comprise storing the plurality of first model parameters representative for the plurality of protein template structures on a database. This database can be used to store and organize information about the protein template structures, including their sequence, structure, and functional properties. Storing the first model parameters on a database has several advantages. First, it allows for easy access to a large amount of protein structure information, which can be used to select appropriate protein template structures for the alignment and adjustment process. Second, it provides a centralized repository for protein structure information, which can be updated and expanded over time as new structures are discovered and characterized.
[0068] The computer-implemented method may generate the first digital resource representative for the MPS based on a generated fourth digital resource representative for the mutated amino acid sequence, this is illustrated in figure 6. In this case the step of providing 100 a first digital resource representative for a mutated protein structure, MPS, related to a target protein type comprises inputting 110 a first input parameter representative for a reference amino acid sequence of a wild type target protein, such as EGFR, BRAF, HER2, TNF-alpha, p53, ALK, KRAS, MYOSIN, etc., and inputting 120 a second input parameter representative for a mutation in the wild type target protein. The term first input parameter representative for a reference amino acid sequence refers to a variable or input used in the computational method that represents a known amino acid sequence for a particular protein. In figure 6 this is illustrated with the string “AKAR” representing an amino acid sequence. This string represents the sequence of the four amino acids alanine (A), lysine (K), alanine (A), and arginine (R) in a specific order. In Figure 6, K2L represents a mutation in the amino acid sequence represented by the string "AKAR". Specifically, the original lysine (K) at the second position of the sequence has been mutated to a leucine (L), shown by the step 130. This mutation can affect the three-dimensional structure and function of the protein, and may be associated with the development of diseases such as cancer. It will be clear that there are a substantial amount of possible mutations that have been identified and will be identified in the future that can be used as examples, for example in the case of the EGFR wild type protein, there are several mutations that are associated with cancer. Some of the most common EGFR mutations include L858R. This mutation involves a substitution of a leucine (L) for an arginine (R) at position 858 of the EGFR protein. It is commonly found in non-small cell lung cancer (NSCLC) patients and is associated with increased sensitivity to EGFR tyrosine kinase inhibitors (TKIs) such as gefitinib and erlotinib. Exon 19 deletion is a mutation that involves a deletion of a portion of exon 19 of the EGFR gene, resulting in a frameshift and altered protein structure. It is also commonly found in NSCLC patients and is associated with increased sensitivity to EGFR TKIs. T790M is a mutation that involves a substitution of a threonine (T) for a methionine (M) at position 790 of the EGFR protein. It is commonly found in NSCLC patients who have developed resistance to EGFR TKIs and is associated with decreased sensitivity to these drugs. G719X is a mutation that involves a substitution of a glycine (G) for either an alanine (A), serine (S), or cysteine (C) at position 719 of the EGFR protein. It is less common than the other mutations but is still found in a significant number of NSCLC patients and is associated with increased sensitivity to EGFR TKIs. In addition to or as an alternative to a reference amino acid sequence, a genetic code and its transcription can also be used as a first input parameter in the computational method. Where an amino acid sequence is the linear arrangement of amino acids that make up a protein, and determine the protein's structure, function, and interactions with other molecules. The genetic code is the set of rules that determines how the nucleotide sequence of a gene is translated into the amino acid sequence of a protein. In other words, input parameters representative for DNA, RNA, genetic code or derivates thereof corresponding to the MPS 1000’ may also be used as a first input parameter.
[0069] To generate 130 a fourth digital resource representative for a mutated amino acid sequence, the method applies the mutation represented by the second input parameter to the reference amino acid sequence of the wild type target protein represented by the first input parameter. The first digital resource representative for the MPS 1000’ is then generated 140 based on the generated fourth digital resource representative for the mutated amino acid sequence. The mutated amino acid sequence and the MPS are two different things, but they are related to each other. The mutated amino acid sequence refers to a change in the sequence of amino acids in a protein due to a mutation, while the MPS refers to the three-dimensional structure of the protein that is affected by the mutation. In the method described, the fourth digital resource representative for a mutated amino acid sequence is generated 130 by applying the mutation represented by the second input parameter to the reference amino acid sequence of the wild type target protein represented by the first input parameter. This generates a new amino acid sequence that reflects the specific mutation that has occurred in the protein. Once the mutated amino acid sequence has been generated, the first digital resource representative for the MPS can be generated 140 based on this mutated sequence. The MPS represents the three-dimensional structure of the protein that is affected by the mutation, and can be generated using computational methods such as molecular modeling and simulation. The MPS is important because it provides insights into the structural changes that occur in the protein due to the mutation. This allows the method to generate the MPS 1000’ for a specific mutation in a wild type target protein, which can be used to optimize the molecular configuration of the ligand and generate protein-ligand complex ensembles that are tailored to the specific mutation in the target protein. By providing a more accurate representation of the mutated protein structure, this approach improves the accuracy of the treatment recommendations generated by the method and increases the likelihood of successful cancer treatment.
[0070] Generating 140 the first digital resource representative for the MPS 1000’ may further comprise providing a plurality of first model parameters representative for a plurality of protein template structures. Generating the first digital resource representative for the MPS 1000’ refers to the process of creating a digital resource representative of the mutated protein structure that is used in the computer-implemented method. The method described involves generating the first digital resource representative for the MPS 1000' by aligning the mutated amino acid sequence of the mutated target protein with a protein template structure and computationally adjusting the positions and orientations of the particles in the digital representation of the mutated target protein sequence to maximize similarity to those of the identified protein template structure. The plurality of first model parameters are inputs of data that describe a set of protein template structures. These protein template structures may be obtained from a variety of sources, such as the Protein Data Bank, PDB, which is a database of experimentally determined protein structures and models. The plurality of first model parameters representative for a plurality of protein template structures means that each of the first model parameters describes a different protein structure that can be used as a template for the mutated protein structure. To do this, the method provides 141 a plurality of first model parameters representative for a plurality of protein template structures. These model parameters can be obtained from databases or other sources of protein structure information. The method then identifies 142 a protein template structure that corresponds to the wild type target protein, which serves as a reference for the alignment and adjustment process. The step of identifying a protein template structure that corresponds to the wild type target protein can comprise selecting a protein structure from a database or other source of protein structure information that is similar in sequence and / or function to the wild type target protein. The protein structure selected serves as a reference for the alignment and adjustment process, and is used to generate a digital representation of the MPS that is representative of the wild type protein's three- dimensional structure. The selection of the protein template structure can be based on various factors such as sequence identity, structural similarity, and functional relevance. For example, a protein template structure may be selected based on its high sequence identity to the wild type target protein, meaning that its amino acid sequence is very similar to that of the wild type protein. Alternatively, a protein template structure may be selected based on its structural similarity to the wild type protein, meaning that its three-dimensional structure is similar to that of the wild type protein. Once a suitable protein template structure has been identified, it is used as a reference for the alignment 143 and adjustment 144 process, which involves matching the amino acid residues in the mutated amino acid sequence to their corresponding positions in the template structure, and computationally adjusting the positions and orientations of the particles in the digital representation of the mutated target protein sequence to maximize the similarity between the structure of the mutated target protein and that of the identified protein template structure. It will be clear that more than one suitable protein template structure can be identified.
[0071] The mutated amino acid sequence of the mutated target protein is then aligned 143 with the identified protein template structure, which may involve matching the amino acid residues in the sequence to their corresponding positions in the template structure. Once the alignment 143 is complete, the method computationally adjusts 144 the positions and orientations of the particles in the digital representation of the mutated target protein sequence to maximize the similarity between the structure of the mutated target protein and that of the identified protein template structure. In the step of computationally adjusting the positions and orientations of the particles in the digital representation of the mutated target protein sequence, the particles being adjusted can be atoms or molecules. Protein structures are typically represented as a collection of atoms or molecules that are connected by chemical bonds. Each atom or molecule has a specific position and orientation within the protein structure, which determines its overall three-dimensional shape and function. The computational adjustment of these particles involves modifying their positions and orientations in order to maximize the similarity between the structure of the mutated target protein and that of the identified protein template structure.
[0072] The computer-implemented method may further comprise tracking first changes of the protein template structure corresponding to the wild type target protein to the generated MPS 1000’. Tracking first changes of the protein template structure involves monitoring for example the alignment of the mutated amino acid sequence with the identified protein template structure and adjusting the digital representation of the MPS 1000’ accordingly to ensure that it accurately represents the mutated protein structure. This can help to improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment. By tracking changes to the protein template structure, the method can ensure that the digital representation of the MPS 1000’ accurately reflects the mutated protein structure and can be used to optimize the molecular configuration of the ligand and generate protein-ligand complex ensembles that are tailored to the specific mutation in the target protein.
[0073] The computer-implemented method may further comprise outputting a three-dimensional representation of the MPS, not shown. The three-dimensional representation of the MPS provides a valuable tool for visualizing and analyzing the structure of the mutated protein, and can be used to gain insights into the molecular mechanisms underlying diseases associated with the mutation. The representation can also be used to guide the design and development of potential drugs or therapeutic agents that target the mutated protein structure.
[0074] The computer-implemented method may include highlighting, not shown, one or more mutations in the three-dimensional representation of the mutated protein structure. This highlighting can be done using various means such as color coding or labeling, and is used to draw attention to the specific regions of the protein that are affected by the mutation. Highlighting the mutations in the three-dimensional representation of the mutated protein structure can be helpful in several ways. First, it can help to identify the specific regions of the protein that are affected by the mutation, and to understand how these changes may affect the protein's structure and function. Second, it can help to guide the design and development of potential drugs or therapeutic agents that target the mutated protein structure.
[0075] The computer-implemented method may comprise storing the plurality of first model parameters representative for the plurality of protein template structures on a database. Storing the first model parameters on a database can help to improve the efficiency and scalability of the computer- implemented method by allowing for rapid access to a large number of protein template structures. This can facilitate the selection of the most appropriate protein template structure for the mutated protein structure and improve the accuracy of the digital representation of the MPS 1000’. Additionally, storing the first model parameters on a database can allow for the method to be easily updated with new protein template structures as they become available. This can help to ensure that the method remains up-to-date and accurate over time, and can improve the overall effectiveness of cancer treatment.
[0076] The plurality of first model parameters representative for the plurality of protein template structures may comprise at least one of an atomic coordinate of particles in the protein structure, distance lengths between particles, in particular the noncovalent bond lengths, bond lengths and / or angles between particles, torsion angles of amino acid side chains and backbone, solvent accessibility of amino acid residues, hydrogen bonding patterns between amino acid residues, secondary structure information such as alpha-helices and beta-sheets, electrostatic potential maps of the protein structure, energy functions that describe the stability and / or interactions of the protein structure. These first model parameters describe various aspects of the protein structure that can be used as a reference to generate a digital representation of the MPS. For example, atomic coordinates of particles in the protein structure describes the position and orientation of each atom in the protein structure, and can be used to generate a three-dimensional representation of the protein. The distance lengths between particles describes the distances between atoms in the protein structure, and can be used to calculate the noncovalent bond lengths, bond lengths, and / or angles between particles. Torsion angles of amino acid side chains and backbone describes the angles between the amino acid residues in the protein structure, and can be used to predict or describe the protein's three-dimensional structure. Solvent accessibility of amino acid residues describes the extent to which each amino acid residue is exposed to the surrounding solvent, and can be used to predict the protein's stability and interactions with other molecules. Secondary structure information such as alpha-helices and beta-sheets describes the local structure of the protein, and can be used to predict the protein's overall three-dimensional structure. Electrostatic potential maps of the protein structure describes the distribution of charged particles in the protein structure, and can be used to predict the protein's interactions with other charged molecules. Energy functions that describe the stability and / or interactions of the protein structure describes the energy landscape of the protein, and can be used to predict the stability and interactions of the protein with other molecules.
[0077] The computer-implemented method may comprise the step of optimizing the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand using molecular docking. Molecular docking is a computational method that can be used to predict the binding modes and affinities of ligands to protein structures. In the context of the computer-implemented method described, molecular docking can be used to predict the optimal molecular configuration of the ligand that will have the most favorable interaction with the MPS 1000’. In this step, the initial molecular configuration of the ligand is optimized by predicting its most favorable binding modes and affinities to the protein structure. This can be done by computationally docking the ligand into the protein structure and predicting the most energetically favorable binding modes and affinities based on the interaction between the ligand and the protein structure. Molecular docking involves several steps, including the preparation of the ligand and protein structures, the generation of a grid of potential binding sites in the protein structure, the docking of the ligand into the potential binding sites, and the scoring of the resulting ligand-protein complex structures based on their predicted binding affinities. The resulting optimized molecular configuration of the ligand can be used to guide the design and development of potential drugs or therapeutic agents that target the protein structure and treat diseases associated with the mutated protein structure. The molecular docking may comprise generating a digital set of possible ligand configurations and evaluating each ligand configuration of the digital set of possible ligand configuration for the ability to bind to or their strength of interaction with the target protein. By evaluating each ligand configuration in the digital set for the ability to bind to the target protein, the method can identify ligands that are more likely to have a favorable interaction with the MPS 1000’. An example of using molecular docking to optimize the molecular configuration of a ligand targeting EGFR is in the development of drugs for non-small cell lung cancer (NSCLC). EGFR is a receptor protein that is overexpressed in many cases of NSCLC, and is an attractive target for developing drugs that can inhibit its activity and prevent cancer cell growth. One example of a drug that targets EGFR is erlotinib, which is a small molecule inhibitor that binds to the ATP- binding site of the EGFR protein and inhibits its activity. The initial molecular configuration of erlotinib can be optimized using molecular docking to predict its most favorable binding modes and affinities to the EGFR protein. In this process, the three-dimensional structure of the EGFR protein is first obtained using techniques such as X-ray crystallography or NMR spectroscopy. The ligand (erlotinib) is then prepared by generating its three-dimensional structure and optimizing its geometry as described above. Next, a grid of potential binding sites is generated in the EGFR protein structure, and the ligand is docked into these binding sites using molecular docking algorithms. The resulting ligand-protein complex structures are scored based on their predicted binding affinities, and the most favorable binding modes and affinities are selected as the optimized molecular configuration of the ligand. This optimized molecular configuration assists a clinician in selecting drugs comprising said therapeutic agent to improve the treatment.
[0078] Additionally, evaluating each ligand configuration may comprise scoring the energy of interaction between the ligand and the MPS 1000’. Scoring the energy of interaction between the ligand and the MPS 1000’ can help to identify ligands that are more likely to have a favorable interaction with the MPS 1000’. By assigning a score to each ligand configuration based on its energy of interaction, the method can identify ligands that are more likely to be effective in inhibiting the activity of the mutated protein structure. Additionally, by using the energy of interaction as a scoring parameter, the method can identify ligands that are more specific to the mutated protein structure and less likely to interact with other proteins in the body. An example of scoring ligands could be using a molecular mechanics force field to calculate the energy of interaction between the ligand and the MPS 1000'. The force field takes into account various parameters such as bond lengths, angles, and torsions, as well as non-bonded interactions such as van der Waals and electrostatic interactions, as previously described. The energy of interaction between the ligand and the MPS 1000' can be calculated by summing up the energies of all the individual interactions between the ligand and the protein structure. The resulting score can be used to rank the ligands based on their predicted binding affinity to the MPS 1000', with higher scores indicating stronger binding affinities. For example, in recommending treatments against cancer, a set of potential ligands can be evaluated using molecular docking and scored based on their energy of interaction with the mutated protein structure. The ligands can be docked into the binding site of the protein structure and the resulting ligand-protein complex structures can be subjected to molecular mechanics simulations to calculate their energies of interaction. The ligands can then be ranked based on their scores, and the ligands with the highest scores can be selected for the recommendation of treatments.
[0079] The computer-implemented method may comprise the step of optimizing the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand using a molecular dynamics simulation. Molecular dynamics simulation is a computational method that can be used to simulate the movement and interactions of atoms and molecules over time. In the context of the computer-implemented method described in the patent application, molecular dynamics simulation can be used to predict the optimal molecular configuration of the ligand that will have the most favorable interaction with the MPS 1000’. By using molecular dynamics simulation to optimize the molecular configuration of the ligand, the method can generate more accurate predictions of the interaction between the ligand and the mutated protein structure, which can lead to more effective treatment recommendations. Additionally, by optimizing the molecular configuration of the ligand, the method can identify ligands that are more likely to have a favorable interaction with the MPS 1000’, which can increase the likelihood of successful cancer treatment.
[0080] In the molecular dynamics simulation step of the computer-implemented method, a digital set of possible ligand configurations may be generated and then a simulation is run to simulate the motion and interaction of particles in the ligand configurations. This is done by solving one or more equations of motion for each particle in the system. The simulation is used to explore configurations and evaluate each ligand configuration in the digital set for the strength of interaction with the target protein, which can help to identify ligands that are more likely to be effective in inhibiting the activity of the mutated protein structure..
[0081] During the optimization of the initial molecular configuration of the ligand, the method may further comprise tracking second changes of the initial molecular configuration to the optimized molecular configuration of the ligand. This involves monitoring the molecular changes that occur during the optimization process and adjusting the molecular configuration of the ligand accordingly. By tracking these changes, the method can ensure that the optimized molecular configuration accurately reflects the most favorable interaction with the MPS 1000’. This can improve the accuracy of the treatment recommendations generated by the method and increase the likelihood of successful cancer treatment.
[0082] The computer-implemented method may further comprise scoring the first and second changes by evaluating the changes in the molecular interactions between the optimized molecular configuration of the ligand and the MPS 1000’. This scoring process can be based on a variety of factors, such as the energy of interaction between the ligand and the MPS 1000’, the degree of similarity between the protein-ligand complex and scientific data, or other relevant information about the MPS 1000’ and the ligand. By scoring the first and second changes in this way, the method can identify the most promising ligand-MPS 1000’ complexes for further testing and development, and provide treatment recommendations based on the predicted interaction between the ligand and the mutated protein structure.
[0083] During the step of optimizing the molecular configuration of the ligand, the first and second changes may be propagated to ensure that the optimization process accurately reflects the most favorable interaction with the MPS 1000’. Propagating the first and second changes involves ensuring that any changes made to the ligand configuration during the optimization process are consistent with the changes in the MPS 1000’. This can help to ensure that the optimized molecular configuration accurately reflects the most favorable interaction with the mutated protein structure and can improve the accuracy of the treatment recommendations generated by the method.
[0084] The third digital resource representative for the first protein-ligand complex ensemble may include one or more interaction parameters that define the interaction between the ligand and the MPS 1000’ in the respective protein-ligand complex ensemble. These interaction parameters may include, for example, the energy of interaction between the ligand and the MPS 1000’, the specific amino acid residues in the MPS 1000’ that are involved in the interaction, or the type of chemical bond between the ligand and the MPS 1000’. By including these interaction parameters in the third digital resource, the method can provide clinicians with a detailed understanding of the molecular interactions between the ligand and the mutated protein structure. This can help to guide treatment decisions and improve the overall effectiveness of cancer treatment. Interaction parameters that can be included in the third digital resource comprise energy of interaction. The energy of interaction between the ligand and the MPS 1000' can be calculated using molecular mechanics force fields or quantum chemical methods. This parameter provides information about the strength of the interaction between the ligand and the protein structure. Amino acid residues in the MPS 1000' that are involved in the interaction with the ligand can be identified using techniques such as molecular docking or molecular dynamics simulations. This parameter provides information about the location and nature of the binding site. The type of chemical bond between the ligand and the MPS 1000' can be determined using spectroscopic techniques such as X-ray crystallography or NMR spectroscopy. This parameter provides information about the mechanism of action of the ligand and how it inhibits the activity of the mutated protein structure. Hydrogen bonding patterns between the ligand and the MPS 1000' can be identified using techniques such as molecular docking or molecular dynamics simulations. This parameter provides information about the nature and strength of the interactions between the ligand and the protein structure. Solvent accessibility of the amino acid residues in the MPS 1000' that are involved in the interaction with the ligand can be calculated using computational methods. This parameter provides information about the extent to which the residues are exposed to the surrounding solvent and can affect the stability and interactions of the protein structure.
[0085] The third digital resource representative of the plurality of protein-ligand complex ensembles may comprise second model parameters that describe the interaction between the MPS 1000’ and the ligand in the respective protein-ligand complex ensembles. These second model parameters may include, for example, empirical potentials that describe the energy of interaction between the ligand and the MPS 1000’. Empirical potentials are mathematical functions that are used to describe the energy of interaction between two molecules based on their distance and orientation. By using these empirical potentials to describe the interaction between the ligand and the MPS 1000’, the method can generate more accurate predictions of the interaction between the two molecules and identify ligands that are more likely to be effective in inhibiting the activity of the mutated protein structure. Examples of protein-ligand complex ensembles that may be modelled using empirical potentials include EGFR:afatinib and EGFR:gefitinib, which are proteinligand complexes involving EGFR and two tyrosine kinase inhibitors that are used in cancer treatment. Additional examples of protein-ligand complex ensembles include but are not limited to: HIV-1 protease:saquinavir - This is a protein-ligand complex involving the protease enzyme of the human immunodeficiency virus (HIV) and the drug saquinavir. The interaction between the drug and the enzyme is essential for blocking the viral replication cycle, making this complex an important target for drug development. ACE2:remdesivir - This is a protein-ligand complex involving the angiotensin-converting enzyme 2 (ACE2) and the antiviral drug remdesivir. This complex is being studied as a potential treatment for COVID-19, as the drug is believed to interfere with viral replication by binding to the ACE2 receptor on human cells. DNA polymerase: acyclovir - This is a protein-ligand complex involving the DNA polymerase enzyme and the antiviral drug acyclovir. The interaction between the drug and the enzyme is essential for blocking viral replication in herpes simplex virus (HSV) infections. PDE5: sildenafil - This is a protein-ligand complex involving the phosphodiesterase type 5 (PDE5) enzyme and the drug sildenafil (Viagra).
[0086] The method may further comprise applying coarse-graining to the third digital resource representative of the plurality of protein-ligand complex ensembles. This involves grouping the particles or molecules of the molecular configuration of the protein and / or the ligand into one or more moieties, respectively. The second model parameters are then used to describe the interaction between the moieties of the protein and the ligand in the respective protein-ligand complex ensembles. The second model parameters are then used to describe the interaction between the moieties of the protein and the ligand in the respective protein-ligand complex ensembles. By using coarse-graining to simplify the molecular models, the method can reduce the computational complexity of the simulation and make it more efficient. This can facilitate the generation of a larger number of protein-ligand complex ensembles and improve the accuracy of the treatment recommendations generated by the method. Additionally, coarse-graining can help to identify important regions of the protein-ligand interface and provide insights into the mechanism of action of the ligand. An example of coarse-graining in the context of protein-ligand complexes can be illustrated as a protein-ligand complex involving a protein structure with 10,000 atoms and a ligand with 1,000 atoms. This complex would be computationally expensive to simulate using molecular dynamics or other methods, due to the large number of particles involved. To simplify the model and reduce computational complexity, the method uses coarse-graining to group the atoms into moieties.
[0087] For example, the protein atoms are grouped into moieties representing the backbone, amino acid side chains, and distinct chemical moieties, and the ligand atoms are grouped into moieties representing the core, peripheral groups, and functional groups. Then the method uses second model parameters such as empirical potentials to describe the interaction between the moieties of the protein and the ligand. By using this coarse-grained model, the number of particles in the simulation is reduced and makes it more computationally efficient. This allows to generate a larger number of protein-ligand complex ensembles and improve the accuracy of the treatment recommendations generated by the method. Additionally, by identifying important regions of the protein-ligand interface using coarse-graining, we can gain insights into the mechanism of action of the ligand and guide the treatment selection of more effective drugs or therapeutic agents.
[0088] A person of skill in the art would readily recognize that steps of various above-described methods can be performed by programmed computers. Herein, some embodiments are also intended to cover program storage devices, e.g., digital data storage media, which are machine or computer readable and encode machine-executable or computer-executable programs of instructions, wherein said instructions perform some or all of the steps of said above-described methods. The program storage devices may be, e.g., digital memories, magnetic storage media such as a magnetic disks and magnetic tapes, hard drives, or optically readable digital data storage media. The program storage devices may be resident program storage devices or may be removable program storage devices, such as smart cards. The embodiments are also intended to cover computers programmed to perform said steps of the above-described methods.
[0089] The description and drawings merely illustrate the principles of the present invention. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the present invention and are included within its scope. Furthermore, all examples recited herein are principally intended expressly to be only for pedagogical purposes to aid the reader in understanding the principles of the present invention and the concepts contributed by the inventor(s) to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass equivalents thereof.
[0090] The functions of the various elements shown in the figures, including any functional blocks labelled as “processors”, may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and non-volatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0091] It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the present invention. Similarly, it will be appreciated that any flowcharts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer.
[0092] It should be noted that the above-mentioned embodiments illustrate rather than limit the present invention and that those skilled in the art will be able to design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word “comprising” does not exclude the presence of elements or steps not listed in a claim. The word “a” or “an” preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware comprising several distinct elements and by means of a suitably programmed computer. In claims enumerating several means, several of these means can be embodied by one and the same item of hardware. The usage of the words “first”, “second”, “third”, etc. does not indicate any ordering or priority. These words are to be interpreted as names used for convenience.
[0093] In the present invention, expressions such as “comprise”, “include”, “have”, “may comprise”, “may include”, or “may have” indicate existence of corresponding features but do not exclude existence of additional features.
[0094] Whilst the principles of the present invention have been set out above in connection with specific embodiments, it is to be understood that this description is merely made by way of example and not as a limitation of the scope of protection which is determined by the appended claims.
Claims
CLAIMS1. A computer-implemented method of treatment recommendation, the method comprising:- providing (100) a first digital resource representative for a mutated protein structure, MPS, related to a target protein type;- providing (200) a second digital resource representative for an initial molecular configuration of at least one ligand;- optimizing (300) the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand based on the provided MPS;- sampling (300) a plurality of alternate conformations of the initial molecular configuration of the at least one ligand;- generating (400) a third digital resource representative of a plurality of protein-ligand complex ensembles, the third digital resource comprising at least- a first protein-ligand complex ensemble representative of the MPS and the optimized molecular configuration of the at least one ligand, and- a plurality of second protein-ligand complex ensembles representative of the MPS and respective sampled alternate conformations of the at least one ligand;- scoring (500) the first and the plurality of second protein-ligand complex ensembles based on a scoring parameter representative for at least an energy of interaction between the ligand and the MPS;- outputting (700) the scored first and second protein-ligand complex ensembles.
2. The computer-implemented method of claim 1 , further comprising classifying (600) the scored first and second protein-ligand complex ensembles into one or more treatment recommendation categories after the step of scoring the first and the plurality of second protein-ligand complex ensembles; and outputting (700) the classified first and second protein-ligand complex ensembles and the respective recommendation categories3. The computer-implemented method of claim 1 or 2, wherein providing (100) the first digital resource representative of the MPS comprises:- inputting (110) a first input parameter representative for a reference amino acid sequence of a wild type target protein such as EGFR, BRAF, HER2, TNF-alpha, p53, Aik, Kras, MYOSIN etc;- inputting (120) a second input parameter representative for a mutation in the wild type target protein;- generate (130) a fourth digital resource representative for a mutated amino acid sequence by applying the mutation represented by the second input parameter to the reference amino acid sequence of the wild type target protein represented by the first input parameter;- generate (140) the first digital resource representative for the MPS based on the generated fourth digital resource representative for the mutated amino acid sequence.
4. The computer-implemented method of claim 3, wherein generating (140) the first digital resource representative for the MPS comprises:- providing (141) a plurality of first model parameters representative for a plurality of protein template structures;- identifying (142) a protein template structure corresponding to the wild type target protein;- aligning (143) the mutated amino acid sequence of the mutated target protein with the identified protein template structure;- generating (144) the first digital resource representative for the MPS by computationally adjusting the positions and orientations of the particles in the digital representation of the mutated target protein sequence to maximize similarity to those of the identified protein template structure.
5. The computer-implemented method of the previous claim, further comprising tracking first changes of the protein template structure corresponding to the wild type target protein to the generated MPS.
6. The computer-implemented method according to any one of the previous claims 4-5, further comprising outputting a three-dimensional representation of the mutated protein structure.
7. The computer-implemented method of the previous claim, wherein the one or more mutations in the three-dimensional representation of the mutated protein structure are highlighted.
8. The computer-implemented method according to any one of the previous claims 4-7, wherein the plurality of first model parameters representative for the plurality of protein template structures are stored on a database.
9. The computer-implemented method according to any one of the previous claims 4-8, wherein the plurality of first model parameters representative for the plurality of protein template structures comprise at least one of an atomic coordinate of particles in the protein structure, distance lengths between particles, bond lengths and / or angles between particles, torsion angles of amino acid sidechains and backbone, solvent accessibility of amino acid residues, hydrogen bonding patterns between amino acid residues, secondary structure information such as alpha-helices and betasheets, electrostatic potential maps of the protein structure, energy functions that describe the stability and / or interactions of the protein structure.
10. The computer-implemented method according to any one of the previous claims, wherein the step of optimizing (300) the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand is performed using molecular docking.
11. The computer-implemented method of the previous claim, wherein the molecular docking comprises generating a digital set of possible ligand configurations and evaluating each ligand configuration of the digital set of possible ligand configuration for the ability to bind to the target protein.
12. The computer-implemented method of the previous claim, wherein the step of evaluating each ligand configuration is performed by scoring the energy of interaction between the ligand and the MPS.
13. The computer-implemented method according to any one of the previous claims, wherein the step of optimizing (300) the initial molecular configuration of the at least one ligand into an optimized molecular configuration of the at least one ligand is performed using a molecular dynamics simulation.
14. The computer-implemented method of the previous claim, wherein the molecular dynamics simulation comprises generating a digital set of possible ligand configurations, simulating a motion and interaction of particles, such as particles and molecules, of the ligand configurations by solving one or more equations of motion for each particle and evaluating each ligand configuration of the digital set of possible ligand configuration for the strength of interaction with the target protein.
15. The computer-implemented method according to any one of the previous claims, wherein during the optimizing (300) of the initial molecular configuration of the at least one ligand, the method further comprises tracking second changes of the initial molecular configuration of the at least one ligand to the optimized molecular configuration of the at least one ligand.
16. The computer-implemented method of claims 4 and 14, further comprising scoring the first and second changes by evaluating the changes of the molecular interactions between the optimizedmolecular configuration of the at least one ligand and the MPS.
17. The computer-implemented method of claims 4 and 14, wherein the first and second changes are propagated during the step of optimizing.
18. The computer-implemented method according to any one of the previous claims, wherein the third digital resource representative for the first protein-ligand complex ensemble comprises one or more interaction parameters which define at least an interaction between the ligand and the MPS in the respective protein-ligand complex ensemble.
19. The computer-implemented method according to any one of the previous claims, wherein the third digital resource representative of the plurality of protein-ligand complex ensembles comprises second model parameters describing at least an interaction between the MPS and the ligand comprised by the protein-ligand complex ensemble, e.g. EGFR:afatinib, EGFR:gefitinib, wherein the interaction is modelled using empirical potentials.
20. The computer-implemented method of the previous claim, further comprising applying coarse graining to the third digital resource representative of the plurality of protein-ligand complex ensembles such that particles or molecules of the molecular configuration of the protein and / or the ligand are grouped into one or more moieties, respectively; and wherein the second model parameters describe at least an interaction between the moieties of the protein and the ligand comprised by the protein-ligand complex ensemble.
21. The computer-implemented method according to any one of the previous claims, further comprising outputting one or more treatment recommendations based on the scored first and second protein-ligand complex ensembles.
22. A computer program product comprising a computer-executable program of instructions for performing, when executed on a computer, the steps of the method of any one of the previous claims.
Citation Information
Patent Citations
Ligand searching device, ligand searching method, program, and recording medium
US20070166760A1