Structural Analysis System and Modeling Method for Protein Degradation Targeting Chimera Complexes

By providing structural analysis systems and modeling methods for protein degradation targeted chimeric complexes, the problems of insufficient accuracy and low availability of structural modeling results in the prior art are solved, and more efficient protein degradation drug design and optimization are achieved.

CN119811534BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510283070.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing structural modeling methods for protein degradation targeted chimera complexes have problems such as insufficient scoring performance, low accuracy of results and poor availability.

Method used

A structural analysis system is provided, including a chimera component analysis module, a chimera-induced complex component analysis module and a linker length distribution analysis module. These modules are used to perform component analysis and conformational search of chimera and ternary complexes, combining energy minimization and random sampling and screening conformational fusion methods to optimize the conformational structure to improve the accuracy of modeling results.

Benefits of technology

It significantly improves the design efficiency and success rate of protein degradation targeted chimera, deeply understands the interaction mechanism between chimera and target protein and ubiquitin ligase, and accelerates the discovery and optimization of new protein degradation drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811534B_ABST
    Figure CN119811534B_ABST
Patent Text Reader

Abstract

The present invention discloses a structural analysis system and a modeling method for proteolysis-targeting chimera complexes, belonging to the field of computer-aided drug molecule design. The structural analysis system includes: a chimera component analysis module, a chimera-induced complex component analysis module, and a linker length distribution analysis module of the chimera to achieve automatic splitting and analysis of the structure of proteolysis-targeting chimera drugs. Based on the component analysis results, through protein-protein docking, conformation search, conformation merging and optimization, conformation clustering, and conformation scoring and ranking, efficient and accurate structure modeling is achieved. The present invention can perform automatic component splitting and structural analysis, so as to comprehensively explore potential drug molecule binding modes, significantly reduce the trial-and-error cost and time of traditional drug structure analysis, significantly improve the design efficiency and success rate of proteolysis-targeting chimeras, and ensure the efficiency and reliability of the structure modeling process through automatic parameter optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer-aided drug molecular design, and particularly to a structural analysis system and modeling method for proteolysis-targeting chimera complexes. Background Art

[0002] Proteolysis-targeting chimeras (PROTACs) are a new type of drug that can achieve the degradation of target proteins (hereinafter referred to as chimeras), which greatly expands the range of targets that drugs can modulate. These drugs can bind to ubiquitin ligases and target proteins simultaneously, thereby inducing the ubiquitination of target proteins and ultimately being recognized and degraded by the proteasome.

[0003] Due to the unique mechanism of these drugs, they have several advantages compared to traditional small molecule inhibitors, including: complete inhibition of function, not limited to the function of enzymes or specific positions on the target; long-term action proportional to protein turnover rate and enhanced selectivity, and since a single proteolysis-targeting chimera molecule can degrade multiple copies of the target, it greatly expands the range of drug targets, so it has become a hot spot in new drug research and development.

[0004] However, the research and development of these drugs is difficult, and their complex mechanism of action makes it difficult to carry out molecular design based on experience or a single protein structure. Therefore, the research and development process of these drugs usually involves large-scale random chemical structure modification, synthesis, and bioactivity detection of drugs. Therefore, a rational design method suitable for this type of drug is urgently needed.

[0005] For example, the invention application with the publication number CN118711703A discloses a PROTAC molecular drug design method based on the DEL platform. Based on the DNA Encoded Library (DEL), bifunctional molecules are quickly generated and evaluated through a molecular generation model and computer-aided drug design, which accelerates the construction of the PROTAC-like DEL bifunctional molecule library. Then, by modeling and simulating the protein-bifunctional molecule ternary complex system mediated by the bifunctional molecule, the binding mode of the bifunctional molecules in the PROTAC-like DEL bifunctional molecule library is predicted, and the bifunctional molecules are further optimized to make the designed bifunctional molecules more reasonable and reduce the synthesis difficulty of the molecules.

[0006] For another example, Reference 2 (Comprehensive Modeling of PROTAC-Mediated Ternary Complexes) discloses a comprehensive modeling method based on FRODOCK and RosettaDock. Through the effective sampling algorithm with coordinate constraints provided by FRODOCK, preliminary structure docking is performed, and each docking structure is optimized in combination with RosettaDock to predict the structure of PROTAC-mediated ternary complexes. This comprehensive modeling method models the ternary complex starting from the unbound structure, making the ternary complex close to the native structure, which is helpful for the design and optimization of new PROTACs.

[0007] However, in the existing work, there are some limitations in the structure modeling methods for chimeric-induced ternary complexes. These structure modeling methods do not provide reliable analysis of chimeras and the components of chimeric-induced ternary complexes, resulting in the inability of the structure modeling results to well approximate the structure of the native ternary complex. Moreover, the structure modeling prediction lacks a reasonable scoring and ranking criterion, leading to insufficient accuracy and low usability of the final modeling results. Summary of the Invention

[0008] The object of the present invention is to provide a structure analysis system and modeling method for proteolysis-targeting chimeric complexes, aiming to solve the problems of insufficient accuracy of the existing modeling scoring performance and structure modeling results and low usability.

[0009] To achieve the above object of the invention, an embodiment provides a structure analysis system for proteolysis-targeting chimeric complexes, including: a chimera component analysis module, a chimeric-induced complex component analysis module, and a linker length distribution analysis module for chimeras;

[0010] The chimera component analysis module is used to identify a series of structural fragments in the chimera molecule, merge adjacent structural fragments by an iterative method; sort each merged fragment, test the rationality of the connection between the target protein ligand and the ubiquitin ligase ligand and the sorted merged fragment, and regard the reasonably sorted merged fragment as the linker part to achieve the splitting and component analysis of the chimera;

[0011] The chimeric-induced complex component analysis module is used to split the chimeric-induced ternary complex, and analyze, identify, and screen the target protein, chimera, and ubiquitin ligase after splitting to achieve the component analysis of the chimeric-induced complex;

[0012] The linker length distribution analysis module for chimeras is used to perform molecular conformation search on the linker part extracted from the chimera in the complex, analyze the distance distribution of the linker endpoints when the chimera is in different three-dimensional conformations, and use it as an index of the linker length range for the structure analysis of proteolysis-targeting chimeric complexes.

[0013] In one embodiment, the identification of a series of structural fragments in the chimeric molecule includes: identifying structural fragments based on a molecular structure pattern or a ring matching pattern;

[0014] The structural fragments matching the molecular structure pattern include: primary amide, thioamide, a carbon atom connected to two hydrogen atoms, hydroxymethyl group, alkyne group, a substructure formed by a ring nitrogen atom connected to an acyclic carbonyl group, or an azo group formed by two nitrogen atoms connected by a double bond;

[0015] The structural fragments of the ring matching pattern include: a six-membered carbon aromatic ring with two substituents, an aliphatic ring system, or a three- to six-membered aromatic azacyclic or oxacyclic ring with two substituents, wherein the aliphatic ring system includes a monocyclic heterocyclic ring containing nitrogen and oxygen and a spiro ring system, and these ring systems must have exactly two substituents. The monocyclic heterocyclic ring is defined as a three- to six-membered aliphatic ring, and the spiro ring can be one of six-membered - six-membered, six-membered - five-membered, or five-membered - five-membered. In the three- to six-membered aromatic azacyclic or oxacyclic ring with two substituents, at most one substituent can be connected to a peptide bond.

[0016] In one embodiment, the adjacent structural fragments are represented as being capable of being connected by a single chemical bond or having partially overlapping chemical structures.

[0017] In one embodiment, the rationality of the connection between the test target protein ligand and the ubiquitin ligase ligand to the linker part includes: if the corresponding ligand target protein ligand and ubiquitin ligase ligand do not meet the conditions, then the fusion fragment is discarded, and the next fusion fragment in the sorting is selected for testing. If the fusion fragment meets the rationality requirements of the corresponding ligand, it is regarded as the linker part of the chimeric molecule, where the conditions are at least one of corresponding to zero or one ligand, the number of atoms of the ligand being less than 5, and the ligand not containing any ring structure.

[0018] In one embodiment, the analysis, identification, and screening of the split target protein, chimeric molecule, and ubiquitin ligase include: for the chimeric molecule part, identifying each non-protein molecule through non-standard amino acid residues, limiting the number of atoms in the non-protein molecule to less than 150 and not being in the form of a polypeptide, sorting them in descending order of the number of atoms, screening these non-protein molecules based on specific characteristics, and selecting the molecule with the largest number of atoms after screening, where the specific characteristics are that the ligands on both sides of the chimeric molecule interact with different protein chains in the target protein and ubiquitin ligase respectively; if no satisfactory ligand is found or all ligands are screened out, then more relaxed conditions are used to re-identify and screen the ligands to achieve the identification and screening of the chimeric molecule part, where the more relaxed conditions are that the number of atoms in the non-protein molecule is less than 300 and it can be in the form of a polypeptide;

[0019] Based on the identification of the chimeric part, two main protein chains that interact with the target protein and the ubiquitin ligase are identified. The protein chains that interact with these two main protein chains are recursively collected, and finally a set of protein chains on both sides is formed, realizing the identification and screening of the target protein and the ubiquitin ligase, and completing the analysis, identification and screening of the split target protein, chimeric body and ubiquitin ligase.

[0020] The present invention also provides a structural modeling method for a protein degradation targeting chimeric complex. The structural modeling method applies the above-mentioned structural analysis system and includes the following steps:

[0021] Based on the chemical structures of the target protein, target protein ligand, ubiquitin ligase protein, ubiquitin ligase protein ligand and linker as input, combined with the linker length distribution analysis module of the chimeric body, the relative positions of the target protein and the ubiquitin ligase protein are adjusted for protein-protein docking, simulating the interaction between proteins to form a set of protein-protein docking conformations;

[0022] Based on the chimeric body component analysis module, a series of three-dimensional structural conformations of the chimeric body to be modeled are generated, and the conformations of the chimeric body are searched and analyzed based on the conformation search strategy to form a set of chimeric body conformations;

[0023] Based on the set of protein-protein docking conformations and the set of chimeric body conformations, they are matched and merged in a conformation merging manner of energy minimization, random sampling and screening, and the results of the matching and merging are optimized by means of fine docking of conformation structures and energy optimization to form a set of chimeric body-induced ternary complex conformations;

[0024] Cluster the set of chimeric body-induced ternary complex conformations, select specific members in each category to form a structure set, combine the chimeric body-induced complex component analysis module, construct a conformation scoring and sorting method, and sort each structure in the result set to complete the structural modeling of the protein degradation targeting chimeric complex.

[0025] In one embodiment, the adjustment of the relative positions of the target protein and the ubiquitin ligase protein for protein-protein docking includes: regarding the protein with a larger number of atoms as the fixed protein, keeping the positions of the backbone atoms unchanged during the docking process, and regarding the protein with a smaller number of atoms as the mobile protein. During the docking process, a conformational search for protein-protein docking is performed through multiple Monte Carlo steps to achieve the lowest total energy of the system and complete the protein-protein docking.

[0026] Optionally, multiple Monte Carlo steps are to randomly select to translate the mobile protein by 0 Å to 0.1 Å or rotate the mobile protein by 0° to 5°, and the rotation direction can be one of the X, Y, and Z axes.

[0027] Optionally, distance constraints are imposed in the conformational search for protein-protein docking. The distance constraint is defined as an energy penalty value for each Monte Carlo step and is represented by a Gaussian function with distance as the independent variable. Based on the linker length distribution analysis module of the chimera, the mean and standard deviation of the linker length distribution are calculated to reduce the search space range.

[0028] Optionally, in one embodiment, the conformational search strategy includes at least one of: performing a conformational search on the complete chimera, performing a conformational search on the linker part, and performing a conformational search on the ligand part of the linker.

[0029] In one embodiment, the energy minimization includes: extracting conformations from the protein-protein docking conformation set and analyzing the protein ligand positions, generating arbitrary conformations of the chimera and placing them in relative positions, and performing a two-step energy minimization process. Among them, the two-step energy minimization process includes: Step 1, ignoring the protein part in the protein-protein docking conformation for energy minimization, and at the same time pulling the ligand in the chimera towards the ligand position in the protein-protein docking conformation; Step 2, re-adding the protein part in the protein-protein docking conformation for energy minimization.

[0030] Optionally, Step 1, ignoring the protein part in the protein-protein docking conformation for energy minimization, includes: ignoring the protein part in the protein-protein docking conformation, only retaining the ligand part, calculating the largest common substructure between each ligand and the chimera to form a mapping relationship; achieving energy minimization through multiple Monte Carlo processes, and minimizing the distance of the mapped atom pairs in the form of an energy penalty value during the process; after energy minimization, calculating the root mean square distance deviation between the chimera and the ligand and setting a threshold for conformation screening.

[0031] Optionally, Step 2, re-adding the protein part in the protein-protein docking conformation for energy minimization, includes: re-adding the protein part in the protein-protein docking conformation and ignoring the ligand part in the protein-protein docking conformation, calculating the position conflict between the protein and the chimera and setting a threshold for screening, and achieving iterative optimization of the position through multiple Monte Carlo processes to achieve conformation screening.

[0032] Further optionally, the position conflict is represented by calculating the overlapping volume of the van der Waals volumes between the protein and the chimera.

[0033] Further optionally, distance constraints for the chimera, pocket atom set, and interface atom set are imposed during the Monte Carlo process. Among them, the pocket atom set is defined as all residue atoms within 5 Å of the chimera molecule, and the interface atom set is defined as between any two protein chains A and B, where there are some protein side chain atoms such that they are located on chain A and there are atoms on chain B within 10 Å.

[0034] Optionally, the iterative optimization of the position by minimizing energy through multiple Monte Carlo processes includes: minimizing energy through multiple Monte Carlo processes, checking the overlapping volume of the van der Waals volumes, and if it is higher than the set threshold, continuing the energy minimization process; otherwise, outputting the merged structure.

[0035] In one embodiment, the random sampling and screening include: pre-searching for the conformations of the chimeras, extracting conformations from the protein-protein docking conformation set, superimposing the conformations of the ligand and each chimera, calculating the root mean square distance deviation of the corresponding atoms between the ligand and the chimera, setting a threshold for preliminary conformation screening; calculating the position conflicts between the protein and the chimera, setting a threshold for secondary screening; and achieving energy minimization convergence through multiple Monte Carlo processes to perform conformation screening.

[0036] Optionally, the setting of the threshold for preliminary conformation screening includes: setting a threshold, and if the root mean square distance deviation of the atoms exceeds this threshold, discarding the corresponding conformation.

[0037] Optionally, the calculation of the position conflicts between the protein and the chimera and the setting of the threshold for secondary screening include: representing the position conflicts by calculating the overlapping volume of the van der Waals volumes between the protein and the chimera, and setting a threshold for conformation screening.

[0038] Further optionally, distance constraints are imposed on the chimera, the pocket atom set, and the interface atom set during the Monte Carlo process, and the intensity of the distance constraints is differentially set for different atom sets.

[0039] Optionally, the achieving of energy minimization convergence through multiple Monte Carlo processes includes: minimizing energy through multiple Monte Carlo processes, checking the overlapping volume of the van der Waals volumes, and if the reduction amount of the overlapping volume compared to that after the previous energy minimization is less than 20%, it is considered that the minimization process has converged, and the merged structure is output. When the minimization converges, if the overlapping volume is higher than 10 cubic angstroms, the corresponding conformation is screened out.

[0040] In one embodiment, the conformation scoring is calculated using multiple metrics or scores of the conformation, including one or more of the following: the number of atoms of the chimera, the number of atoms of the linker, the number of movable dihedral angles of the linker, the mean and variance calculated by fitting the length distribution of the linker with a Gaussian function, the docking score of the chimera, the total energy of the system after energy minimization, the pairwise binding energy components between the chimera and the two proteins, the number of member conformations of the conformation class where it is located, the ranking from the central conformation, the protein-protein interaction interface energy, the difference in solvent accessible surface area between the linker in the dissociated state and the bound state, and the strain energy of the drug conformation of the chimera.

[0041] Compared with the prior art, the beneficial effects of the present invention at least include:

[0042] (1) By integrating chimeric component analysis, ternary complex analysis, and linker length analysis, this system can comprehensively evaluate the structural characteristics of chimeras, significantly improve the design efficiency and success rate of proteolysis-targeting chimeras, and contribute to a deeper understanding of the interaction mechanisms between chimeras and target proteins and ubiquitin ligases, accelerating the discovery and optimization of novel proteolysis drugs and promoting the development of targeted proteolysis therapies.

[0043] (2) Based on the structure modeling method provided by the structure analysis system, the efficient structure modeling of the proteolysis-targeting chimera complex is achieved through steps such as protein-protein docking, conformational search, conformational merging and optimization, conformational clustering, and conformational scoring and ranking. It is accurate, efficient, and practical, and is of great significance for drug research and development and the study of proteolysis mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0045] Figure 1 It is a schematic structural diagram of the structure analysis system for proteolysis-targeting chimera complexes provided by the present invention.

[0046] Figure 2 It is a schematic diagram of a proteolysis-targeting chimera and the ternary complex and sub-structures induced thereby.

[0047] Figure 3 It is a schematic structural diagram of the chimeric component analysis module provided by the present invention.

[0048] Figure 4 It is a schematic structural diagram of the chimeric-induced complex component analysis module provided by the present invention.

[0049] Figure 5 It is a flowchart of the structure modeling method for proteolysis-targeting chimera complexes provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further elaborates on the present invention with reference to the

[0051] embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0052] To improve the accuracy of structural modeling of chimeric-induced ternary complexes, the embodiments provide a structural analysis system for proteolysis-targeting chimeric complexes, which is used to process the structural data of a large number of proteolysis-targeting chimeric molecules and their complexes, including splitting, property analysis, structure generation, etc., and is used in relevant modeling processes or is used alone or called by other computer systems for tasks such as data analysis. As Figure 1 shown, the structural analysis system includes: a chimeric component analysis module, a chimeric-induced complex component analysis module, and a linker length distribution analysis module for chimeras. Each part of the structural analysis system will be introduced in detail below.

[0053] First, the embodiments provide a chimeric component analysis module. The splitting of chimeric components mainly refers to analyzing the chemical structure of the chimera and automatically identifying and selecting its respective substructures (target protein ligand, ubiquitin ligase ligand, linker). As Figure 2 shown, the proteolysis-targeting chimera 100 usually consists of three parts: a target protein ligand part 101, a ubiquitin ligase ligand part 102, and a linker part 103 connecting the two. In drug design, the target protein ligand part 101 and the ubiquitin ligase ligand part 102 usually select known binder structures.

[0054] Therefore, after these two ligands are assembled into a chimera, the three-dimensional structure similar to the original binder is still retained in the induced ternary complex. In contrast, the linker part 103 usually has high flexibility, and the differences in its length and composition may significantly affect the assembly mode of the ternary complex. Therefore, an automated component analysis method needs to be developed to parse the chemical structure of the chimera into three substructure parts: target protein ligand, ubiquitin ligase ligand, and linker, in order to implement a differential processing strategy in the complex modeling process.

[0055] As Figure 3 shown, a series of structural fragments 301 are first identified in the chimeric component analysis module. Specifically, first, the following substructures 302 are matched based on the molecular structure pattern: primary amide, or thio-primary amide, or a carbon atom connected to two hydrogen atoms, or hydroxymethyl group, or alkyne group, or a substructure formed by a cyclic nitrogen atom connected to an acyclic carbonyl, or an azo group formed by two nitrogen atoms connected by a double bond. The above substructures can be represented by the SMARTS chemical structure search language as "[NH]C=O", "[NH]C=S", "[H2&C]", "[H2&C]O", "C#C", "[R1&N][R0&C]=O", "[N]=[N]".

[0056] Second, based on the ring matching pattern, match the following substructures 303: a six-membered carbon aromatic ring with two substituents, or a three- to six-membered aromatic azacyclic or oxacyclic ring with two substituents, where at most only one substituent is connected to the peptide bond; based on the ring matching pattern, match the following substructure 304: an aliphatic ring system, including monocyclic heterocyclic rings and spiro ring systems containing nitrogen and oxygen, and these ring systems must have exactly two substituents. The monocyclic heterocyclic ring is defined as a three- to six-membered aliphatic ring, and the spiro ring can be one of the following: six-membered - six-membered, six-membered - five-membered, five-membered - five-membered.

[0057] Subsequently, an iterative method is used to merge adjacent structural fragments. The adjacent structural fragments refer to those connected by a single chemical bond or having partial chemical structure overlap. If the outermost fragment in the merged fragment is an aromatic ring, then this ring is removed from the merged fragment. Subsequently, each merged fragment 305 is sorted from largest to smallest according to the approximate bond number length index, which can be calculated by the number of ring atoms and non-ring atoms in the fragment: approximate bond number length index = 0.5 × number of ring atoms + number of non-ring atoms. This index is used to quickly indicate the approximate length when the linker is fully extended. This bond number length is only used to quickly determine the approximate length of the candidate linker part during the splitting of the chimera components, and is different from the linker length distribution described later.

[0058] Next, a rationality test is performed on each sorted merged fragment 305 in turn. Specifically, it is to test whether the target protein ligand part 101 and the ubiquitin ligase ligand part 102 in the chimera meet specific requirements when the merged fragment 305 is regarded as the linker part.

[0059] Specifically, first identify two anchor atoms 306 in the merged fragment 305 that covalently connect to atoms outside the fragment, and then extract the atom sets that can be extended by covalent bonds starting from these two anchor atoms respectively. For each anchor atom 306, the extracted atom set is further split according to connectivity, and the larger atom set is selected as the ligand 307 on this side, that is, the target protein ligand or the ubiquitin ligase ligand. During the process of taking each merged fragment 305 as the linker 103 for testing, if the corresponding ligand 307 meets one of the following conditions, then this merged fragment is discarded, and the next merged fragment in the sorting is selected for testing: corresponding to zero or one ligand only, or the number of atoms in the ligand is less than 5, or the ligand does not contain any ring structure. If a merged fragment 305 meets the rationality requirements of the corresponding ligand, it is regarded as the linker part of the chimera.

[0060] To verify the effectiveness of the chimeric component analysis module in automatically splitting chimeras, it was compared with the chimeric structure and sub-structure information manually collected in the existing public database ProtacDB, and the error items with large differences in the number of sub-structure atoms and the failed items that did not recognize some sub-structures were analyzed to quantitatively analyze the applicability of the above chimeric component analysis module to diverse chimeric drug molecular structures. As shown in Table 1, the generalizability of the automatic chimera splitting method provided by the chimeric component analysis module was interpreted, and by deleting some key steps or modifying some key parameters, the key nature of each step and parameter was demonstrated.

[0061] Table 1

[0062]

[0063] In the embodiment, a chimeric induced complex component analysis module was provided. The splitting of the chimeric ternary complex components mainly refers to analyzing the three-dimensional structure of the ternary complex induced by the chimeric drug and automatically identifying and selecting its respective sub-structures (target protein, chimera, ubiquitin ligase).

[0064] As Figure 2 shown, the protein degradation-targeting chimera ternary complex 200 is formed by the mutual binding of the target protein 201, the chimeric drug 100, and the ubiquitin ligase 202. Among them, both the target protein 201 and the ubiquitin ligase 202 may participate in the form of multi-protein complexes and contain multiple protein chains, making it difficult to split them only according to the protein chains in the structure. For example, the von Hippel-Lindau (VHL) E3 ubiquitin ligase usually functions in the form of the VHL-ElongB-ElangC complex. Therefore, an automated component analysis method needs to be developed to split the three-dimensional molecular structure model of the chimeric ternary complex into three sub-structure parts: target protein, chimeric drug, and ubiquitin ligase. Each part may contain several molecules and protein chains to implement its differential processing strategy in the complex modeling process.

[0065] As Figure 4 shown, in the chimeric induced complex component analysis module, identifying the chimera part specifically involves, first, identifying each non-protein molecule through non-standard amino acid residues, restricting the number of atoms in the chimera molecule to less than 150, and the chimera molecule not being in the form of a polypeptide, and sorting them in descending order of the number of atoms (301). Second, screening these chimera molecules based on a specific feature and selecting the chimera molecule with the largest number of atoms after screening. The specific feature is that the ligands on both sides of the chimera molecule will interact with different protein chains in the target protein and the ubiquitin ligase respectively.

[0066] To quickly calculate this specific feature, each atom and bond in the chimeric molecule needs to be represented in the form of an undirected graph, and the minimum eccentricity algorithm is used to calculate the central atom 313 and the set of peripheral atoms of the chimeric molecule. Subsequently, the shortest path lengths between all pairs of peripheral atom nodes are calculated, and the pair of nodes corresponding to the longest length is selected as the terminal atom pair 314. Next, the shortest path lengths between the central atom and the two terminal atoms a and b are calculated, denoted as and . Subsequently, for each atom in the molecule, the shortest path lengths from it to the terminal atoms a and b are calculated, denoted as and . For each atom, the following formula can be used to calculate which end side of the molecule the atom biases towards:

[0067] ,

[0068] According to the positive or negative value of, the atoms in the chimeric molecule can be split into a set A that biases towards the terminal atom a and a set B that biases towards the terminal atom b. Subsequently, for each set, all non-drug atoms within 4 angstroms of the atoms in the set are counted, and the protein chain to which the most non-drug atoms belong is counted as the main interacting protein chain of this end-side atom set. If the main interacting protein chains of these two atom sets are different, it can be determined that this specific feature (315) is satisfied.

[0069] The first ligand that satisfies this specific feature will be recognized as the chimeric part in the structural model. If no ligand that satisfies the condition is found, or all ligands are screened out, more relaxed conditions, that is, the number of atoms in the molecule is less than 300 and it can be in the form of a polypeptide, are used to re-identify and screen the ligands.

[0070] It should be noted that if this specific feature is not used for screening and the number of interacting chains of the molecule is directly judged, non-expected molecules may be mis-identified, especially polypeptide fragments, molecular glues, or drugs occupied by protein-protein interaction interfaces parsed due to precision limitations in the protein structure.

[0071] Based on the recognition of the chimeric part, the two main protein chains with which the drug interacts can be recognized in a similar manner as above. Subsequently, the protein chains that interact with these two chains are recursively collected, and finally, the protein chain sets on both sides (316) are formed as the target protein and ubiquitin ligase substructure parts.

[0072] In the embodiments, a linker length distribution analysis module for chimeras is also provided. The linker part in a proteolysis-targeting chimera usually has high flexibility and greatly affects the distance limitation between the target protein and the ubiquitin ligase protein during the ternary complex assembly process. Therefore, a method for analyzing the linker length distribution of chimeras is required to analyze the distance distribution of the linker endpoints when the chimeric drug is in different three-dimensional conformations, as an index of the linker length range.

[0073] First, perform a molecular conformation search on the linker part extracted from the chimera in the complex. Specifically, a method based on the distance geometry method is used to generate multiple three-dimensional conformations of the molecule. First, define the distance range between atoms through a distance constraint matrix, then use randomized sampling to generate an initial conformation that satisfies the constraints, and finally use an optimization algorithm to minimize the energy of the conformation. For each three-dimensional conformation, calculate the straight-line distance between the atoms at both ends of the linker and perform statistical analysis.

[0074] Generally, this distance distribution presents a Gaussian distribution 305, but the specific chemical composition of the linker may affect the distribution. Using a Gaussian function to fit this distribution, the mean and variance of the distance distribution can be calculated as indexes of the linker length and length variability.

[0075] The ternary complex is the key structure for the proteolysis-targeting chimera drug to exert its degradation effect. The chimera forms a ternary complex by simultaneously binding to the target protein and the ubiquitin ligase, allowing the ubiquitin on the ubiquitin ligase to be transferred to the cysteine or lysine of the target protein and be recognized and degraded by the proteasome. Therefore, ternary complex modeling is the key to the design of proteolysis-targeting chimera drugs.

[0076] Based on the above structural analysis system, the embodiments provide a method for structural modeling of a proteolysis-targeting chimera complex. As Figure 5 shown, it includes the following steps:

[0077] S1. Using the chemical structures of the target protein, the target protein ligand, the ubiquitin ligase protein, the ubiquitin ligase protein ligand, and the linker as inputs, and combining with the linker length distribution analysis module of the chimera, adjust the relative positions of the target protein and the ubiquitin ligase protein for protein-protein docking, simulate the interaction between proteins, and form a set of protein-protein docking conformations.

[0078] In the embodiments, using the chemical structures of the target protein, the target protein ligand, the ubiquitin ligase protein, the ubiquitin ligase protein ligand, and the linker as inputs, specify the chemical structure of a target chimera during the modeling process to generate the corresponding ternary complex. The strength of the protein-protein interaction significantly affects the structural stability of the ternary complex, and further affects the degradation activity of the proteolysis-targeting chimera.

[0079] Therefore, the first step in ternary complex modeling is to perform protein-protein docking on the target protein and the ubiquitin ligase protein. Both the target protein and the ubiquitin ligase protein used for docking should contain their respective ligand parts. This docking process can be described as placing the three-dimensional structures of the two proteins in a reasonable relative position to achieve a high level of energetic evaluation of the protein-protein interactions. During this process, the relative three-dimensional positions of the backbone atoms of the same protein (including the nitrogen atom of each amino acid amino group, the carbon atom connected to the nitrogen atom, and the carbonyl atom in the peptide bond formed by each amino acid and the next amino acid) generally remain unchanged, while the relative three-dimensional positions of the side chain atoms (i.e., atoms other than the backbone atoms) can either remain unchanged or be reallocated.

[0080] In protein-protein docking, the protein with a larger number of atoms in the target protein and the ubiquitin ligase protein will be regarded as the fixed protein, and the positions of its backbone atoms will remain unchanged during the docking process. The protein with a smaller number of atoms will be regarded as the mobile protein, and it will undergo random translation or rotation during the docking process. The conformational search process consists of multiple Monte Carlo steps, that is, randomly selecting a translation of the mobile protein by 0 Å - 0.1 Å or a rotation of the mobile protein by 0° - 5°, and the rotation direction can be one of the X, Y, and Z axes. After each transformation, calculate the energy value of the current docking conformation based on the physical force field. If the energy of the new conformation is lower, it will be accepted and added to the structure set for continued search.

[0081] After several steps, when the ratio of the energy value decrease reaches a certain threshold, stop the search process, select each conformation in the current structure set, and output them sorted by the total energy value of the system.

[0082] Optionally, distance constraints are imposed during the search process of protein-protein docking to reduce the search space. This is because the chimera limits the distance and the binding angle between the target protein and the ubiquitin ligase protein. This distance constraint will be imposed between the ligand anchor atoms of the two proteins. The anchor atoms can be described as: the junction atoms between the ligand and the linker molecule of the chimera to be modeled. When imposing the constraint, first, based on the linker length distribution analysis module of the chimera, calculate the mean and standard deviation of the length distribution. Subsequently, define the distance constraint as an additive value to the energy value of each Monte Carlo step, which is in the form of a Gaussian function with distance as the independent variable, and the mean and standard deviation are equal to the mean and standard deviation of the length of the linker to be modeled. This energy value additive can make the distance between the anchor atoms gradually deviate towards the mean length of the linker to be modeled during the Monte Carlo steps, forming a set of protein-protein docking conformations.

[0083] S2. Based on the chimera component analysis module, generate a series of three-dimensional structural conformations of the chimera to be modeled, search for and analyze the conformations of the chimera based on the conformational search strategy, and form a set of chimera conformations.

[0084] In an embodiment, a series of three-dimensional structural conformations are generated for the protein degradation targeting chimera to be modeled, and unreasonable atom-atom distances and steric structures inconsistent with chemical substructures are avoided. This process is usually achieved by a series of Monte Carlo steps.

[0085] In the three-dimensional structure of the chimera, the conformation of the ligand part is usually very similar to the binding mode of the corresponding independent ligand and can be determined in advance and avoided from changing during the conformation search process, while the conformation of the linker part can be freely searched. Therefore, when performing the conformation search of the chimera, any of the following strategies can be optionally used.

[0086] (1) Perform a conformation search on the complete protein chimera drug molecule without imposing other restrictions.

[0087] (2) Split the linker part for conformation search, and then splice the ligand part onto each conformation of the linker. Specifically, first use the chimera component analysis module to extract the linker part in the chimera and split it into independent molecules for conformation search. Subsequently, attach the three-dimensional conformations of the ligand structures on both sides at the bond positions broken during the splitting.

[0088] (3) Only search the linker part. Specifically, first use the chimera component analysis module to extract each atom in the linker part of the chimera and screen all atoms that do not belong to any ring system. Subsequently, calculate the dihedral angles involved by these atoms and use multiple Monte Carlo steps based on the dihedral angles to achieve the conformation search.

[0089] (4) Impose conformation restrictions on the ligand part. Specifically, first use the chimera component analysis module to extract each atom of the ligand parts on both sides in the molecule. In order to keep the conformation of the ligand part unchanged during the search process, calculate the dihedral angles involved by these atoms as the limit values, and calculate the difference between each restricted dihedral angle and the limit value after the Monte Carlo steps of the search as an added value to the energy value of this Monte Carlo step. This strategy will make the conformation of the ligand part in the conformation result tend to the original conformation, and at the same time, small changes can be achieved to adapt to the possible changes in the binding mode after the ligand is assembled into the chimera drug.

[0090] The set of three-dimensional structures of the protein degradation targeting chimera obtained by the above method forms a chimera conformation set.

[0091] S3. Based on the protein-protein docking conformation set and the chimera conformation set, perform matching and splicing in the way of conformation merging with energy minimization, random sampling and screening, and optimize the result of the matching and splicing by means of fine docking of the conformation structure and energy optimization to form a chimera-induced ternary complex conformation set.

[0092] Conformational merging refers to the process of matching and splicing two sets of conformations, namely the protein-protein docking conformation set and the chimera conformation set, after obtaining them, in order to form the conformation of the chimera-induced ternary complex. In the protein-protein docking conformation set, each protein-protein docking conformation contains its bound ligand, which is used to search for the molecular conformation of the chimera so that the atomic positions of the ligand part therein are close to the atomic positions of the bound ligand in the protein-protein docking conformation. There are mainly two implementation methods for conformational merging based on the protein-protein docking conformation set and the chimera conformation set: based on energy minimization traction and based on random sampling and screening.

[0093] Both methods can adapt to most common protein degradation-targeting chimera modes, but there are significant performance differences for different chimera molecule sizes. For molecules with a smaller atomic weight and fewer rotatable bonds, using random sampling and screening is more efficient. Conversely, energy minimization traction can avoid the problem of simultaneously sampling a large number of rotatable bonds and has significantly higher performance. In the examples, the methods and parameters of these two implementations are elaborated in detail.

[0094] (1) Conformational merging based on energy minimization traction: This method does not require prior conformational search of the chimera molecule. Protein conformations are sequentially taken out from the protein-protein docking conformation set, and the positions of the bound ligands therein are analyzed. Subsequently, 64 arbitrary conformations (including planar conformations or arbitrary three-dimensional conformations) of the chimera molecule are generated, and their centroids are placed at the centroid center point position of the bound ligand. Subsequently, a two-step energy minimization process is carried out.

[0095] In the first step, the protein part in the protein-protein docking conformation is ignored, and only the ligand part is retained. The maximum common substructure between each ligand and the chimera molecule is calculated, and the atomic mapping relationship between the two molecules is formed. Subsequently, 100 Monte Carlo processes of energy minimization are carried out, and distance restrictions between atom pairs in the mapping are imposed during the process. This distance restriction is usually a Gaussian function with distance as the independent variable and is used as an additive value to the energy value to affect the Monte Carlo minimization process. The purpose of this restriction is to pull the ligand part in the chimera towards the ligand position in the docking conformation during the above process. After minimization, the root mean square distance deviation between the two needs to be calculated and screened out with a threshold of 3 angstroms.

[0096] In the second step, the protein part in the protein-protein docking conformation is reintroduced, and the ligand part in the protein-protein docking conformation is ignored. First, the positional conflict between the protein structure and the chimera molecular structure is calculated, which is calculated by the overlapping volume of their van der Waals volumes, and screened with a threshold of 150 cubic angstroms. Subsequently, several atomic sets are defined for defining different distance constraints: the pocket atomic set is defined as all residue atoms within 5 Å of the chimera molecule, and the interface atomic set is defined as that between any two protein chains A and B, where there are some protein side-chain atoms that are located on chain A and there are atoms of chain B within 10 Å. A Monte Carlo process of 100 energy minimizations is performed on this system, and optionally, distance constraints on the chimera molecule, the pocket atomic set, and the interface atomic set are applied so that they do not deviate significantly from the initial positions. Subsequently, the overlapping volume of the van der Waals volumes is rechecked. If it is higher than the threshold of 5 cubic angstroms, the above energy minimization process is performed again on this structure. Otherwise, the merged structure is output as the ternary complex conformation. If the overlapping volume threshold still cannot be reached after 10 energy minimization processes, this conformation is discarded.

[0097] This method is particularly applicable to cases where the complexity of the binding interface cavity is low and the number of rotatable bonds of the chimera molecule is large. In such cases, there may be a 10- to 100-fold efficiency improvement compared to methods based on random sampling and screening.

[0098] (2) Conformation merging based on random sampling and screening: This method requires prior conformational search of the chimera molecule. Conformations are sequentially taken out from the protein-protein docking conformation set. Subsequently, the ligand part and the conformations of each chimera molecule are superimposed. Superimposition is a method of minimizing the distance between specific atom pairs in two structures, and in this process, each structure is a rigid body, that is, the conformation remains unchanged. Subsequently, the root mean square distance deviation of the corresponding atoms between the ligand in the protein-protein docking conformation and the chimera molecule is calculated. If this deviation is higher than 1 Å, this chimera molecule conformation is discarded.

[0099] Subsequently, the positional conflict between the protein structure and the chimera molecular structure is calculated, which is calculated by the overlapping volume of their van der Waals volumes, and screened with a threshold of 300 cubic angstroms.

[0100] Define the pocket side-chain and backbone atomic sets for defining different distance constraints: the pocket atomic set is defined as all residue atoms within 5 Å of the chimera molecule, and the side-chain atoms and backbone atoms are respectively assigned to different atomic sets.

[0101] The system is subjected to at most 5 energy minimizations, each containing 20 Monte Carlo processes. Optionally, distance constraints are imposed on the chimeric molecule, the set of pocket side-chain atoms, and the set of backbone atoms during the process to prevent them from deviating significantly from their initial positions. The strength of the distance constraints can be set differently for different atom sets. Usually, the backbone atoms are frozen, a constraint of 1.0 relative unit is imposed on the chimeric molecule, and a constraint of 0.5 relative unit is imposed on the side-chain atoms. After each energy minimization, the overlapping volume between the protein structure and the chimeric molecule structure is rechecked. If the reduction in the overlapping volume compared to the previous energy minimization is less than 20%, the minimization process is considered to have converged, and no further energy minimization is performed.

[0102] After the minimization process converges or exceeds the maximum number of energy minimization thresholds, the overlapping volume between the protein structure and the chimeric molecule structure is recalculated. If it is higher than 10 cubic angstroms, the conformation is discarded; otherwise, the merged structure is output as the conformation of the chimeric-induced ternary complex.

[0103] This method is particularly applicable to cases where the binding interface is narrow and the chimeric molecule has few rotatable bonds. In such cases, methods based on energy minimization traction may not be able to search for conformations in a short time, while this method is more efficient in finding reasonable conformations.

[0104] For the conformation of the chimeric-induced ternary complex after merging, further structural optimization and energy minimization can be performed.

[0105] Define the set of pocket side-chain and backbone atoms for defining different distance constraints: The pocket atom set is defined as all residue atoms within 5 angstroms of the chimeric molecule, where the side-chain atoms and backbone atoms are assigned to different atom sets respectively.

[0106] Define the set of interface side-chain and backbone atoms for defining different distance constraints: The side-chain atom set is defined as between any two protein chains A and B, where there are some residue atoms that are located on chain A and there are atoms on chain B within 10 Å. The side-chain atoms and backbone atoms are assigned to different atom sets respectively.

[0107] For each conformation of the merged ternary complex, a Monte Carlo process of up to 100 energy minimizations is performed. Optionally, distance constraints are imposed on the ligand part and linker part of the chimeric molecule, the pocket side-chain and backbone atom sets, and the interface side-chain and backbone atom sets during the process to prevent them from deviating significantly from their initial positions. The strength of the distance constraints can be set differently for different atom sets. Usually, the backbone atoms of the pocket and interface are frozen, a constraint of 1.0 relative unit is imposed on the ligand part of the chimeric molecule, and a constraint of 0.5 relative unit is imposed on the side-chain atoms of the pocket and interface and the linker part of the chimeric molecule.

[0108] After this energy minimization step, an optimized structure and the total energy of the system are output, thereby forming a conformational set of the chimera-induced ternary complex.

[0109] S4. Cluster the conformational set of the chimera-induced ternary complex, select specific members in each category to form a structure set, combine the chimera-induced complex component analysis module, construct a conformational scoring and ranking method, rank each structure in the result set, and complete the structure modeling for the protein degradation targeting chimera complex.

[0110] In the embodiment, the common contact point ratio clustering algorithm (FCC) is used to cluster the conformational set of the chimera-induced ternary complex to select representative conformations therein. Specifically, during clustering, 75% common contact similarity is used as the class threshold, and the minimum number of conformations in each conformation class is limited to 4. After clustering, the number of members in each class is calculated, and the central conformation of each class is selected. Among them, the central conformation of the class is defined as the conformation with the highest sum of common contact similarities with other conformations in the class.

[0111] During the conformational clustering process, the conformation of the chimera molecule is not considered because the linker part has high flexibility and the ligand part is anchored to the protein position, so no other clustering metrics are set outside the protein position.

[0112] The structure modeling method provided in the embodiment will predict 10 - 1000 three-dimensional structures. Among them, it is expected that there are reasonable results similar to the structure obtained from the crystal diffraction experiment of the chimera-induced ternary complex (hereinafter referred to as the crystal structure), and there are also unreasonable results with large deviations. Therefore, constructing a conformational scoring and ranking method to rank each structure in the result set can further optimize the modeling prediction results, and the final result conformation can be used in other new drug R & D tasks such as molecular design and optimization to complete the structure modeling for the protein degradation targeting chimera complex.

[0113] In the embodiment, conformational scoring needs to be calculated using multiple metrics or scores of the conformation, including one or more of the following: the number of atoms of the chimera drug, the number of atoms of the linker, the number of movable dihedral angles of the linker, the mean and variance calculated after Gaussian function fitting of the length distribution of the linker, the docking score of the chimera drug, the total energy of the system after energy minimization, the pairwise binding energy components of the chimera drug and the two proteins, the number of member conformations in the conformation class where it is located, the ranking from the central conformation, the energy of the protein-protein interaction interface, the difference in solvent accessible surface area of the linker in the dissociated state and the bound state, and the strain energy of the chimera drug molecule conformation. The sources or calculation methods of these metrics or scores are elaborated in detail below:

[0114] (1) The number of atoms of the chimera can be directly obtained from the specified chemical structure, and hydrogen atoms are not included in the calculation.

[0115] (2) The number of atoms in the linker can be obtained by the chimera component analysis module extracting the linker part and visually calculating the number of atoms through the chemical structure, excluding hydrogen atoms during the calculation.

[0116] (3) The number of mobile dihedral angles in the linker is achieved through publicly available chemical structure analysis methods. The principle is to determine the mobile dihedral angles by calculating whether the bonds in the aromatic ring system and conjugated system can be twisted.

[0117] (4) The length distribution of the linker of the chimera and its fitting can be obtained according to the aforementioned linker length distribution analysis module of the chimera.

[0118] (5) The docking score of the chimera is calculated by using the protein-drug molecule docking method, which calculates the docking score obtained by in-situ docking of the chimera molecule to the target protein-ubiquitin ligase complex structure. In-situ docking means not performing the molecular conformation search in the docking process, but directly using the position of the chimera molecule in the ternary complex conformation for molecular docking scoring.

[0119] (6) The total energy of the system after energy minimization can be calculated and output during the conformational structure optimization and energy minimization process.

[0120] (7) The calculation of the binding energy component can be described as follows: for two structures A and B, each structure may include one or more protein chains or molecules. When there is only one structure in the calculation system, the total energy of the system and , and then the total energy of the system when the two are combined in the system is calculated , then the binding energy component between the two can be defined as . Four groups of binding energy components can be calculated between the chimera and the ubiquitin ligase, the chimera and the target protein, the ubiquitin ligase and the target protein, and the complex of the chimera with the ubiquitin ligase and the target protein.

[0121] (8) The number of member conformations of the conformational class where it is located can be calculated and output during the conformational clustering process.

[0122] (9) The calculation of the sorting of the distance from the central conformation during the clustering process can refer to the calculation method of the central conformation during the conformational clustering process.

[0123] (10) The calculation of the protein-protein interaction interface energy is usually performed after conformational structure optimization and energy minimization. Compared with the protein-protein binding energy component, this indicator has stronger adaptability for the predicted structure of protein-protein docking.

[0124] (11) The binding of the target protein and the ubiquitin ligase usually forms a cavity, and the linker of the chimera is usually located within this cavity. The difference in solvent accessible surface area between the dissociated state and the bound state can be used to determine the degree to which the linker part is close to the surface of the cavity pocket. The solvent accessible surface area refers to the surface area of the protein molecule that can be contacted by the solvent and can be calculated by the rolling ball algorithm, which uses a sphere with a specific radius to simulate the solvent to detect the surface of the biomolecule.

[0125] (12) The calculation of the strain energy of the chimera molecular conformation can be described as follows: First, perform a conformational search of the molecule, simulate water as the solvent, calculate the total energy of the system in the lowest energy conformation, and find the difference from the current molecular conformation energy. This energy difference can be used as the strain energy and characterize the likelihood of the natural occurrence of this molecular conformation in the polar solvent state.

[0126] In the examples, the conformation score is used to characterize the similarity of the conformation to the crystal structure. Specifically, it is characterized by the difference in the binding mode between the target protein and the ubiquitin ligase protein. To further quantify this similarity, three key interface quality indicators need to be calculated: the root mean square deviation of the main chain atoms at the interface (iRMS), the root mean square deviation of the main chain atoms of the ligand protein (LRMS), and the recurrence rate of protein-protein interactions (Fnat). These indicators evaluate the matching degree of the conformation to the crystal structure from different perspectives.

[0127] Among them, iRMS measures the deviation of the main chain atoms (such as atoms) at the binding interface between the target protein and the ubiquitin ligase protein. A lower iRMS value indicates a higher similarity in the main chain conformation between the conformation and the crystal structure, and a better matching degree in the interface region. LRMS measures the deviation of the main chain atoms of the ligand protein (i.e., the ubiquitin ligase protein) in the whole complex. It not only focuses on the interface region but evaluates the deviation of the overall main chain conformation of the ligand protein from the crystal structure. A lower LRMS value indicates that the overall conformation of the ligand protein is more consistent with the crystal structure. Fnat measures the recurrence rate of the real contacts between the target protein and the ubiquitin ligase protein in the conformation. It is evaluated by calculating the matching ratio of the predicted interface atom pairs to the real interface atom pairs in the crystal structure. A higher Fnat value indicates that the conformation better reproduces the real interaction mode between proteins.

[0128] By calculating these indicators, the precise matching degree of the conformation relative to the crystal structure can be obtained. However, a single iRMS, LRMS, or Fnat may not comprehensively reflect the quality of the conformation because they focus on different aspects respectively. Therefore, the DockQ algorithm is selected to calculate the comprehensive score to more comprehensively evaluate the quality of the conformation, and its calculation formula is as follows:

[0129] ,

[0130] ,

[0131] ,

[0132] wherein, and are predefined threshold parameters, represents the root mean square deviation standardized by the predefined parameters, therefore, it can be calculated by iRMS, LRMS and Fnat and used to evaluate the similarity between the modeled conformation and the crystal structure, and the result is between 0 and 1. A larger value indicates a higher similarity between the modeled conformation and the crystal structure.

[0133] For each conformation output by one modeling, the scoring training target value of the -th conformation can be expressed by the following formula, that is, the ratio of the cube of the conformation score to the cube of the optimal conformation :

[0134] ,

[0135] The similarity to the crystal structure can also include the conformational and positional differences of the chimeric molecule. However, since the linker usually has high flexibility and the ligand part of the chimeric molecule has been anchored in the original ligand binding site of the protein, it is usually not considered additionally.

[0136] In order to train the weights of each index in the conformation score through the model, it is necessary to collect conformation score training data in the examples. First, the ternary complex structure of the known crystal structure in Table 2 below is disassembled, chimeric-induced ternary complex modeling is performed, and each index or score of each conformation is calculated according to the foregoing method, and at the same time, the similarity between the conformation and the crystal structure is calculated. The known crystal structure 701 can be obtained from the PDB (Protein Data Bank) and has a corresponding database number.

[0137] Table 2

[0138]

[0139] For each known crystal structure, it is necessary to automatically split the complex structure to generate input data for validation modeling. Based on the chimera-induced complex component analysis module, the target protein, ubiquitin ligase, and proteolysis-targeting chimera in the ternary complex are identified. Further using the above chimera component analysis module, the linker part in the proteolysis-targeting chimera is identified and removed, and only the ligand part is retained. The two proteins and their ligand parts are randomly moved and rotated so that the distance is at least 30 Å to ensure that the original binding mode between proteins cannot be transmitted to the modeling algorithm through the input data.

[0140] Subsequently, validation modeling is performed using the generated input data. For each conformation, various metrics and scores are calculated and stored, and at the same time, its similarity to the crystal structure is calculated and stored as a training dataset for constructing conformation scores.

[0141] Since there are a large number of conformations dissimilar to the crystal structure and fewer similar conformations in this dataset, there is an issue of dataset imbalance, which may affect the training accuracy. Therefore, molecular dynamics methods are used for data augmentation. For each ternary complex modeling result set, the 3 conformations with the highest similarity to the crystal structure are selected for molecular dynamics simulation, and the structures after kinetic perturbation are extracted.

[0142] Specifically, the structure is placed in an aqueous solution environment, 0.15 mol / L sodium chloride ions are added to simulate the physiological environment, and the TIP3P water model is used to construct a periodic water box. The system first performs 5000 steps of energy minimization, and then performs a 10-nanosecond molecular dynamics simulation under the NPT ensemble. During this process, the Langevin thermostat is used to maintain the temperature at 300 Kelvin and the pressure at 1 atmosphere. The time step for kinetic calculation is 2 femtoseconds, and all hydrogen bonds are constrained using the SHAKE algorithm. In the last 5-nanosecond kinetic trajectory of the simulation, one conformation is taken every 10 picoseconds, and RMSD (Root Mean Square Deviation) clustering analysis is performed on these conformations. The number of clusters is set to 10, and the central conformation of each cluster is selected as the representative conformation. Energy minimization is performed on these representative conformations, and the resulting conformations are used as supplementary samples for data augmentation, thus expanding the number of conformations similar to the crystal structure in the dataset. This collection process can form a dataset containing 8524 modeling conformations and their metrics.

[0143] Next, the random forest algorithm is used for the training of conformational scoring. Before training, the conformational index dataset is standardized. The Z-score standardization method is adopted to eliminate the dimensional differences between different features, ensuring that each feature has the same weight in model training. During model training, the number of decision trees is set to 500, the maximum number of features is set to the square root of the total number of features, the maximum depth is set to 10, and the minimum number of samples in leaf nodes is set to 2 to avoid the model being sensitive to noise caused by too small leaf nodes. The stratified cross-validation technique is adopted to ensure that the model can achieve good performance on different data subsets. The early stopping strategy is used to avoid overfitting.

[0144] Finally, through the above training process of step-by-step iteration, optimization and evaluation, a robust random forest model with good generalization ability is constructed, and the root mean square error (MSE) of conformational scoring prediction is 0.11. This model can accurately predict the multi-factor scores of new conformations, thus providing reliable data support for the ranking of the conformational set of ternary complexes.

[0145] To determine the relationship between conformational scoring and ranking accuracy, the PDB database entry numbered 6HR2 is used as a reference for ternary complex modeling. This entry is a ternary complex structure and is not involved in the dataset. Conformational ranking is performed through its predicted scores, which has stronger ranking performance compared to ranking only by protein-protein interaction energy or score, or only by the size of ternary complex conformational clustering, as shown in Table 3.

[0146] Table 3

[0147]

[0148] The data in Table 3 show that the top several models ranked using conformational scoring have higher similarity to the 6HR2 reference structure in terms of structure, especially in the key interaction regions. The ranking results of conformational scoring can better reflect the true structural features. This result indicates that conformational scoring can not only effectively distinguish the energy states of different conformations, but also comprehensively consider factors such as structural rationality and interaction complexity, thus providing more accurate ranking results.

[0149] Further analysis shows that the method of ranking using conformational scoring can effectively distinguish the quality of different conformations. Especially in the complex interactions of ternary complexes, it can capture more subtle structural differences. In contrast, the method of ranking relying solely on protein-protein interaction energy or scoring, although it can provide reasonable ranking results in some cases, often fails to comprehensively consider the effects of all interactions when dealing with complex ternary complexes, resulting in a decrease in the accuracy of the ranking results. In addition, the method of ranking only by cluster size, although it can reflect the prevalence of certain conformations, cannot accurately evaluate the structural rationality and energy characteristics of the conformations.

[0150] In summary, conformational scoring shows significant advantages in the ranking of ternary complexes, and can provide a more reliable tool for structural biology research, helping researchers to more accurately identify the optimal conformational model in complex protein-protein interactions for the structural modeling of proteolysis-targeting chimera complexes.

[0151] To verify the effectiveness of the structural analysis system and modeling method for proteolysis-targeting chimera complexes provided by the present invention, an example was carried out for the research and development of proteolysis-targeting chimera drugs.

[0152] The structural modeling method for proteolysis-targeting chimera complexes involved in the present invention can be applied to the research and development of proteolysis-targeting chimera drugs for protein kinase B (AKT). In the initial stage, a lead compound with weak degradation activity was obtained. To improve the degradation efficacy of this compound, the structural modeling method described in the present invention was used to conduct an in-depth analysis of the interaction between AKT and ubiquitin ligase.

[0153] By modeling the ternary complex of the lead compound with AKT and E3 ligase, it was found that there was a phenylalanine group adjacent to the protein surface on the linker spatial path, and potential π-π interactions could be formed through chemical structure design. This interaction is considered to be a key factor in improving the drug degradation activity. Therefore, the structure of the linker part of the drug was optimized, and a benzene ring structure was introduced at an appropriate position, thereby enhancing the interaction between the linker and the protein-protein interaction cavity and improving the stability of the ternary complex.

[0154] The optimized chimera showed a significantly improved degradation activity in cell experiments. Compared with the unoptimized lead compound, the optimized molecule could achieve effective degradation of AKT at a lower concentration. These results fully demonstrate the effectiveness of the ternary complex modeling method described in the present invention in guiding the structural optimization of proteolysis-targeting chimera molecules and improving the degradation efficacy.

[0155] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A structural analysis system for a protein degradation targeted chimera complex, characterized in that: The structural analysis system includes: a chimera component analysis module, a chimera induced complex component analysis module and a chimera connector length distribution analysis module; The chimera component analysis module is used to identify a series of structural fragments in the chimera, and fuse adjacent structural fragments using an iterative method; sort each fused fragment, test the rationality of the connection between the target protein ligand and the ubiquitin ligase ligand and the sorted fused fragments, and regard the rationally sorted fused fragments as the connector part to achieve the splitting and component analysis of the chimera; The chimera-induced complex component analysis module is used to split the chimera-induced ternary complex, and analyze, identify and screen the split target protein, chimera and ubiquitin ligase to achieve component analysis of the chimera-induced complex; The connector length distribution analysis module of the chimera is used to perform molecular conformation search on the connector part extracted from the chimera in the complex, analyze the connector endpoint distance distribution when the chimera is in different three-dimensional conformations, and use it as an indicator of the connector length range for structural analysis of protein degradation targeted chimera complexes.

2. The structural analysis system according to claim 1, characterized in that: The identification of a series of structural fragments in the chimera comprises: identifying structural fragments based on a molecular structure pattern or based on a loop matching pattern; The structural fragments matched by the molecular structure pattern include: primary amide, primary thioamide, a carbon atom connected to two hydrogen atoms, a hydroxymethyl group, an alkynyl group, a substructure formed by a ring-containing nitrogen atom connected to a non-cyclic carbonyl group, or an azo group formed by two nitrogen atoms connected by a double bond; The structural fragments of the ring matching pattern include: a six-membered carbon aromatic ring containing two substituents, an aliphatic ring system, or a three- to six-membered aromatic nitrogen heterocycle or oxygen heterocycle containing two substituents, wherein the aliphatic ring system includes a monocyclic heterocycle and a spirocyclic system containing nitrogen and oxygen, and these ring systems must have exactly two substituents, the monocyclic heterocycle is limited to a three- to six-membered aliphatic ring, and the spirocyclic ring can be one of six-membered-six-membered, six-membered-five-membered, and five-membered-five-membered, and at most one substituent in the three- to six-membered aromatic nitrogen heterocycle or oxygen heterocycle containing two substituents can be connected to the peptide bond.

3. The structural analysis system according to claim 1, characterized in that: The adjacent structural fragments refer to fragments that can be connected by a single chemical bond or have partially overlapping chemical structures.

4. The structural analysis system according to claim 1, characterized in that: The test of the rationality of the connection between the target protein ligand and the ubiquitin ligase ligand and each fusion fragment includes: if the corresponding target protein ligand and the ubiquitin ligase ligand meet the conditions, the fusion fragment is discarded, and the next fusion fragment in the sorting is selected for testing; if the fusion fragment meets the rationality requirements of the corresponding ligand, it is regarded as the connector part of the chimera, wherein the conditions are corresponding to only zero or one ligand, the number of atoms of the ligand is less than 5, or the ligand does not contain any ring structure.

5. The structural analysis system according to claim 1, characterized in that: The analysis, identification and screening of the target protein, chimera and ubiquitin ligase after splitting include: for the chimera part, identifying each non-protein molecule through non-standard amino acid residues, limiting the number of atoms in the non-protein molecule to less than 150 and not in the form of polypeptides, and sorting them from large to small by the number of atoms, screening these non-protein molecules based on specific features, and selecting the molecule with the largest number of atoms after screening, wherein the specific feature is the ligands on both sides of the chimera molecule, which interact with different protein chains in the target protein and ubiquitin ligase respectively; if no satisfying ligand is found or all ligands are screened out, re-identifying and screening the ligands under more relaxed conditions to achieve identification and screening of the chimera part, wherein the more relaxed condition is that the number of atoms in the non-protein molecule is less than 300 and can be in the form of polypeptides; On the basis of identifying the chimera part, the two main protein chains that interact with the target protein and ubiquitin ligase are identified, the protein chains that interact with these two main protein chains are recursively collected, and finally a set of protein chains on both sides is formed to realize the identification and screening of the target protein and ubiquitin ligase, and complete the analysis, identification and screening of the split target protein, chimera and ubiquitin ligase.

6. A method for structural modeling of a protein degradation targeted chimera complex, characterized in that: The structural modeling method uses the structural analysis system according to any one of claims 1 to 5, and comprises the following steps: Based on the chemical structures of target protein, target protein ligand, ubiquitin ligase protein, ubiquitin ligase protein ligand and linker as input, combined with the linker length distribution analysis module of the chimera, the relative positions of target protein and ubiquitin ligase protein are adjusted to perform protein-protein docking, simulate the interaction between proteins, and form a protein-protein docking conformation set; Based on the chimera component analysis module, a series of three-dimensional structural conformations are generated for the chimera to be modeled, and the conformations of the chimera are searched and analyzed based on the conformation search strategy to form a chimera conformation set; Based on the protein-protein docking conformation set and the chimera conformation set, matching and splicing are performed by energy minimization, random sampling and screening conformation fusion, and the matching and splicing results are optimized by conformational structure fine docking and energy optimization to form a chimera-induced ternary complex conformation set; The chimera-induced ternary complex conformation sets are clustered, and specific members in each category are selected to form a structure set. Combined with the chimera-induced complex component analysis module, a conformation scoring and sorting method is constructed, and each structure in the result set is sorted to complete the structural modeling of the chimera complex targeted for protein degradation. After clustering, the number of members in each category is calculated, and the central conformation of each category is selected. The central conformation of a class is defined as the conformation with the highest sum of common contact similarities with other conformations in the class, that is, the specific member is obtained.

7. The structural modeling method according to claim 6, characterized in that: The adjustment of the relative positions of the target protein and the ubiquitin ligase protein for protein-protein docking includes: treating the protein with a large number of atoms as a fixed protein, keeping the position of the skeleton atoms unchanged during the docking process, treating the protein with a small number of atoms as a mobile protein, performing random translation or rotation during the docking process, and performing position transformation through multiple Monte-Carlo steps to achieve the lowest total energy of the system and complete the protein-protein docking.

8. The structural modeling method according to claim 6, characterized in that: The energy minimization includes: extracting conformations from the protein-protein docking conformation set and analyzing the ligand position, generating an arbitrary conformation of the chimera and placing the center of mass of the chimera at the center point of the ligand center of mass, and performing a two-step energy minimization process, wherein the two-step energy minimization process includes: step 1 ignoring the protein part in the protein-protein docking conformation for energy minimization, and at the same time pulling the ligand in the chimera to move to the ligand position in the protein-protein docking conformation; step 2 re-adding the protein part in the protein-protein docking conformation for energy minimization.

9. The structural modeling method according to claim 6, characterized in that: The random sampling and screening include: searching the conformation of the chimera in advance, extracting the conformation from the protein-protein docking conformation set, superimposing the conformation of the ligand and each chimera, calculating the root mean square distance deviation of the corresponding atoms between the ligand and the chimera, setting a threshold for preliminary conformation screening; calculating the position conflict between the protein and the chimera, setting a threshold for secondary screening; performing multiple Monte-Carlo processes to achieve energy minimization convergence, and completing the screening.

10. The structural modeling method according to claim 6, characterized in that: The conformational score is calculated using multiple indicators or scores of the conformation, including one or more of the following: the number of atoms in the chimera, the number of atoms in the linker, the number of movable dihedral angles of the linker, the mean and variance of the length distribution of the linker calculated after Gaussian function fitting, the docking score of the chimera, the total energy of the system after energy minimization, the pairwise binding energy components of the chimera and the two proteins, the number of member conformations in the conformational category, the ranking of the distance to the central conformation, the protein-protein interaction interface energy, the difference in solvent accessible surface area of ​​the linker in the dissociated and bound states, and the strain energy of the chimera conformation.

Citation Information

Patent Citations

  • PROTAC molecular drug design method based on DEL platform

    CN118711703A

  • Biomacromolecule targeted protein hydrolysis chimera BioPROTAC capable of degrading PD-L1 as well as preparation method and application of biomacromolecule targeted protein hydrolysis chimera BioPROTAC

    CN114891119A

  • PROTAC micromolecule preparation method and application

    CN116747316A