Transient state searching method and system for complex biomacromolecule conformational change and application

Through reversible dimensionality reduction and low-dimensional search methods, searching for transition states of complex biological macromolecules in low-dimensional space, the efficiency and accuracy of transition state search in high-dimensional space is solved, and efficient and accurate identification of transition state structure is achieved.

CN120356509AActive Publication Date: 2025-07-22THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510306467.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-22
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently search for transition states of complex biological macromolecules in high-dimensional spaces, especially due to the problem of loss of information during dimensionality disasters and dimensionality reduction.

Method used

Reversible dimensionality reduction method is used to project high-dimensional molecular dynamics data into low-dimensional space, combined with low-dimensional search methods such as GAD to search candidate transition state samples in low-dimensional space, and verify the transition state structure by reversible mapping back to high-dimensional space.

Benefits of technology

High-dimensional transition state structure search with high computing efficiency and high accuracy is achieved, which significantly reduces calculation costs, avoids false positive problems, and ensures the accuracy of transition state recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356509A_ABST
    Figure CN120356509A_ABST
Patent Text Reader

Abstract

The invention provides a transition state searching method and system for complex biomacromolecule conformational change and application. The method comprises the steps that an input data set related to molecular dynamics simulation trajectory data is acquired; dimensionality reduction is carried out on an input data set through a reversible dimensionality reduction method, trajectory data is projected from a high-dimensional space to a low-dimensional space on the basis of reserving transition state information, and a low-dimensional free energy surface is generated; searching a low-dimensional candidate transition state sample on the low-dimensional free energy surface by adopting a low-dimensional search method; inversely projecting the low-dimensional candidate transition state sample back to a high-dimensional space through a reversible dimension reduction method to obtain a complete candidate transition state structure; and verifying the candidate transition state structure, and screening out a real transition state structure according to a verification result. According to the method, high-dimensional transition state structure search with high calculation efficiency and high accuracy is realized through reversible dimension reduction and low-dimensional search based on generative artificial intelligence, and the method is high in expandability, high in physical reliability and suitable for a complex biomolecular system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computational simulation of biomolecular systems, and more particularly, to a method, system and application for searching for transition states of conformational changes of complex biomacromolecules. Background Art

[0002] When biomolecules perform their functions, they are often accompanied by huge structural changes, that is, functional conformational changes of biomolecules. The transition state is the key for physical chemists to understand and regulate the microscopic mechanism of the functions of biomacromolecules. Since its existence time is extremely short and it is difficult to be captured by experimental means, its structure must be comprehensively characterized by simulation calculations driven by physical laws. However, different from the chemical reaction process which only involves a small number of atoms, the functional conformational changes of biomacromolecules involve a huge number of atoms and coordinates, and searching for their transition states will inevitably encounter the curse of dimensionality, that is, the reaction coordinate problem. Since the transition state is usually in a low-probability distribution region in high-dimensional space, directly searching for the transition state has extremely high computational costs.

[0003] The conformational changes of complex macromolecules involve a huge number of atoms, and in extreme cases, they can include all the atoms of the solute, and even the atoms of lipid and solvent molecules in the environment. The numerous atoms (and their three-dimensional coordinates) pose great challenges for the automated analysis of molecular dynamics and accurately finding the transition state.

[0004] To determine whether the state (conformation or structure) of a molecule is the transition state of its motion, it is mainly done through Committor Analysis (CA), that is, starting from this structure, running multiple (tens to thousands) independent molecular dynamics simulations (with different initial rates). If nearly 50% of the simulations fall into the stable state A first, and the other 50% fall into the stable state B, then this molecule can be determined to be the transition state (or at least not far from the transition state). However, due to the excessive possible structures in high-dimensional space, it is impossible to perform CA determination one by one. Therefore, CA cannot be used for transition state search, but can only be used as a post-detection method after searching for some candidate transition state samples.

[0005] Existing transition state search methods, such as Gentlest Ascent Dynamics (GAD), have major defects: 1. It can only be performed on a potential energy function described by an analytical formula; 2. It can only be performed in a space with a lower dimension. Due to the huge number of atoms and extremely high dimensions of complex biomacromolecules, it directly leads to the reaction coordinate problem, that is: 1. It is impossible to directly run GAD in high-dimensional space and dimensionality reduction is required first; 2. The free energy surface formed after dimensionality reduction of the potential energy surface cannot be directly expressed by an analytical formula; 3. If there is a slight mistake during the dimensionality reduction process, the transition state region will be distorted and relevant information will be lost.

[0006] GAD needs to frequently calculate the Hessian matrix, so its fast search is usually limited to lower-dimensional spaces. This characteristic makes it difficult for GAD to be directly applied to biomacromolecular systems and requires other methods to reduce the dimension first. However, existing dimensionality reduction algorithms have significant defects. For example, the commonly used time-lagged Independent Component Analysis (tICA) method uses the kinetic information contained in the data for dimensionality reduction, but the low-dimensional space coordinates are only linear combinations of the original high-dimensional space, while the actual requirements may be non-linear combinations. Other non-linear dimensionality reduction algorithms only perform dimensionality reduction based on steady-state information and density information and completely ignore kinetic information (the conditional probability that one conformation transfers to another within a certain relaxation time, Transition Probability) during the learning process. Therefore, it cannot guarantee the preservation of kinetic information during dimensionality reduction and cannot be combined with GAD to search for transition states.

[0007] Therefore, a solution is needed to achieve the search for transition states of conformational changes in complex biomacromolecules. Summary of the Invention

[0008] Based on the problems existing in the prior art, the present application provides a method, system and application for searching transition states of conformational changes in complex biomacromolecules. The specific solutions are as follows:

[0009] In the first part, the present application proposes a method for searching transition states of conformational changes in complex biomacromolecules, including the following:

[0010] Obtain an input data set involving molecular dynamics simulation trajectory data;

[0011] Reduce the dimension of the input data set through a preset reversible dimensionality reduction method, project the trajectory data from a high-dimensional space to a low-dimensional space while retaining the transition state information of the input data set, and generate a low-dimensional free energy surface;

[0012] Use a preset low-dimensional search method to search for low-dimensional candidate transition state samples on the low-dimensional free energy surface;

[0013] Inverse-project the low-dimensional candidate transition state samples back to the high-dimensional space through the reversible dimensionality reduction method to obtain candidate transition state structures in a complete high-dimensional representation;

[0014] Verify the candidate transition state structures and screen out the true transition state structures according to the verification results.

[0015] In some specific embodiments, the dimensionality reduction process specifically includes:

[0016] Set the relaxation time and the dimension reduction dimension based on a preset target molecular system, so as to map the trajectory data of the input data set to the latent space corresponding to the dimension reduction dimension through a preset first dimension reduction method;

[0017] Set the relaxation time and the number of dimensions of the low-dimensional space based on the current molecular system, so as to map the data in the latent space to the low-dimensional space through a preset second dimension reduction method.

[0018] In some specific embodiments, the search process for the low-dimensional candidate transition state samples includes:

[0019] Determine the starting position and the initial search direction of the search, and optimize each starting position according to the initial search direction;

[0020] Search at all starting positions according to the set search step size and the number of iteration steps to obtain the search trajectory of each step;

[0021] Extract the corresponding data points as low-dimensional candidate transition state samples by evaluating the convergence of each search trajectory.

[0022] In some specific embodiments, the first dimension reduction method includes the time-structured independent component analysis method;

[0023] And / or, the second dimension reduction method includes the reaction coordinate flow method, the diffusion model or the flow model based on the variational autoencoder.

[0024] In some specific embodiments, the low-dimensional search method includes GAD, the GAD variant optimized based on the Hessian matrix, and the combination of GAD and gradient descent.

[0025] In some specific embodiments, calculate the committor probability of each candidate transition state structure. If the committor probability of a certain candidate transition state structure meets the preset interval, then determine that this candidate transition state structure is the true transition state structure.

[0026] In some specific embodiments, the process of obtaining the true transition state structure includes:

[0027] Perform molecular dynamics simulations on each candidate transition state structure to obtain multiple molecular dynamics trajectories;

[0028] Determine the metastable state that each molecular dynamics trajectory approaches first, and calculate the committor probability according to the number of different metastable states that all molecular dynamics trajectories approach first;

[0029] If the committor probability of a certain candidate transition state structure is within the preset interval, then determine that this candidate transition state structure is the true transition state structure.

[0030] In the second part, the present application proposes a transition state search system for conformational changes of complex biomacromolecules, including the following:

[0031] An input unit for obtaining an input data set involving molecular dynamics simulation trajectory data;

[0032] A dimensionality reduction unit for reducing the dimensionality of the input data set by a preset reversible dimensionality reduction method, projecting the trajectory data from a high-dimensional space to a low-dimensional space while retaining the transition state information of the input data set, and generating a low-dimensional free energy surface;

[0033] A search unit for searching for low-dimensional candidate transition state samples on the low-dimensional free energy surface by using a preset low-dimensional search method;

[0034] A candidate unit for inversely projecting the low-dimensional candidate transition state samples back to the high-dimensional space by the reversible dimensionality reduction method to obtain candidate transition state structures in a complete high-dimensional representation;

[0035] A verification unit for verifying the candidate transition state structures and screening out the true transition state structures according to the verification results.

[0036] In the third part, the present application proposes a computer device, which includes:

[0037] One or more processors;

[0038] A memory for storing one or more programs;

[0039] When the one or more programs are executed by the one or more processors, the one or more processors implement the transition state search method for conformational changes of complex biomacromolecules as described in any one of the first part.

[0040] In the fourth part, the present application proposes a computer program product, including executable instructions for implementing the transition state search method for conformational changes of complex biomacromolecules as described in any one of the first part when being executed by a processor.

[0041] Advantageous effects: The present application proposes a transition state search method, system and application for conformational changes of complex biomacromolecules. Based on generative artificial intelligence, it realizes the search for high-dimensional transition state structures with high computational efficiency and high accuracy through reversible dimensionality reduction and low-dimensional search, and has strong scalability and high physical reliability, and is applicable to complex biomolecular systems. Projecting high-dimensional molecular conformation data into a low-dimensional space and performing parallel transition state search in the low-dimensional space makes the search complexity much lower than that in the high-dimensional space, significantly reducing the computational cost. The search process can not only detect the transition state, but also generate reasonable transition state conformations, accurately display the key structures corresponding to the conformational change process, avoid the false positive problem caused by abnormal data distribution, and effectively reduce the false positive rate of transition state recognition.

[0042] To make the above objects, features, and advantages of the present application more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and describes them in detail as follows. Description of the Drawings

[0043] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0044] Figure 1 It is a schematic flowchart of the transition state search method of the present application;

[0045] Figure 2 It is a schematic flowchart of the specific process of the RCF-GAD method of the present application;

[0046] Figure 3 It is a schematic diagram of the RCF principle of the present application;

[0047] Figure 4 It is a metastable state distribution diagram during the transition process of the T4 lysozyme L99A variant from the ground state to the excited state in the verification experiment of the present application;

[0048] Figure 5 It is a three-dimensional visualization schematic diagram of the transition state distribution in the 4 reaction coordinate space in the verification experiment of the present application;

[0049] Figure 6 It is a schematic diagram of the key conformational changes related to the transition state identified in the verification experiment of the present application;

[0050] Figure 7 It is a schematic diagram of the working role simulation system module of the present application.

[0051] Reference numerals: 1 - input unit; 2 - dimensionality reduction unit; 3 - search unit; 4 - candidate unit; 5 - verification unit. Detailed Embodiments

[0052] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0053] The present application proposes a method for searching the transition states of conformational changes of complex biomacromolecules. By combining a reversible dimensionality reduction method and a low-dimensional search method, first, the reversible dimensionality reduction method that can retain the kinetic information of the high-dimensional space in the low-dimensional space during the dimensionality reduction process is used to reduce the dimensionality of the molecular dynamics dataset of complex biomacromolecules. Then, possible candidate transition state sample points are searched on the low-dimensional free energy surface. Finally, the low-dimensional candidate samples are reversely projected back to the original high-dimensional space, and the candidate transition states are verified in the original space to ensure the accuracy of the obtained high-dimensional transition state structures. The schematic flow diagram of the method for searching the transition states of conformational changes of complex biomacromolecules is as shown in the appendix Figure 1 as shown, and the principle is as shown in the appendix Figure 2 as follows:

[0054] A method for searching the transition states of conformational changes of complex biomacromolecules includes the following:

[0055] 101. Obtain an input dataset involving molecular dynamics simulation trajectory data;

[0056] 102. Reduce the dimensionality of the input dataset by a preset reversible dimensionality reduction method, project the trajectory data from the high-dimensional space to the low-dimensional space while retaining the transition state information of the input dataset, and generate a low-dimensional free energy surface;

[0057] 103. Use a preset low-dimensional search method to search for low-dimensional candidate transition state samples on the low-dimensional free energy surface;

[0058] 104. Reverse-project the low-dimensional candidate transition state samples back to the high-dimensional space through the reversible dimensionality reduction method to obtain the candidate transition state structures in the complete high-dimensional representation;

[0059] 105. Verify the candidate transition state structures and screen out the true transition state structures according to the verification results.

[0060] The present application combines a reversible dimensionality reduction method with a low-dimensional search method such as GAD for transition state search. The dimensionality reduction process can retain both density information and transition pair information simultaneously, avoiding the loss of transition states caused by traditional dimensionality reduction. Performing GAD search in the low-dimensional space has highly parallelized search and significantly reduced time costs. The low-dimensional transition states are reversibly mapped back to the high-dimensional space to ensure the accurate reconstruction of the transition state structures. Taking the reversible dimensionality reduction method as RCF and the low-dimensional search method as GAD as an example, the principle is as shown in the appendix Figure 2 as shown. After verification, the combination of RCF + GAD reduces the transition state search time from several hours in high dimensions to a few minutes.

[0061] In this application, the reversible dimensionality reduction method uses the transition state information as the loss function during the training process to retain the transition state information in the original input dataset. Among them, the transition state information is hidden in the transition (transfer, conditional) probability of the molecule transitioning from one structure to another within a relaxation time in the high-dimensional trajectory and cannot be directly obtained. However, during the training process of the reversible neural network, the transition probability itself is used as part of the loss, that is, the training objective of the neural network is to make the transition probability in the low-dimensional space after dimensionality reduction as consistent as possible with the transition probability in the original high-dimensional space. Therefore, this hidden information will be retained during the dimensionality reduction process.

[0062] Step 101 involves preparing the input data. After collecting sufficient molecular dynamics simulation trajectory data, it is necessary to preprocess the input data. For example, perform structural alignment on the molecular conformations in the trajectory to eliminate the influence of structural changes. Subsequently, extract key input information from the aligned trajectory data, such as the coordinates of all heavy atoms, and organize it into a dataset for subsequent training.

[0063] Step 102 involves dimensionality reduction of the input dataset. Use the reversible dimensionality reduction method to project the trajectory data into a low-dimensional space to generate a low-dimensional free energy surface. This dimensionality reduction process can simultaneously retain transition state information such as the density information of the system and the transition pair information.

[0064] In some specific embodiments, the dimensionality reduction process specifically includes: setting the relaxation time and the dimensionality reduction dimension based on a preset target molecular system to map the trajectory data of the input dataset to the latent space corresponding to the dimensionality reduction dimension through a preset first dimensionality reduction method; setting the relaxation time and the number of dimensions of the low-dimensional space based on the current molecular system to map the data in the latent space to the low-dimensional space through a preset second dimensionality reduction method.

[0065] In some specific embodiments, the first dimensionality reduction method includes the time-structured independent component analysis method; and / or, the second dimensionality reduction method includes the reaction coordinate flow method, the diffusion model, or the flow model based on the variational autoencoder. Using reversible dimensionality reduction methods such as RCF and the diffusion model ensures that important transition state information is not lost during the dimensionality reduction process. Project the high-dimensional molecular conformation data into a low-dimensional latent space and perform transition state search in parallel in the low-dimensional space. The search complexity in the low-dimensional space is much lower than that in the high-dimensional space, significantly reducing the computational cost. Compared with traditional dimensionality reduction methods, the present invention can accurately recover the key state conformations in the original high-dimensional space after dimensionality reduction, ensuring the integrity of the transition state information.

[0066] Specifically, according to the characteristics of the target molecular system, an appropriate relaxation time (lag time) and the dimensionality reduction dimension d are selected, and time-lagged independent component analysis (TICA) is applied to the input data set for dimensionality reduction. During the dimensionality reduction process of TICA, a transformation matrix will be calculated and the original high-dimensional data will be mapped to the latent space. Generally, the dimension of the latent space can be set to retain about 80% of the original space information. In addition, the calculated transformation matrix and the dimensionality-reduced latent space data need to be collected for subsequent analysis and modeling.

[0067] Specifically, on the latent space data after preliminary dimensionality reduction, the reaction coordinate flow (RCF) method is further used for dimensionality reduction. First, select a relaxation time suitable for the current molecular system (the value of the previous step can be used), and set the number of dimensions of the final required low-dimensional space (reaction coordinates, RCs) (such as 1-20). Subsequently, train the RCF model and save the trained model and its corresponding low-dimensional space data for subsequent analysis and application.

[0068] Among them, the reaction coordinate flow (RCF) is a normalized flow variant that learns kinetic information, and its basic principle is as shown in the appendix Figure 3 and includes: RCF first performs the following operations on the input data x t Using TICA, according to the obtained transformation matrix M D×m Obtain the data in the latent space where s t = x t ·M + b t , m < D, D is the dimension of the original space, m is the dimension of the set latent space, and the optional range is D*(2 - 80%).

[0069] Then, an RCF model is constructed by optimizing the loss function . According to the set τ, d, the trajectory s r is further reduced to the RCs space as z t , where:

[0070]

[0071] is the average log-likelihood of transition pairs, which characterizes the kinetic information of the trajectory;

[0072]

[0073] For s t the average marginal log-likelihood of s t ), which characterizes the density distribution information of the trajectory; d is the dimensionality of the RCs space; τ is the minimum relaxation time that enables TICA.

[0074] In the RCs space, the reduced-dimensional kinetic information is Brownian dynamics is expressed, and here it can also be replaced by other kinetic models.

[0075] Step 103 involves the search for low-dimensional candidate transition state samples. The low-dimensional search methods include GAD, a variant of GAD optimized based on the Hessian matrix, and the combination of GAD and gradient descent. Preferably, the Gentlest Ascent Dynamics (GAD) is used to search for low-dimensional candidate transition state samples. GAD is a representative algorithm for transition state search. Starting from a metastable state or an arbitrary state in a preset low-dimensional space, it can directly complete the transition state search in the low-dimensional potential energy surface space. The principle of the algorithm is as follows: Starting from an arbitrary point in the low-dimensional potential energy surface space, according to the following rules

[0076]

[0077] to determine the displacement direction for moving to the next step in each iteration, that is, making a small-step movement along the direction with the smallest rate of change of the potential energy function gradient, and finally converging to the saddle point position (i.e., the transition state). Here in is the force calculated for the molecular system based on the potential energy gradient in the current low-dimensional CV space; and is set to approach the eigenvector corresponding to the minimum eigenvalue of the Hessian matrix of the potential energy function, that is, pointing to the direction with the smallest curvature, which needs to converge based on repeated iterations. During this period, γ controls the influence ability of H on the change, so as to eliminate the noise in the potential energy function. Simply put, the rules of the formula will guide the molecule to continuously climb against the trend along the direction with the gentlest potential energy slope until convergence stagnates at the transition state. It should be noted that GAD needs to calculate the Hessian matrix frequently, so its fast search is usually limited to a relatively low-dimensional space. This characteristic makes it difficult for GAD to be directly applied to biological macromolecule systems, and other methods need to be used to reduce the dimension first.

[0078] After the dimensionality reduction in step 102 of this application, due to the relatively low dimensionality of the reduced space, the calculation efficiency of GAD is significantly improved, avoiding the computational bottleneck of high-dimensional search. GAD can directly run on the equivalent potential energy surface to find candidate transition state sample points, and then inverse-project them back to the original high-dimensional space to complete the candidate CA test and accurately determine the transition state structure.

[0079] In some specific embodiments, the search process for low-dimensional candidate transition state samples includes: determining the starting position and initial search direction of the search, and optimizing each starting position according to the initial search direction; performing searches at all starting positions according to the set search step size and number of iteration steps to obtain the search trajectory of each step; and extracting corresponding data points as low-dimensional candidate transition state samples by evaluating the convergence of each search trajectory.

[0080] Selecting the starting position: The starting position can be any point in the low-dimensional space, usually the position where metastable states (MS) are located. The position of the metastable state can be obtained by clustering methods such as Density Peaks, or determined based on prior knowledge.

[0081] Initializing the search direction for the first step: To ensure as comprehensive a search as possible in the low-dimensional space, the initial search direction n needs to cover a variety of possibilities. For example, in a low-dimensional space of dimension n, 3 initial directions can be generated for each dimension, and the specific rules are as follows: move one unit in the negative direction of this dimension; move one unit in the positive direction of this dimension; stay at the origin unchanged. By combining the above three choices in different dimensions, 3^n initial directions can be generated, thus efficiently covering the possible search paths in the low-dimensional space. For example, in a 2RCs (two reaction coordinates) reduced space, 9 initial directions can be generated; in a 4RCs reduced space, 81 initial directions can be generated.

[0082] Optimizing the starting position: The following two methods can be used to improve the search efficiency: (a) Set the end point of the first step of the search path to the point farthest from the initial metastable state (MS) within the cluster where the initial MS is located, along the specified initial search direction. (b) Set the end point of the first step of the search path to the point that moves multiple unit distances along the specified initial direction, which can be set to 1 - 5 units.

[0083] Setting GAD parameters: Select a suitable search step size (such as 0.05), set γ = 10, and set the number of iteration steps, which can be set to 300.

[0084] Performing GAD search: Perform GAD searches at all starting positions.

[0085] Screen out candidate transition states: Evaluate the convergence of the GAD trajectory, for example, compare the Root-Mean-Square-Distance (RMSD) of step 200 and step 300. For GAD search paths with RMSD less than the critical value, extract the data points corresponding to step 200 as candidate transition states in the low-dimensional space.

[0086] Step 104 involves obtaining a candidate transition state structure based on a low-dimensional candidate transition state sample, and a variety of methods can be used to obtain the candidate transition state structure. In some specific embodiments, the process of obtaining the candidate transition state structure includes: using a reversible dimensionality reduction method to restore the low-dimensional candidate transition state sample from a low-dimensional space to a high-dimensional space to reconstruct the candidate transition state structure; specifically, reconstructing the candidate transition state by RCF and TICA back projection: for the candidate transition state in the low-dimensional space, use RCF and TICA to back-project to the high-dimensional conformation space to reconstruct the candidate transition state structure. The low-dimensional transition state is inversely mapped back to the high-dimensional space to ensure accurate reconstruction of the transition state structure.

[0087] Step 105 involves verifying the candidate transition state structure. In some specific embodiments, the committor probability of each candidate transition state structure is calculated, and if the committor probability of a candidate transition state structure meets a preset interval, the candidate transition state structure is determined to be a true transition state structure. By calculating the committor probability of the transition state structure, it is verified whether it is a true transition state.

[0088] In some specific embodiments, the process of obtaining the true transition state structure includes: performing molecular dynamics simulation on each candidate transition state structure to obtain multiple molecular dynamics trajectories; determining the metastable state that each molecular dynamics trajectory approaches first, and calculating the committor probability based on the number of different metastable states that all molecular dynamics trajectories approach first; if the committor probability of a candidate transition state structure is within a preset interval, the candidate transition state structure is identified as the true transition state structure. Dynamic verification is performed by combining Committor Analysis (CA) to ensure that the identified transition state is indeed located at the dynamic boundary between two stable states. This method significantly improves the physical reliability of the model. When performing molecular dynamics simulation, ensure that the initial position of each trajectory is the same, but the initial velocity (direction of movement) is different, and multiple molecular dynamics trajectories are obtained at one time.

[0089] First, run CMD (Classical Molecular Dynamics) to generate candidate transition state verification data: for each candidate transition state, run CMD with multiple molecular dynamics trajectories (50 to 100), and the simulation time range of each trajectory can be set to 50ps-2ns.

[0090] Next, analyze the CA results to judge the authenticity of the candidate transition states: First, the metastable state that each molecular dynamics trajectory approaches first can be determined. For each molecular dynamics trajectory, identify the moment when its distance from the metastable state is less than the critical value (RMSD ), so as to identify the metastable state that the trajectory tends to. Then, count the trajectories that tend to and calculate the committor probability: Count the number of all trajectories that first fall into metastable state A and metastable state B, and calculate the committor probability according to the formula:

[0091]

[0092] If the committor probability of a candidate transition state structure is between 40% and 60%, then this candidate transition state is considered a real transition state.

[0093] Since the search scheme of the present application combines generative modeling and can model the stable state and transition state more clearly in the low-dimensional latent space, even in the face of complex biomolecular systems, it can accurately distinguish the key transition regions between different states.

[0094] In some embodiments, the search scheme of the present application can be extended to other fields, including:

[0095] Chemical reaction kinetics: It can be used for chemical reaction path search, predicting reaction intermediate states and rate-determining steps.

[0096] Materials science: It can be used for predicting material phase transitions and interface migration paths.

[0097] Protein-ligand interaction prediction: It can be used for small molecule drug design to find the key intermediate states of receptor-ligand binding.

[0098] In addition, in addition to biological macromolecules such as proteins, the search scheme of the present application is also applicable to small molecule systems, nanomaterials or polymer systems. By adjusting the dimensionality reduction strategy, it can be applicable to ultra-high dimensional systems (such as modeling complex biological networks).

[0099] In practical applications, the search method of the present application is mainly implemented on a GPU / high-performance computing cluster, but it can be replaced by quantum computing optimization or low-power AI chips to achieve efficient computing. A distributed computing scheme is further used to improve the feasibility of large-scale simulations.

[0100] To verify the effectiveness of the search method of the present application, RCF is used for dimensionality reduction and GAD is used for search, and it is applied to the conformational transition of a relatively complex biomolecular system - the T4 lysozyme L99A variant (T4L-L99A) from the ground state (G state) to the excited state (E state) in an explicit solvent environment, as Figure 4 shown. Figure 4Shows the metastable state distribution during the transition of T4 lysozyme L99A variant from the ground state to the excited state. The ground state (MS1) and excited state (MS5) structures of T4L-L99A, highlighting the key residues M106, F114, and W138, as well as the α0 and α1 helices. During the transition from the ground state to the excited state, flipping of F114 and W138, elevation of M106, and rearrangement of the α0 and α1 helices occur. Among them, the α0 helix corresponds to residues 114 - 122 in the ground state (G state) and residues 119 - 122 in the excited state (E state), and the α1 helix corresponds to residues 107 - 113 in the G state and residues 107 - 118 in the E state. In previous studies, a low free energy path (LFEP) containing five metastable states (MSs) and four transition states (TSs) was determined using the Traveling Salesman Automatic Path Search (TAPS) protocol. In this experimental test study, the LFEP nodes identified by TAPS (i.e., conformations containing the entire system and solvent) were used as the initial conformations, and 4 sets of unbiased molecular dynamics (MD) simulations were performed for each path node, each lasting 50 ns, resulting in a total of 2650 ns of unbiased trajectory data. Subsequently, we trained the RCF model on this T4L-L99A dataset and reduced its dimension to 4 reaction coordinates (RCs).

[0101] After performing GAD search in 4-dimensional space, three transition states verified by Committor Analysis (CA) were successfully located, namely TS12 (MS1 / 2), TS23 (MS2 / 3), and TS45 (MS4 / 5), as Figure 5 shown. Figure 6 Each of the small plots A–E of shows the key conformational changes related to the identified transition states. (A) TS12, located between MS1 and MS2, shows the downward movement of the benzene ring of F114. (B) TS23, located between MS2 and MS3, depicts the movement of the benzene ring of F114 flipping to the right. (C) TS23, located between MS2 and MS3, shows a slight increase in the angle between the α0 and α1 helices. (D) TS45, located between MS4 and MS5, captures the flipping transition of W138. (E) TS45, highlights the upward movement of M106. This method successfully identified transition states that can clearly show the conformational transition process and is highly consistent with the transition states obtained by the Taps method. Although TS34 (MS3 / 4) was not identified, this may be due to the high structural similarity or low energy barrier height of MS3 / 4, but the obtained RCF-GAD transition states (TSs) are highly consistent with the TSs previously identified by TAPS, specifically as Figure 6as described in A-E. Overall, the TS structure parsed by RCF-GAD is more reasonable than that of the TAPS method. For example, in TS23 parsed by RCF-GAD, the position of F114 is closer to the middle region between MS2 / 3, which is more reasonable than TS23 parsed by TAPS, as shown in Figure 6 B as shown in Figure 6 D of

[0102] The experimental results show that the TS structure parsed by RCF-GAD is more accurate than TAPS because it combines unbiased MD trajectory data and reaction coordinate mapping, enabling the TS structure to form naturally in a real kinetic environment. These results indicate that the RCF-GAD method can be applied to actual large biomolecular systems and can be used to optimize the TS structures obtained by bias sampling-based path methods (such as the Finite Temperature String method, the Fast Tomographic method, the Path Metadynamics method, and the TAPS method).

[0103] This application provides a transition state search system for conformational changes of complex biomacromolecules. The schematic diagram of the system modules is shown in the appendix Figure 7 as follows, and the system includes the following:

[0104] Input unit 1, used to obtain an input data set involving molecular dynamics simulation trajectory data;

[0105] Dimensionality reduction unit 2, used to reduce the dimensionality of the input data set by a preset reversible dimensionality reduction method, project the trajectory data from a high-dimensional space to a low-dimensional space while retaining the transition state information of the input data set, and generate a low-dimensional free energy surface;

[0106] Search unit 3, used to search for low-dimensional candidate transition state samples on the low-dimensional free energy surface by using a preset low-dimensional search method;

[0107] Candidate unit 4, used to inversely project the low-dimensional candidate transition state samples back to the high-dimensional space by the reversible dimensionality reduction method to obtain candidate transition state structures under a complete high-dimensional representation;

[0108] Verification unit 5, used to verify the candidate transition state structures and screen out the real transition state structures according to the verification results.

[0109] The present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute a method for searching for transition states of conformational changes of complex biomacromolecules. Applying a method for searching for transition states of conformational changes of complex biomacromolecules to a computer program product facilitates execution.

[0110] The present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of a method for searching for transition states of conformational changes of complex biomacromolecules as described above.

[0111] The computer storage medium of the present application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. The present application applies a method for searching for transition states of conformational changes of complex biomacromolecules to a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the clothing simulation method provided by the present application, which is simple, fast, easy to store, and not easily lost.

[0112] The present application proposes a method, system, and application for searching for transition states of conformational changes of complex biomacromolecules, which realizes high-dimensional transition state structure search with high computational efficiency and high accuracy based on reversible dimensionality reduction and low-dimensional search, and has strong scalability and high physical reliability, and is applicable to complex biomolecular systems. Projecting high-dimensional molecular conformation data into a low-dimensional space and performing parallel transition state search in the low-dimensional space makes the search complexity much lower than that in the high-dimensional space, significantly reducing the computational cost. The search process can not only detect the transition state, but also generate reasonable transition state conformations, accurately display the key structures corresponding to the conformational change process, avoid false positive problems caused by abnormal data distribution, and effectively reduce the false positive rate of transition state recognition.

[0113] Those of ordinary skill in the art should understand that the various modules of the present application described above can be implemented using a general-purpose computing system. They can be centralized on a single computing system or distributed across a network composed of multiple computing systems. Optionally, they can be implemented using program code executable by a computer system, so that they can be stored in a storage system and executed by the computing system, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0114] Note that the above is only a preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments here, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments only. Without departing from the concept of the present application, more other equivalent embodiments can be included, and the scope of the present application is determined by the scope of the appended claims.

[0115] The above discloses only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present application.

Claims

1. A method for searching the transition state of conformational changes of complex biological macromolecules, characterized in that, It includes the following: Obtain an input data set involving molecular dynamics simulation trajectory data; Reduce the dimension of the input data set by a preset reversible dimensionality reduction method, project the trajectory data from a high-dimensional space to a low-dimensional space while retaining the transition state information of the input data set, and generate a low-dimensional free energy surface; Use a preset low-dimensional search method to search for low-dimensional candidate transition state samples on the low-dimensional free energy surface; Inverse-project the low-dimensional candidate transition state samples back to the high-dimensional space through the reversible dimensionality reduction method to obtain candidate transition state structures under a complete high-dimensional representation; Verify the candidate transition state structures, and screen out the true transition state structures according to the verification results.

2. The method for searching the transition state of the conformational change of complex biological macromolecules according to claim 1, characterized in that, The dimensionality reduction process specifically includes: Set the relaxation time and the dimensionality reduction dimension based on a preset target molecular system, so as to map the trajectory data of the input data set to the latent space corresponding to the dimensionality reduction dimension through a preset first dimensionality reduction method; Set the relaxation time and the number of dimensions of the low-dimensional space based on the current molecular system, so as to map the data in the latent space to the low-dimensional space through a preset second dimensionality reduction method.

3. The method for searching the transition state of conformational changes of complex biomacromolecules according to claim 1, wherein The search process of the low-dimensional candidate transition state samples includes: Determine the starting position and the initial search direction of the search, and optimize each starting position according to the initial search direction; Search at all starting positions according to the set search step size and the number of iteration steps to obtain the search trajectory of each step; Extract the corresponding data points as low-dimensional candidate transition state samples by evaluating the convergence of each search trajectory.

4. The method for searching the transition state of conformational changes of complex biomacromolecules according to claim 2, wherein The first dimensionality reduction method includes the time-structured independent component analysis method; And / or, the second dimensionality reduction method includes the reaction coordinate flow method, the diffusion model, or the flow model based on the variational autoencoder.

5. The method for searching the transition state of conformational changes of complex biomacromolecules according to claim 2, wherein The low-dimensional search method includes GAD, the variant of GAD optimized based on the Hessian matrix, and the combination of GAD and gradient descent.

6. The method for searching the transition state of conformational changes of complex biological macromolecules according to claim 1, characterized in that, Calculate the committor probability of each candidate transition state structure. If the committor probability of a certain candidate transition state structure meets the preset interval, then it is determined that this candidate transition state structure is the true transition state structure.

7. The method for searching the transition state of the conformational change of complex biomacromolecules according to claim 1, characterized in that, The process of obtaining the true transition state structure includes: Perform molecular dynamics simulations on each candidate transition state structure to obtain multiple molecular dynamics trajectories; Determine the metastable state that each molecular dynamics trajectory approaches first, and calculate the committor probability according to the number of different metastable states that all molecular dynamics trajectories approach first; If the committor probability of a certain candidate transition state structure is within the preset interval, then it is determined that this candidate transition state structure is the true transition state structure.

8. A transition state search system for conformational changes of complex biological macromolecules, characterized in that, It includes the following: An input unit for obtaining an input data set involving molecular dynamics simulation trajectory data; A dimensionality reduction unit for reducing the dimension of the input data set by a preset reversible dimensionality reduction method, projecting the trajectory data from a high-dimensional space to a low-dimensional space while retaining the transition state information of the input data set, and generating a low-dimensional free energy surface; A search unit for using a preset low-dimensional search method to search for low-dimensional candidate transition state samples on the low-dimensional free energy surface; A candidate unit for inverse-projecting the low-dimensional candidate transition state samples back to the high-dimensional space through the reversible dimensionality reduction method to obtain candidate transition state structures under a complete high-dimensional representation; A verification unit, configured to verify candidate transition state structures and screen out true transition state structures according to the verification results.

9. A computer device, characterized in that, The computer device includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method for searching for transition states of conformational changes of complex biological macromolecules according to any one of claims 1-7.

10. A computer program product, characterized in that, Including executable instructions that, when executed by a processor, implement the method for searching for transition states of conformational changes of complex biological macromolecules according to any one of claims 1-7.

Citation Information

Patent Citations

  • Transient state determination method and device

    CN115684499A

  • Chemical reaction transition state search system, chemical reaction transition state search method, and chemical reaction transition state search program

    US20110191390A1

  • Computer implemented method and system for small molecule drug discovery

    US20230317202A1