Molecular motion trajectory generation method, device, terminal and storage medium

By pre-training the generative model to generate molecular motion trajectories frame by frame, the problem of difficulty in efficiently and accurately predicting ligand escape trajectories in existing technologies is solved, and fast and accurate molecular motion trajectory simulation is achieved.

CN120564898BActive Publication Date: 2025-10-03GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE (INTERNATIONAL ADVANCED TECHNOLOGY APPLICATION PROMOTION CENTER (SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511062460.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-03
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing methods for simulating the escape trajectories of ligands from protein pockets rely on force field calculations and molecular dynamics, which make it difficult to predict the escape trajectories efficiently and accurately.

Method used

Through the pre-trained generative model, based on the escape trajectory training of the existing protein-ligand complex, the molecular motion trajectory of the target protein-ligand complex is generated frame by frame, and the generative model is used to simulate the entire process of the protein-ligand complex from the initial state to the separation state.

Benefits of technology

It accelerates the simulation process of molecular motion trajectories, reduces computing costs, improves simulation efficiency and prediction accuracy, and realizes fast and accurate generation of molecular motion trajectories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564898B_ABST
    Figure CN120564898B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, terminal and storage medium for generating molecular motion trajectories, and relates to drug discovery technology. The method comprises: using a pre-trained generative model, generating the atomic coordinate information of several subsequent frames frame by frame for the atomic category sequence of the target protein-ligand complex and the atomic coordinate information of the first frame; the pre-trained generative model is trained based on the escape trajectory of an existing protein-ligand complex; and generating the molecular motion trajectory of the target protein-ligand complex according to the atomic category sequence of the target protein-ligand complex and the atomic coordinate information of all frames. The present invention simulates the molecular motion trajectory by adopting a paradigm of generating coordinates frame by frame by a generative model, and learns the inter-frame coordinate evolution law of the ligand in the process of escaping the protein pocket through the existing escape trajectory. In the inference stage, it is only necessary to input the atomic category sequence and the atomic coordinate information of the initial state of the first frame, and the atomic coordinate information of the subsequent frames can be iteratively generated quickly and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug discovery, and in particular to a method, device, terminal and storage medium for generating molecular motion trajectories. Background Art

[0002] Molecular dynamics simulations are a core tool for analyzing the microscopic dynamics of interactions between biomacromolecules (such as proteins) and drug molecules. The trajectory of ligand escape from protein pockets is crucial for analyzing key properties such as drug binding stability, dissociation rate, and effective residence time, thus influencing drug design optimization and metabolic prediction.

[0003] Traditional escape trajectory simulation methods rely on force field calculations and molecular dynamics. This method not only has low simulation efficiency but also requires complex parameter settings, making it difficult to efficiently and accurately predict the motion trajectory of ligands escaping from protein pockets.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, device, terminal and storage medium for generating molecular motion trajectories in response to the above-mentioned defects of the prior art, aiming to solve the problem that the existing trajectory simulation method for ligand escape from protein pocket relies on force field calculation and molecular dynamics, making it difficult to predict the escape trajectory efficiently and accurately.

[0006] The technical solutions adopted by the present invention to solve the problem are as follows:

[0007] In a first aspect, an embodiment of the present invention provides a method for generating molecular motion trajectories, the method comprising:

[0008] Obtain the atomic class sequence and atomic coordinate information of the first frame of the target protein-ligand complex;

[0009] Generate atomic coordinate information for several subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generative model; the pre-trained generative model is trained based on the escape trajectory of an existing protein-ligand complex;

[0010] The molecular motion trajectory of the target protein-ligand complex is generated according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of all frames.

[0011] In one embodiment, a method for obtaining an escape trajectory of an existing protein-ligand complex comprises:

[0012] Obtaining the original escape trajectory of an existing protein-ligand complex from an initial state to a separated state from a public data set, and normalizing each frame of information in the original escape trajectory to obtain a standardized file;

[0013] From the standardized files of all frames, the atomic class sequence of the existing protein-ligand complex and the atomic coordinate information of all frames are extracted to compose the escape trajectory of the existing protein-ligand complex.

[0014] In one embodiment, the training method of the pre-trained generative model includes:

[0015] Use the flow model to build the initial generation model between two adjacent frames;

[0016] The initial generative model is trained using the escape trajectory of the existing protein-ligand complex as training data to fit the vector field of the ordinary differential equation (ODE) to obtain the pre-trained generative model.

[0017] In one embodiment, the initial generative model is trained using the escape trajectories of existing protein-ligand complexes as training data to fit the vector field of ordinary differential equations (ODEs), including:

[0018] From the escape trajectory of the existing protein-ligand complex, the atomic coordinate information of the adjacent first training frame and the second training frame is sampled and interpolated to obtain the atomic coordinate information of the training intermediate state;

[0019] Performing a differential operation on the atomic coordinate information of the training intermediate state to obtain a first vector field;

[0020] Predicting a second vector field based on the atomic coordinate information of the training intermediate state, the atomic category sequence of the escape trajectory, and time information through an initial generative model;

[0021] The initial generation model is trained according to the first vector field and the second vector field to fit the vector field.

[0022] In one embodiment, the initial generative model is trained using the escape trajectories of existing protein-ligand complexes as training data to fit the vector field of ordinary differential equations (ODEs), including:

[0023] From the escape trajectory of the existing protein-ligand complex, the atomic coordinate information of the adjacent first training frame and the second training frame is sampled and interpolated to obtain the atomic coordinate information of the training intermediate state;

[0024] Performing a differential operation on the atomic coordinate information of the training intermediate state to obtain a first vector field;

[0025] When the first training frame is not the first frame, extracting training history information based on atomic coordinate information of several historical frames before the first training frame;

[0026] Predicting a second vector field based on the atomic coordinate information of the training intermediate state, the atomic category sequence of the escape trajectory, the training history information, and the time information through an initial generation model;

[0027] The initial generation model is trained according to the first vector field and the second vector field to fit the vector field.

[0028] In one embodiment, training the initial generative model based on the first vector field and the second vector field to fit the vector field includes:

[0029] When training the initial generation model, using the first vector field as a supervisory signal to determine whether the motion trend reflected by the second vector field predicted by the initial generation model is accurate;

[0030] The gap between the first vector field and the second vector field is converged to fit the vector field.

[0031] In one embodiment, converging the gap between the first vector field and the second vector field to fit the vector field includes:

[0032] Calculating a model loss value according to the first vector field and the second vector field;

[0033] Determine whether the training termination condition is currently met. If not, backpropagate the initial generation model according to the model loss value to update the network parameters;

[0034] The step of sampling atomic coordinate information of adjacent first training frames and second training frames from the escape trajectory of the existing protein-ligand complex and performing interpolation processing is continued until the training termination condition is reached, and the current initial generation model is used as the pre-trained generation model.

[0035] In one embodiment, the atomic coordinate information of the subsequent frames is generated frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generative model, including:

[0036] Generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and time information using the pre-trained generative model;

[0037] Determine whether the iteration termination condition is currently satisfied, and if not, use the atomic coordinate information of the second frame as the atomic coordinate information of the first frame;

[0038] Continue to execute the step of generating the atomic coordinate information of the second frame according to the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and the time information through the pre-trained generation model until the iteration termination condition is met.

[0039] In one embodiment, generating atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and time information by using the pre-trained generative model includes:

[0040] By using the pre-trained generative model and the numerical integration solution method, based on the atomic category sequence, the atomic coordinate information of the first frame, and the time information, the atomic coordinate information of the intermediate state is gradually solved based on time increments; the time increments are obtained based on the frame spacing;

[0041] When the accumulated time increment reaches the frame interval, the atomic coordinate information of the second frame is determined according to the atomic coordinate information of the intermediate state currently solved.

[0042] In one embodiment, the atomic coordinate information of the subsequent frames is generated frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generative model, including:

[0043] When the first frame is not the first frame, extracting historical information based on atomic coordinate information of several historical frames before the first frame;

[0044] Generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information, and the time information through the pre-trained generative model;

[0045] Determine whether the iteration termination condition is currently satisfied, and if not, use the atomic coordinate information of the second frame as the atomic coordinate information of the first frame;

[0046] Continue to execute the step of generating the atomic coordinate information of the second frame according to the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information and the time information through the pre-trained generation model until the iteration termination condition is met.

[0047] In one embodiment, generating atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information, and the time information by using the pre-trained generative model includes:

[0048] By using the pre-trained generative model and numerical integration solution, based on the atomic category sequence, the atomic coordinate information of the first frame, the historical information, and the time information, the atomic coordinate information of the intermediate state is gradually solved based on time increments; the time increments are obtained based on the frame spacing.

[0049] When the accumulated time increment reaches the frame interval, the atomic coordinate information of the second frame is determined according to the atomic coordinate information of the intermediate state currently solved.

[0050] In one embodiment, when the task of the pre-trained generative model is to generate the escape trajectory of the target protein-ligand complex, the iteration termination condition is: the distance value between the protein surface and the ligand calculated based on the atomic coordinate information of the second frame is less than a preset value.

[0051] In a second aspect, an embodiment of the present invention further provides a molecular motion trajectory generating device, the device comprising:

[0052] An acquisition module is used to obtain the atomic class sequence and atomic coordinate information of the first frame of the target protein-ligand complex;

[0053] a generation module, configured to generate atomic coordinate information for a plurality of subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generation model; the pre-trained generation model is trained based on the escape trajectory of an existing protein-ligand complex;

[0054] The summarizing module is used to generate the molecular motion trajectory of the target protein-ligand complex according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of all frames.

[0055] In a third aspect, an embodiment of the present invention further provides a terminal comprising a memory and one or more processors; the memory stores one or more programs; the programs include instructions for executing any of the molecular motion trajectory generation methods described above; and the processor is used to execute the programs.

[0056] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded and executed by a processor to implement the steps of any of the above-described methods for generating molecular motion trajectories.

[0057] Beneficial effects of the present invention: The embodiment of the present invention equates the problem of generating molecular motion trajectories to the problem of generating video frames. One frame corresponds to one step length, and each frame can reflect the three-dimensional spatial coordinates of the protein and small molecules under the corresponding step length. By simulating the entire process of the protein-ligand complex from the initial state to the separation state through the generative model, the simulation process of the molecular motion trajectory can be accelerated. During the training phase, the pre-trained generative model learns the inter-frame coordinate evolution law of the ligand in the process of escaping the protein pocket through the existing escape trajectory. Therefore, during the inference phase, only the atomic category sequence and the atomic coordinate information of the initial state of the first frame need to be input, and the atomic coordinate information of the subsequent frames can be iteratively generated quickly and accurately, which greatly reduces the computational cost and improves the simulation efficiency and prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 4 is a flow chart of a method for generating molecular motion trajectories provided in an embodiment of the present invention.

[0060] Figure 2 3 is a schematic diagram of a frame-by-frame generation process of an escape trajectory provided by an embodiment of the present invention.

[0061] Figure 3 Schematic diagram of the protein-ligand complex from the binding state to the dissociation state provided by the embodiment of the present invention.

[0062] Figure 4 Schematic diagram of a module of a molecular motion trajectory generating device provided by an embodiment of the present invention.

[0063] Figure 5 This is a principle block diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The present invention discloses a method, apparatus, terminal, and storage medium for generating molecular motion trajectories. To clarify the objectives, technical solutions, and effects of the present invention, the present invention is further described below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention.

[0065] Those skilled in the art will understand that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" as used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0066] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0067] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method for generating molecular motion trajectories, the method comprising: obtaining an atomic category sequence and atomic coordinate information of a first frame of a target protein-ligand complex; generating atomic coordinate information of several subsequent frames frame by frame based on the atomic category sequence and atomic coordinate information of the first frame of the target protein-ligand complex using a pre-trained generative model; the pre-trained generative model is trained based on the escape trajectory of an existing protein-ligand complex; and generating the molecular motion trajectory of the target protein-ligand complex based on the atomic category sequence and atomic coordinate information of all frames of the target protein-ligand complex. The present invention equates the molecular motion trajectory generation problem to the video frame generation problem, where one frame corresponds to a step length, and each frame can reflect the three-dimensional spatial coordinates of the protein and small molecule at the corresponding step length. By simulating the entire process of the protein-ligand complex from the initial state to the separated state through the generative model, the simulation process of the molecular motion trajectory can be accelerated. During the training phase, the pre-trained generative model learns the inter-frame coordinate evolution law of the ligand in the process of escaping the protein pocket through the existing escape trajectory. Therefore, in the inference phase, only the atomic category sequence and the atomic coordinate information of the initial state of the first frame need to be input to quickly and accurately iteratively generate the atomic coordinate information of subsequent frames, greatly reducing the computational cost and improving the simulation efficiency and prediction accuracy.

[0068] like Figure 1 As shown, the method specifically includes the following steps:

[0069] Step S100: Acquire the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame.

[0070] Specifically, the target protein-ligand complex in this embodiment can be any protein-ligand complex for which molecular motion trajectory simulation is required. The atomic class sequence includes the atomic class of each atom in the protein-ligand complex; the atomic coordinate information includes the three-dimensional spatial coordinates of each atom in the protein-ligand complex. This embodiment equates the molecular motion trajectory generation problem to the problem of video frame generation, with one frame corresponding to a step size, and each frame reflects the three-dimensional spatial coordinates of the protein and small molecule at the corresponding step size. This embodiment simulates and generates molecular motion trajectories by generating video frames using a generative model, saving significant time and computing power. When generating molecular motion trajectories, the initial state of the target protein-ligand complex is first extracted to obtain the atomic class sequence and the atomic coordinate information for the first frame required for modeling. This serves as the basic data for predicting the atomic coordinate information for subsequent frames. The atomic class sequence provides key information about the protein-ligand complex, such as atomic species, atomic chemical properties, and interaction rules. During the subsequent molecular motion trajectory simulation, the number and type of atoms in the atomic class sequence remain the same, with only the coordinate information differing. The atomic coordinate information of the first frame is the three-dimensional spatial coordinate of the initial state of the protein-ligand complex. The coordinate information of other frames can be gradually predicted based on the atomic coordinate information of the first frame.

[0071] Step S200: Generate atomic coordinate information of several subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generative model; the pre-trained generative model is trained based on the escape trajectory of an existing protein-ligand complex.

[0072] This embodiment equates the molecular motion trajectory generation problem to the problem of video frame generation. Therefore, a pre-trained generative model, based on a given first frame, can predict all subsequent frames, ultimately predicting the complete ligand escape process from the pocket. The trajectory of a ligand escaping a protein pocket is related to the pocket depth and the complexity of the path. The total number of frames in the entire escape trajectory typically varies for different target protein-ligand complexes. This embodiment generates molecular motion trajectories based on a pre-trained generative model, making it applicable to escape paths of varying pocket depths and complexities.

[0073] Step S300 : generating a molecular motion trajectory of the target protein-ligand complex according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of all frames.

[0074] Specifically, this embodiment adopts a frame-by-frame generation strategy to construct the molecular motion trajectory of the ligand escaping from the protein pocket frame by frame. The generative model can effectively implement this generation task, such as deep generative models such as the diffusion model and the flow matching model. The pre-trained generative model can capture the atomic coordinate change rules and patterns of the ligand in the process of escaping from the protein pocket structure by learning the escape trajectories of a large number of existing protein-ligand complexes. The target protein-ligand complex is input into the generative model after training. The generative model can predict the atomic coordinate information of each subsequent frame frame by frame based on the input atomic category sequence and the atomic coordinate information of the first frame, thereby forming the molecular motion trajectory of the target protein-ligand complex.

[0075] In one implementation, a method for obtaining an escape trajectory of an existing protein-ligand complex includes:

[0076] Obtaining the original escape trajectory of an existing protein-ligand complex from an initial state to a separated state from a public data set, and normalizing each frame of information in the original escape trajectory to obtain a standardized file;

[0077] From the standardized files of all frames, the atomic class sequence of the existing protein-ligand complex and the atomic coordinate information of all frames are extracted to compose the escape trajectory of the existing protein-ligand complex.

[0078] Specifically, the public dataset contains a large number of known raw escape trajectories of existing protein-ligand complexes, which record all atomic states of the protein-ligand complex from the bound to the dissociated state. Each frame of each raw escape trajectory corresponds to all atomic states for a time step. Each frame is saved as a standardized file. Alternatively, all raw escape trajectories are normalized to produce a standardized file that records the correspondence between step length and three-dimensional coordinates. This involves saving all frames (all states) of each raw escape trajectory as a single standardized file. This step aims to isolate the core data contained in the raw escape trajectory, reduce redundant information, and achieve structured data storage for subsequent chronological access. The atom class sequence is a fixed sequence of the types of all atoms in the protein-ligand complex that does not change over time. The atomic coordinate information refers to the three-dimensional spatial coordinates of all atoms in the protein-ligand complex at each time step, which changes dynamically over time. The atomic coordinate information for each frame can be extracted from the standardized file. Finally, the raw escape trajectory is simplified into the atom class sequence and the atomic coordinate information for all frames. This embodiment can transform the complex original escape trajectory into standardized time series data that can be understood and learned by the generation model by decomposing and reassembling the original escape trajectory, thereby reducing the amount of data and alleviating the learning burden of the subsequent generation model.

[0079] For example, raw escape trajectories of existing protein-ligand complexes were obtained from the public dataset DD-13M, which contains nearly 20,000 raw escape trajectories of protein-ligand complexes. Each raw escape trajectory was recorded and saved using atomic coordinate files in the h5md format and atom and bond type files in the topology.pdb format. In this example, after obtaining each escape trajectory of the protein-ligand complex, the MDanalysis package (an analysis software package for molecular dynamics simulations) in Python (a programming language) was used to save the information of each frame into a single PDB file (a file storing structural data in the Protein Data Bank).

[0080] Then use Python's biopython (an open source bioinformatics toolkit based on Python) and rdkit package (an open source chemical information toolkit) to read the three-dimensional spatial coordinates of each atom of the ligand and protein in the PDB file and atomic types It is understandable that each three-dimensional space coordinate Including ligand coordinates and protein surface coordinates .

[0081] Taking the original escape trajectory of an existing protein-ligand complex as an example, the three-dimensional spatial coordinates of each atom contained in the original escape trajectory in different frames can be obtained through the above steps. , and the sequence of amino acid types and ligand atom types ,in, is the initial structure, For the The three-dimensional spatial coordinates of each atom in the frame, is the number of atoms, is the number of frames, is a multidimensional array space composed of real numbers. The three-dimensional spatial coordinates of all atoms in each frame are defined as the atomic coordinate information for that frame. Finally, the information contained in the original escape trajectory is simplified to the atomic coordinate information and atom type sequence for all frames, thus obtaining the escape trajectory of the existing protein-ligand complex. It should be noted that the number and type of atoms in each frame of the protein-ligand complex in a single escape trajectory are the same; only the three-dimensional spatial coordinates of the atoms differ.

[0082] In one implementation, the training method of the pre-trained generative model includes:

[0083] Use the flow model to build the initial generation model between two adjacent frames;

[0084] The initial generative model is trained using the escape trajectory of the existing protein-ligand complex as training data to fit the vector field of the ordinary differential equation (ODE) to obtain the pre-trained generative model.

[0085] Specifically, the initial generative model refers to an untrained generative model. This embodiment uses a flow model to construct an initial generative model between two adjacent frames, which can provide an explainable computational framework for the generation process of the molecular motion trajectory of the protein-ligand complex from the binding state to the dissociation state. It learns the inter-frame coordinate evolution law through the escape trajectory of the existing protein-ligand complex, and transforms the molecular trajectory generation into a vector field estimation problem. The purpose of training the initial generative model is to fit the vector field so that the vector field predicted by the model is close to the real vector field. The generative model obtained after training is the pre-trained generative model.

[0086] For example, the pre-trained generative model in this embodiment can be designed based on a diffusion model or a flow model.

[0087] Take the stream model as an example. The stream model is a model used for distributed transformation. Its design principles are as follows:

[0088] A flow is defined as: , the vector field of the associated ordinary differential equation (ODE) is: :

[0089] ;

[0090] Where, Represents a flow; represents the time parameter in the flow model, ; Represents the three-dimensional space coordinates of the molecule; Indicates trajectory; Denote the initial structure of the protein-ligand complex (i.e., the initial value of the flow model), let ; Represents the vector field to be fitted, which is used to guide the model in time parameters How should the current coordinates change when ? Represents the differential operation symbol, and the left side of the equation describes the "coordinate parameter with time The rate of change expresses the coordinate change per time parameter unit.

[0091] If known , you can use Transform the two distributions by integrating the differential equations. , , by solving Available to The transformation relationship (such as Figure 2 shown):

[0092] ;

[0093] Where, represents the initial frame structure of the protein-ligand complex, Indicates the next frame structure; express distribution of express distribution.

[0094] It can be seen that the vector field provides an intuitive way to describe molecular motion. The ordinary differential equation establishes a mathematical connection from the vector field to the change of molecular position. By solving the integral of the differential equation, the position at different times can be calculated. Therefore, it is necessary to determine a suitable vector field. ( After the model is formed into a neural network, the flow model can be used as the generative model between two adjacent frames. Since the exact form of the vector field in the molecular system is difficult to obtain through theoretical analysis, this embodiment uses a neural network to fit the vector field in the ODE. Therefore, this embodiment needs to make full use of the escape trajectory of the existing protein-ligand complex to train the generative model to fit the vector field. , which enables it to accurately describe the inter-frame coordinate motion of the ligand in the process of escaping from the protein pocket.

[0095] In one implementation, the initial generative model is trained using the escape trajectory of an existing protein-ligand complex as training data to fit the vector field of an ordinary differential equation (ODE), including:

[0096] From the escape trajectory of the existing protein-ligand complex, the atomic coordinate information of the adjacent first training frame and the second training frame is sampled and interpolated to obtain the atomic coordinate information of the training intermediate state;

[0097] Performing a differential operation on the atomic coordinate information of the training intermediate state to obtain a first vector field;

[0098] Predicting a second vector field based on the atomic coordinate information of the training intermediate state, the atomic category sequence of the escape trajectory, and time information through an initial generative model;

[0099] The initial generation model is trained according to the first vector field and the second vector field to fit the vector field.

[0100] During the training phase, the untrained generative model is defined as the initial generative model, and the trained generative model is defined as the pre-trained generative model. The escape trajectory of an existing protein-ligand complex is used as training data, and through vector field fitting, the initial generative model learns the dynamic changes in the three-dimensional spatial coordinates of atoms between frames. The first training method provided in this embodiment includes sampling two adjacent frames from the existing escape trajectory, with the earlier one defined as the first training frame and the later one as the second training frame. The atomic coordinate information of the two frames is interpolated to obtain the atomic coordinate information of the intermediate training state that transitions between the two frames. The atomic coordinate information of the intermediate training state allows the initial generative model to learn the conformational changes between the two frames, thereby learning more detailed and continuous dynamic changes. Each atom is represented by a three-dimensional vector representing its next movement trend. The three-dimensional vectors of all atoms together form a vector field, which can reflect the evolution from the intermediate training state to the next frame. Differentiation of the atomic coordinate information of the intermediate training state yields the first vector field. The atomic coordinates, atomic class sequence, and time information of the training intermediate state are input into the initial generative model, which can predict the second vector field. The parameters of the initial generative model are adjusted based on the first and second vector fields to fit the vector field.

[0101] For example, to realize the atomic coordinate information based on the current frame Generate atomic coordinate information for the next frame , we need to build a series of Map to The ordinary differential equation for :

[0102] ;

[0103] , ;

[0104] ;

[0105] make , in The value range is , Indicates the time parameter Next, from Initial coordinates of the frame Starting from, the intermediate coordinates after flow transformation. Initial and end conditions: : Time parameter =0, the initial coordinates are (starting point). : Time parameter =1, the coordinates become the next frame (end).

[0106] But the real vector field It is difficult to obtain directly, this embodiment uses a neural network To fit the vector field ,in, is the time step, represents a sequence of atomic types, From the initial coordinates Departure to time parameters The atomic coordinates of the intermediate state (corresponding to the flow model ).

[0107] For the first vector field, design the intermediate state to complete the flow Among them, from the time parameter =0 to =1, the atomic coordinate information of the current frame can be completed Atomic coordinate information to the next frame transformation.

[0108] This embodiment uses linear interpolation as an example to complete the flow transformation. Based on the atomic coordinate information of the current frame and the atomic coordinate information of the next frame Perform linear interpolation to obtain the atomic coordinate information of the intermediate state :

[0109] , formula (1);

[0110] In addition to linear interpolation, other interpolation methods may also be used, which are not specifically limited in this embodiment.

[0111] According to the atomic coordinate information of the intermediate state , and get the first vector field :

[0112] , formula (2).

[0113] During the training phase, the neural network The vector field predicted based on the input data is defined as the second vector field. The initial generative model is trained using the first vector field and the second vector field to implement the vector field fitting process.

[0114] In another implementation, the initial generative model is trained using the escape trajectory of an existing protein-ligand complex as training data to fit the vector field of an ordinary differential equation (ODE), including:

[0115] From the escape trajectory of the existing protein-ligand complex, the atomic coordinate information of the adjacent first training frame and the second training frame is sampled and interpolated to obtain the atomic coordinate information of the training intermediate state;

[0116] Performing a differential operation on the atomic coordinate information of the training intermediate state to obtain a first vector field;

[0117] When the first training frame is not the first frame, extracting training history information based on atomic coordinate information of several historical frames before the first training frame;

[0118] Predicting a second vector field based on the atomic coordinate information of the training intermediate state, the atomic category sequence of the escape trajectory, the training history information, and the time information through an initial generation model;

[0119] The initial generation model is trained according to the first vector field and the second vector field to fit the vector field.

[0120] During the training phase, this embodiment also provides a training method that combines historical information, so that the initial generation model can combine past states to more accurately predict future atomic motions. The specific steps of the second training method provided in this embodiment include: sampling two adjacent frames from the existing escape trajectory, the one in front is defined as the first training frame, and the one in the back is defined as the second training frame, and interpolating the atomic coordinate information of the two to obtain a training intermediate state for transition between the two frames. Differentiation operation is performed on the atomic coordinate information of the training intermediate state to obtain a first vector field. When the first training frame is not the first frame, it means that there is at least one frame before the first training frame, which can be used to extract training history information. The preset number of historical frames before the first training frame is used as the basic data for extracting training history information, and the specific value of the preset number can be dynamically adjusted based on actual conditions. Among them, the preset number of historical frames can be continuous historical frames, historical frames with a certain step size pattern, or randomly selected historical frames. The atomic coordinates, atomic class sequence, training history, and time information from the intermediate training states are input into the initial generative model. This model then predicts the second vector field. The training history serves as conditional information for the initial generative model, providing constraints that ensure the initial generative model maintains a motion trend consistent with the atomic coordinate changes in the historical frame when predicting the second vector field. Finally, the parameters of the initial generative model are adjusted using the first and second vector fields to fit the vector field.

[0121] For example, the input data is the atomic coordinate information of the first training frame And the atomic coordinate information of the second training frame , The value range of , and the atomic type sequence . The atomic coordinate information of the training intermediate state is calculated through formula (1) .

[0122] The atomic coordinate information of the past consecutive K frames (for example, K = 9) corresponding to the first training frame is used as conditional information (condition) and input into the neural network. If t < K, the selected historical frames only go up to , that is, the historical information .

[0123] and will be input into the neural network together . The neural network can be set as a neural network directly used to predict three-dimensional coordinates, such as a multi-layer perceptron MLP. Stochastic gradient descent using the Adam optimizer is used for iterative training until convergence reaches the training termination condition. For example, the training termination condition is that the loss between two adjacent iterations is less than a preset threshold, such as 0.1. The generated model obtained after training is used as the pre-trained generated model . <00​​​​​​​​​​​​​​​​​​​​Specifically, during the training of the initial generative model, the primary training goal is to converge the gap between the first and second vector fields in order to achieve a fitted vector field. In this embodiment, the first vector field, reflecting the molecular motion trend, is used as the true motion trend. During training, the first vector field is used as a supervisory signal to determine whether the second vector field predicted by the initial generative model accurately reflects the atomic motion trend. By continuously adjusting the model parameters of the initial generative model, the gap between the first and second vector fields is gradually narrowed. The final model parameters are then used to generate the final generative model, ultimately achieving accurate prediction of the molecular motion trend by the generative model.

[0129] In one implementation, converging a gap between the first vector field and the second vector field to fit the vector field includes:

[0130] Calculating a model loss value according to the first vector field and the second vector field;

[0131] Determine whether the training termination condition is currently met. If not, backpropagate the initial generation model according to the model loss value to update the network parameters;

[0132] The step of sampling atomic coordinate information of adjacent first training frames and second training frames from the escape trajectory of the existing protein-ligand complex and performing interpolation processing is continued until the training termination condition is reached, and the current initial generation model is used as the pre-trained generation model.

[0133] Specifically, during the training phase, the model loss value is calculated by comparing the first vector field and the second vector field. The smaller the model loss value, the better the model performance. The training termination condition may be that the model loss value reaches a preset threshold, or the number of iterations reaches a preset number. If the training termination condition is not currently met, the backpropagation algorithm is used to adjust the network parameters of the initial generative model based on the calculated model loss value. Repeat the above steps, resampling different adjacent frames from the escape trajectory to train the initial generative model in each iteration until the training termination condition is met. When the training is terminated, the initial generative model at this time is used as the pre-trained generative model, and the pre-trained generative model can be directly used to predict atomic coordinate information.

[0134] For example, the calculation principle of the model loss value of the fitted vector field is as follows: :

[0135] Where, express distribution of express distribution of Express expectations.

[0136] Based on the first training method, the training goal is to calculate the model loss value using formula (3):

[0137] , formula (3);

[0138] Where, express distribution of express express expectations; Indicates the frame number; , Represents the time parameter.

[0139] Based on the second training method, the training goal is to calculate the model loss value using formula (4):

[0140] , formula (4);

[0141] Where, express distribution of express express expectations; Represents historical information; Frame number; , Represents the time parameter.

[0142] In one implementation, the pre-trained generative model is used to generate atomic coordinate information of subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame, including:

[0143] Generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and time information using the pre-trained generative model;

[0144] Determine whether the iteration termination condition is currently satisfied, and if not, use the atomic coordinate information of the second frame as the atomic coordinate information of the first frame;

[0145] Continue to execute the step of generating the atomic coordinate information of the second frame according to the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and the time information through the pre-trained generation model until the iteration termination condition is met.

[0146] Specifically, the pre-trained generative model is a fully trained generative model that has learned the movement patterns of atoms between adjacent frames through a large number of escape trajectories. If the pre-trained generative model uses the first training method, then during the inference phase, the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and the time information are input into the pre-trained generative model. The pre-trained generative model then performs a cyclic process of predicting the second frame from the first frame, generating the molecular motion trajectory of the target protein-ligand complex frame by frame until the iteration termination condition is met.

[0147] Furthermore, generating atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame through the pre-trained generative model includes:

[0148] By using the pre-trained generative model and the numerical integration solution method, based on the atomic category sequence, the atomic coordinate information of the first frame, and the time information, the atomic coordinate information of the intermediate state is gradually solved based on the time increment; the time increment is obtained based on the frame spacing;

[0149] When the accumulated time increment reaches the frame interval, the atomic coordinate information of the second frame is determined according to the atomic coordinate information of the intermediate state currently solved.

[0150] Frame spacing refers to the time interval between two adjacent frames. Time increment refers to dividing the time interval between two adjacent frames into multiple small time periods, each of which is a time increment. By gradually calculating the atomic coordinate information within each time increment, the complete motion process is finally obtained. This embodiment combines time increment subdivision with integral solution to make the transition between frames more consistent with physical continuity. Specifically, the atomic category sequence, the atomic coordinate information of the first frame, and the time information are used as input data of the pre-trained generative model, and the pre-trained generative model performs continuous calculations of multiple time increments based on the input data to simulate the trajectory of the continuous motion of each atom in the target protein-ligand complex. At the same time, the vector field predicted by the pre-trained generative model is converted into actual three-dimensional space coordinate changes through the numerical integration solution method. When the cumulative time increment reaches the frame spacing, the atomic coordinate information of the intermediate state solved at this time is the atomic coordinate information of the second frame.

[0151] In another implementation, the pre-trained generative model is used to generate atomic coordinate information of several subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame, including:

[0152] When the first frame is not the first frame, extracting historical information based on atomic coordinate information of several historical frames before the first frame;

[0153] Generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information, and the time information through the pre-trained generative model;

[0154] Determine whether the iteration termination condition is currently satisfied, and if not, use the atomic coordinate information of the second frame as the atomic coordinate information of the first frame;

[0155] Continue to execute the step of generating the atomic coordinate information of the second frame according to the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information and the time information through the pre-trained generation model until the iteration termination condition is met.

[0156] Specifically, if the pre-trained generative model adopts the second training method, and the first frame is not the first frame, indicating that there is at least one frame of data before the first frame, then in the inference stage, the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information extracted based on the preset number of historical frames before the first frame, and the time information are required to be input into the pre-trained generative model, and the pre-trained generative model is used to predict the second frame from the first frame in a cycle, and the molecular motion trajectory of the target protein-ligand complex is generated frame by frame until the iteration termination condition is met. In this embodiment, the preset number of historical frames can be continuous historical frames, or historical frames with a certain step size pattern, or randomly selected historical frames. Historical information is extracted from the historical frame information as conditional information and input into the pre-trained generative model, avoiding simplifying the molecular motion trajectory generation into a static mapping "determined only by the current coordinates", so that the molecular motion trajectory generated frame by frame is more consistent with the continuity of the change.

[0157] Furthermore, the pre-trained generative model is used to generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information, and the time information, including:

[0158] By using the pre-trained generative model and numerical integration solution, based on the atomic category sequence, the atomic coordinate information of the first frame, the historical information, and the time information, the atomic coordinate information of the intermediate state is gradually solved based on time increments; the time increments are obtained based on the frame spacing.

[0159] When the accumulated time increment reaches the frame interval, the atomic coordinate information of the second frame is determined according to the atomic coordinate information of the intermediate state currently solved.

[0160] Specifically, the pre-trained generative model is fed with the atomic class sequence, the atomic coordinates of the first frame, historical information, and time information. The model then performs continuous calculations at multiple time increments based on this input data to simulate continuous motion. Simultaneously, a numerical integration method is used to convert the vector field predicted by the pre-trained generative model into actual three-dimensional spatial coordinate changes. When the cumulative time increment reaches the frame interval, the atomic coordinates of the intermediate state calculated are used as the atomic coordinates of the second frame.

[0161] For example, to obtain the structural information of a protein-ligand complex in the binding state, that is, the initial structure , and get the atomic type sequence As shown in formula (5), solve the following N numerical integrals to sample the molecular motion trajectory of the protein-ligand complex: :

[0162] , formula (5);

[0163] Where, Represents historical information; Represents the atomic coordinate information of the first frame; Represents the atomic coordinate information of the second frame; represents the pre-trained generative model; represents the time parameter of the flow model, ; Indicates the step size of the current frame.

[0164] The method for solving the numerical integral may be iterative sampling by using the Euler method, the explicit center point method in the Runge-Kutta method, etc., which is not specifically limited in this embodiment.

[0165] Taking the Euler method as an example, the atomic coordinate information of the first frame is , the atomic coordinate information of the second frame is obtained according to the following formula (6): :

[0166] , formula (6);

[0167] Where, , Represents a time increment.

[0168] In one implementation, when the task of the pre-trained generative model is to generate the escape trajectory of the target protein-ligand complex, the iteration termination condition is: the distance value between the protein surface and the ligand calculated based on the atomic coordinate information of the second frame is less than a preset value.

[0169] Specifically, the iteration termination conditions of pre-trained generative models vary for different tasks. For example, if the task is to generate the molecular motion trajectory of the target protein-ligand complex within a preset time period, the iteration termination condition is set based on time or the number of iterations. If the task is to generate the escape trajectory of the target protein-ligand complex, the iteration termination condition is set based on the distance between the protein surface and the ligand. Whether the ligand has escaped the pocket can be determined by whether the current ligand has touched the protein surface.

[0170] For example, a complete escape trajectory refers to the entire process of the protein-ligand complex from the initial state to the separated state. Assume that the task goal is to sample the complete escape trajectory of the ligand escaping from the outside of the protein. The generation problem of molecular motion trajectory is equivalent to the generation problem of video frames. One frame corresponds to one step length. The trajectory length can be defaulted to N, and N can be dynamically set to determine the first step length. Whether the protein-ligand complex is in a dissociated state at the time of the frame.

[0171] Each three-dimensional space coordinate Including ligand coordinates and protein surface coordinates, the current ligand coordinates are defined as , the protein surface coordinates are expressed as By calculating the distance between the ligand coordinates and the protein surface coordinates , to determine whether the ligand touches the protein surface. , is angstroms, the protein-ligand complex is considered to be The frame is in a separated state, which can interrupt the iteration process of the current model generation and determine the three-dimensional space coordinates of the separated state according to the last frame. , the molecular motion trajectory at this time This is the complete escape trajectory. Figure 2 As shown, is the three-dimensional coordinate of the t-th frame, which shows that the ligand escapes from the protein pocket in the protein-ligand complex at this frame. Figure 3 As shown, the initial structure of the protein-ligand complex In the combined state, at the 20th frame (corresponding to Figure 3 in ) to the 30th frame (corresponding to Figure 3 in ) shows gradual separation. The total number of frames in the escape trajectory of different protein-ligand complexes varies, and is related to the depth of the protein pocket and the complexity of the path. The method of this embodiment is applicable to various protein pocket depths.

[0172] Based on the above embodiments, the present invention also provides a molecular motion trajectory generating device, such as Figure 4As shown, the device includes:

[0173] Acquisition module 01, used to obtain the atomic class sequence and atomic coordinate information of the first frame of the target protein-ligand complex;

[0174] Generation module 02, configured to generate atomic coordinate information for a plurality of subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generation model; the pre-trained generation model is trained based on the escape trajectory of an existing protein-ligand complex;

[0175] The summarizing module 03 is configured to generate a molecular motion trajectory of the target protein-ligand complex according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of all frames.

[0176] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be shown as follows: Figure 5 As shown. The terminal includes a processor, a memory, a network interface, and a display screen connected via a system bus. The processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a method for generating molecular motion trajectories. The display screen of the terminal can be a liquid crystal display or an electronic ink display.

[0177] Those skilled in the art will understand that Figure 5 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0178] In one implementation, the terminal has one or more programs stored in its memory, and is configured to be executed by one or more processors. The one or more programs include instructions for performing a method for generating molecular motion trajectories.

[0179] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0180] In summary, the present invention discloses a method, device, terminal, and storage medium for generating molecular motion trajectories, relating to the field of drug discovery technology. The method comprises: obtaining an atomic class sequence and atomic coordinate information for a first frame of a target protein-ligand complex; generating atomic coordinate information for several subsequent frames frame by frame based on the atomic class sequence and atomic coordinate information for the first frame of the target protein-ligand complex using a pretrained generative model; the pretrained generative model is trained based on existing escape trajectories of protein-ligand complexes; and generating the molecular motion trajectory of the target protein-ligand complex based on the atomic class sequence and atomic coordinate information for all frames of the target protein-ligand complex. The present invention employs a generative model to generate coordinates frame by frame to simulate molecular motion trajectories. During the training phase, the pretrained generative model learns the inter-frame coordinate evolution patterns during ligand escape from a protein pocket using a dataset containing existing escape trajectories. Therefore, during the inference phase, only the atomic class sequence and the atomic coordinate information for the initial state of the first frame need be input to quickly and accurately iteratively generate atomic coordinate information for subsequent frames, significantly reducing computational costs and improving simulation efficiency and prediction accuracy.

[0181] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for generating molecular motion trajectories, characterized in that: The method comprises: Obtain the atomic class sequence and atomic coordinate information of the first frame of the target protein-ligand complex; The pre-trained generative model is used to generate atomic coordinate information of several subsequent frames frame by frame based on the atomic category sequence of the target protein-ligand complex and the atomic coordinate information of the first frame, including: using the pre-trained generative model and a numerical integration solution method, based on the atomic category sequence, the atomic coordinate information of the first frame and time information, the atomic coordinate information of the intermediate state is gradually solved based on time increments; the time increment is obtained based on the frame spacing; when the cumulative time increment reaches the frame spacing, the atomic coordinate information of the second frame is determined based on the atomic coordinate information of the intermediate state currently solved; the pre-trained generative model is trained based on the escape trajectory of the existing protein-ligand complex; The molecular motion trajectory of the target protein-ligand complex is generated according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of all frames.

2. The method for generating molecular motion trajectories according to claim 1, wherein: Existing methods for obtaining escape trajectories of protein-ligand complexes include: Obtaining the original escape trajectory of an existing protein-ligand complex from an initial state to a separated state from a public data set, and normalizing each frame of information in the original escape trajectory to obtain a standardized file; From the standardized files of all frames, the atomic class sequence of the existing protein-ligand complex and the atomic coordinate information of all frames are extracted to compose the escape trajectory of the existing protein-ligand complex.

3. The method for generating molecular motion trajectories according to claim 1, wherein: The training method of the pre-trained generative model includes: Use the flow model to build the initial generation model between two adjacent frames; The initial generative model is trained using the escape trajectory of the existing protein-ligand complex as training data to fit the vector field of the ordinary differential equation (ODE) to obtain the pre-trained generative model.

4. The method for generating molecular motion trajectories according to claim 3, wherein: The initial generative model is trained using the escape trajectories of existing protein-ligand complexes as training data to fit the vector field of ordinary differential equations (ODEs), including: From the escape trajectory of the existing protein-ligand complex, the atomic coordinate information of the adjacent first training frame and the second training frame is sampled and interpolated to obtain the atomic coordinate information of the training intermediate state; Performing a differential operation on the atomic coordinate information of the training intermediate state to obtain a first vector field; Predicting a second vector field based on the atomic coordinate information of the training intermediate state, the atomic category sequence of the escape trajectory, and time information through an initial generative model; The initial generation model is trained according to the first vector field and the second vector field to fit the vector field.

5. The method for generating molecular motion trajectories according to claim 3, wherein: The initial generative model is trained using the escape trajectories of existing protein-ligand complexes as training data to fit the vector field of ordinary differential equations (ODEs), including: From the escape trajectory of the existing protein-ligand complex, the atomic coordinate information of the adjacent first training frame and the second training frame is sampled and interpolated to obtain the atomic coordinate information of the training intermediate state; Performing a differential operation on the atomic coordinate information of the training intermediate state to obtain a first vector field; When the first training frame is not the first frame, extracting training history information based on atomic coordinate information of several historical frames before the first training frame; Predicting a second vector field based on the atomic coordinate information of the training intermediate state, the atomic category sequence of the escape trajectory, the training history information, and the time information through an initial generation model; The initial generation model is trained according to the first vector field and the second vector field to fit the vector field.

6. The method for generating molecular motion trajectories according to claim 4 or claim 5, wherein: Training the initial generative model according to the first vector field and the second vector field to fit the vector field includes: When training the initial generation model, using the first vector field as a supervisory signal to determine whether the motion trend reflected by the second vector field predicted by the initial generation model is accurate; The gap between the first vector field and the second vector field is converged to fit the vector field.

7. The method for generating molecular motion trajectories according to claim 6, wherein: Converging the gap between the first vector field and the second vector field to fit the vector field includes: Calculating a model loss value according to the first vector field and the second vector field; Determine whether the training termination condition is currently met. If not, backpropagate the initial generation model according to the model loss value to update the network parameters; The step of sampling atomic coordinate information of adjacent first training frames and second training frames from the escape trajectory of the existing protein-ligand complex and performing interpolation processing is continued until the training termination condition is reached, and the current initial generation model is used as the pre-trained generation model.

8. The method for generating molecular motion trajectories according to claim 1, wherein: Generate atomic coordinate information of several subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame through a pre-trained generative model, including: Generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and time information using the pre-trained generative model; Determine whether the iteration termination condition is currently satisfied, and if not, use the atomic coordinate information of the second frame as the atomic coordinate information of the first frame; Continue to execute the step of generating the atomic coordinate information of the second frame according to the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, and the time information through the pre-trained generation model until the iteration termination condition is met.

9. The method for generating molecular motion trajectories according to claim 1, wherein: Generate atomic coordinate information of several subsequent frames frame by frame based on the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of the first frame through a pre-trained generative model, including: When the first frame is not the first frame, extracting historical information based on atomic coordinate information of several historical frames before the first frame; Generate atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information, and the time information through the pre-trained generative model; Determine whether the iteration termination condition is currently satisfied, and if not, use the atomic coordinate information of the second frame as the atomic coordinate information of the first frame; Continue to execute the step of generating the atomic coordinate information of the second frame according to the atomic category sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information and the time information through the pre-trained generation model until the iteration termination condition is met.

10. The method for generating molecular motion trajectories according to claim 9, wherein: Generating atomic coordinate information of a second frame according to the atomic class sequence of the target protein-ligand complex, the atomic coordinate information of the first frame, the historical information, and the time information through the pre-trained generation model, including: By using the pre-trained generative model and numerical integration solution, based on the atomic category sequence, the atomic coordinate information of the first frame, the historical information, and the time information, the atomic coordinate information of the intermediate state is gradually solved based on time increments; the time increments are obtained based on the frame spacing. When the accumulated time increment reaches the frame interval, the atomic coordinate information of the second frame is determined according to the atomic coordinate information of the intermediate state currently solved.

11. The method for generating molecular motion trajectories according to claim 8 or claim 9, wherein: When the task of the pre-trained generative model is to generate the escape trajectory of the target protein-ligand complex, the iteration termination condition is: the distance value between the protein surface and the ligand calculated based on the atomic coordinate information of the second frame is less than a preset value.

12. A molecular motion trajectory generating device, characterized in that: The device comprises: An acquisition module is used to obtain the atomic class sequence and atomic coordinate information of the first frame of the target protein-ligand complex; A generation module is configured to generate, frame by frame, atomic coordinate information of several subsequent frames based on the atomic category sequence of the target protein-ligand complex and the atomic coordinate information of the first frame using a pre-trained generation model, including: using the pre-trained generation model and a numerical integration solution method, gradually solving the atomic coordinate information of the intermediate state based on time increments according to the atomic category sequence, the atomic coordinate information of the first frame, and time information; the time increment is obtained based on the frame spacing; when the cumulative time increment reaches the frame spacing, determining the atomic coordinate information of the second frame based on the atomic coordinate information of the intermediate state currently solved; the pre-trained generation model is trained based on the escape trajectory of an existing protein-ligand complex; The summarizing module is used to generate the molecular motion trajectory of the target protein-ligand complex according to the atomic class sequence of the target protein-ligand complex and the atomic coordinate information of all frames.

13. A terminal, characterized in that: The terminal includes a memory and one or more processors; the memory stores one or more programs; the programs include instructions for executing the molecular motion trajectory generation method according to any one of claims 1 to 11; and the processor is used to execute the programs.

14. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that: The instructions are suitable for being loaded and executed by a processor to implement the steps of the method for generating molecular motion trajectories as claimed in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Molecular dynamics simulation acceleration method, electronic equipment and storage medium

    CN117116368A

  • Molecular generation method and device

    CN118280477A