Molecular processing method, device, electronic device, storage medium, and program product
Patent Information
- Application Number
- CN202211165022.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-09-23
AI Technical Summary
相关技术中基于分子动力学的分子状态模拟需要高性能计算机来辅助算法的加速,随着科学技术的发展,分子动力学应用日益广泛,基于分子动力学的分子状态模拟的需求越来越多,但是相关技术中针对大分子直接进行分子状态模拟的计算耗时非常长,而且目标分子不处于理想的满足拉格朗日力学的状态,因此在模拟精度上存在损失
通过对目标分子的第一状态随机分布进行对应第一时刻和第二时刻的状态采样处理,得到目标分子在第一时刻的第一状态以及目标分子在第二时刻的第二状态,相当于进行了粗粒度采样,从而再从由第一时刻以及第二时刻构成的第一时间段中确定出多个第三时刻,相当于对时间进行更加细粒度的划分,并对第一状态以及第二状态分别进行编码处理,得到第一状态的第一编码以及第二状态的第二编码,相当于将分子状态从状态空间映射到编码隐空间,根据第一编码、第二编码以及第三时刻的时刻标识,生成对应第三时刻的第二状态随机分布,从而引入第二状态随机分布的随机过程来对细粒度的分子状态进行约束,后续对第二状态随机分布进行随机状态采样处理,得到第三时刻的第三编码,进而对第三编码进行解码处理,得到目标分子在第三时刻的第三状态,将分子状态从编码隐空间还原到状态空间,从而提高分子状态的计算速度以及计算精度。
Smart Images

Figure CN116992748B_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and more particularly to a molecular processing method, apparatus, electronic device, computer-readable storage medium, and computer program product based on artificial intelligence. Background Technology
[0002] Artificial Intelligence (AI) is a comprehensive technology within computer science that studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. AI technology is a multidisciplinary field, encompassing a wide range of areas, including natural language processing and machine learning / deep learning. With technological advancements, AI will be applied in more fields and play an increasingly important role.
[0003] Molecular state simulation based on molecular dynamics is of paramount importance to computational chemistry and pharmaceutical fields. Related technologies require high-performance computers to accelerate algorithms. With the development of science and technology, the applications of molecular dynamics are becoming increasingly widespread, leading to a growing demand for molecular state simulation based on molecular dynamics. However, direct molecular state simulation of macromolecules is computationally very time-consuming, and the target molecules are not in ideal Lagrangian states, resulting in a loss of simulation accuracy. Summary of the Invention
[0004] This application provides an artificial intelligence-based molecular processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can simultaneously improve the calculation accuracy and speed of molecular states.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides an artificial intelligence-based molecular processing method, including: The first state of the target molecule is randomly distributed and sampled at the corresponding first and second time points to obtain the first state of the target molecule at the first time point and the second state of the target molecule at the second time point. Multiple third moments are determined from the first time period consisting of the first moment and the second moment, and the first state and the second state are encoded respectively to obtain the first code of the first state and the second code of the second state. Based on the first code, the second code, and the time identifier of each third time moment, a second state random distribution corresponding to each third time moment is generated, and random state sampling processing is performed on each second state random distribution to obtain a third code for each third time moment, wherein the time identifier represents the temporal order of the plurality of third times moments; For each of the third moments, the following processing is performed: the third encoding of the third moment and the third encoding of the historical third moment are decoded to obtain the third state of the target molecule at the third moment, wherein the historical third moment is the third moment among the plurality of third moments that was decoded before the third moment.
[0006] This application provides an artificial intelligence-based molecular processing device, comprising: The first sampling module is used to perform state sampling processing on the random distribution of the first state of the target molecule at the corresponding first time and second time, so as to obtain the first state of the target molecule at the first time and the second state of the target molecule at the second time. The encoding module is used to determine a plurality of third moments from a first time period consisting of the first moment and the second moment, and to encode the first state and the second state respectively to obtain a first code for the first state and a second code for the second state. The second sampling module is used to generate a second state random distribution corresponding to each third time moment based on the first code, the second code, and the time identifier of each third time moment, and to perform random state sampling processing on each second state random distribution to obtain a third code for each third time moment, wherein the time identifier represents the temporal order of the plurality of third times moments; The decoding module is configured to perform the following processing for each of the third moments: decode the third encoding of the third moment and the third encoding of the historical third moment to obtain the third state of the target molecule at the third moment, wherein the historical third moment is the third moment among the plurality of third moments that was decoded before the third moment.
[0007] In the above scheme, before performing state sampling processing on the first state random distribution of the target molecule at corresponding first and second time moments to obtain the first state of the target molecule at the first time moment and the second state of the target molecule at the second time moment, the first sampling module is further configured to: obtain the first sample state of the target molecule at multiple first sample moments, wherein the time interval between two adjacent first sample moments is the same as the length of the first time period; and perform distribution fitting processing based on the first sample states at the multiple first sample moments to obtain the random distribution of the first state of the target molecule.
[0008] In the above scheme, the encoding module is further configured to: perform any one of the following processes: equally divide the first time period based on the target time interval to obtain multiple division points, and take the time corresponding to the multiple division points as the third time; randomly divide the first time period to obtain multiple division points, and take the time corresponding to the multiple division points as the third time.
[0009] In the above scheme, the encoding process is implemented through an encoding network, which includes N cascaded encoding layers. The encoding module is further configured to: acquire first state data of the first state, wherein the first state data includes bond length data and bond angle data of the target molecule at the first moment; perform feature extraction processing on the first input of the nth encoding layer in the N cascaded encoding layers to obtain the nth feature extraction result at the first moment, and transmit the nth feature extraction result at the first moment to the (n+1)th encoding layer for further feature extraction processing; wherein the value of N satisfies 2≤N, n is an integer starting from 1 and increasing, and the value of n satisfies 1≤n≤N-1; when n is 1, the first input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the first moment, and when n is N-1, the first input of the (n+1)th encoding layer is the first... The (n+1)th feature extraction result at time n is the first code of the first state; the second state data of the second state is obtained, wherein the second state data includes the bond length data and bond angle data of the target molecule at the second time; through the nth encoding layer in N cascaded encoding layers, feature extraction processing is performed on the second input of the nth encoding layer to obtain the nth feature extraction result at the second time, and the nth feature extraction result at the second time is transmitted to the (n+1)th encoding layer for further feature extraction processing; wherein the value of N satisfies 2≤N, n is an integer starting from 1 and increasing, and the value of n satisfies 1≤n≤N-1; when n is 1, the second input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the second time, and when n is N-1, the (n+1)th feature extraction result at the second time output by the (n+1)th encoding layer is the second code of the second state.
[0010] In the above scheme, the second sampling module is further configured to: obtain the length of the first time period; perform the following processing for each third time moment: determine a target mean based on the length of the first time period, the time identifier of the third time moment, the first code, and the second code; determine a target variance based on the length of the first time period and the time identifier of the third time moment; and determine the random distribution of the target mean and the target variance constraints as a second state random distribution corresponding to the third time moment.
[0011] In the above scheme, the second sampling module is further configured to: determine a first ratio between the length of the first time period and the time identifier of the third time period; obtain a first difference between the value and the first ratio; multiply the first difference with the first code to obtain a first multiplication result; multiply the first ratio with the second code to obtain a second multiplication result; and determine the sum of the first multiplication result and the second multiplication result as the target mean.
[0012] In the above scheme, the second sampling module is further configured to: determine a second difference between the time marker of the third time moment and the length of the first time period; multiply the time marker of the third time moment and the second difference to obtain a third multiplication result; and determine the second ratio between the third multiplication result and the length of the first time period as the target variance.
[0013] In the above scheme, the decoding process is implemented through a decoding network, which includes a mask self-attention layer and a forward propagation layer. The decoding module is further configured to: perform mask-based self-attention processing on the third encoding at the third time step and the third encoding at the historical third time step through the mask self-attention layer to obtain a mask-based self-attention processing result; and perform multilayer sensing processing on the mask-based self-attention processing result through the forward propagation layer to obtain the third state of the target molecule at the third time step.
[0014] In the above scheme, the decoding module is further configured to: perform bond length mapping processing on the mask-based self-attention processing result to obtain multiple bond length distributions and distribution weights for each bond length distribution; perform weighted processing on the multiple bond length distributions based on the distribution weights of each bond length distribution to obtain a synthetic bond length distribution; perform bond angle mapping processing on the mask-based self-attention processing result to obtain multiple bond angle distributions and distribution weights for each bond angle distribution; perform weighted processing on the multiple bond angle distributions based on the distribution weights of each bond angle distribution to obtain a synthetic bond angle distribution; perform bond length sampling processing on the synthetic bond length distribution to obtain the bond length of the target molecule at the third time step; and perform bond angle sampling processing on the synthetic bond angle distribution to obtain the bond angle of the target molecule at the third time step.
[0015] In the above scheme, before encoding the first state and the second state respectively to obtain the first code of the first state and the second code of the second state, the encoding module is further configured to: obtain positive sample combinations of sample molecules and at least one negative sample combination corresponding to the positive sample combination; wherein, the positive sample combination includes a first sample time, a second sample time, and a third sample time, the third sample time being randomly determined from a first sample time period consisting of the first sample time and the second sample time; the negative sample combination includes the first sample time, the second sample time, and a fourth sample time, the fourth sample time being within the second sample time period; and determine, through the encoding network, a first prediction code corresponding to the first sample time, a second prediction code corresponding to the second sample time, a third prediction code corresponding to the third sample time, and a fourth prediction code corresponding to the fourth sample time. Encoding; based on the first sample time, the second sample time, and the third sample time, a first predicted code corresponding to the first sample time, a second predicted code corresponding to the second sample time, and a third predicted code corresponding to the third sample time are used to determine a first distance corresponding to the positive sample combination; for each negative sample combination, based on the first sample time, the second sample time, and the fourth sample time of the negative sample combination, a second distance corresponding to the first predicted code corresponding to the first sample time, a second predicted code corresponding to the second sample time, and a fourth predicted code corresponding to the fourth sample time are used to determine a negative sample combination; based on the first distance corresponding to the positive sample combination and the second distance corresponding to each negative sample combination, the loss corresponding to the encoding network is determined; the parameter change value of the encoding network when the loss reaches its minimum value is determined, and the parameters of the encoding network are updated based on the parameter change value.
[0016] In the above scheme, the encoding module is further configured to: determine a first sample ratio of the length of the third sample time period to the length of the first sample time period; determine a first sample difference between the first sample ratio and the first sample difference, and multiply the first sample difference with the first prediction code to obtain a first sample product; multiply the first sample ratio with the second prediction code to obtain a second sample product, and sum the first sample product and the second sample product to obtain a first sample sum result; and use the distance between the third prediction code and the first sample sum result as the first distance corresponding to the positive sample combination.
[0017] In the above scheme, the encoding module is further configured to: obtain a standardized distance that is positively correlated with each of the second distances; sum the multiple standardized distances to obtain a second sample summation result; and obtain the loss of the corresponding encoding network that is positively correlated with the first distance and negatively correlated with the second sample summation result.
[0018] This application provides an electronic device, including: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the artificial intelligence-based molecular processing method provided in the embodiments of this application.
[0019] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the artificial intelligence-based molecular processing method provided in this application.
[0020] The embodiments of this application have the following beneficial effects: By sampling the first state of the target molecule at the first and second time points according to the random distribution of the first state, we obtain the first state of the target molecule at the first time point and the second state of the target molecule at the second time point, which is equivalent to coarse-grained sampling. Then, we determine multiple third time points from the first time period consisting of the first and second time points, which is equivalent to fine-grained division of time. We encode the first and second states respectively to obtain the first code of the first state and the second code of the second state, which is equivalent to mapping the molecular state from the state space to the coding latent space. Based on the first code, the second code, and the time identifier of the third time point, we generate the random distribution of the second state at the third time point, thus introducing the random process of the random distribution of the second state to constrain the fine-grained molecular state. Subsequently, we perform random state sampling on the random distribution of the second state to obtain the third code of the third time point, and then decode the third code to obtain the third state of the target molecule at the third time point. This restores the molecular state from the coding latent space to the state space, thereby improving the calculation speed and accuracy of the molecular state. Attached Figure Description
[0021] Figure 1 This is a schematic diagram illustrating the evolution of molecular states in related technologies; Figure 2 This is a schematic diagram of the architecture of the artificial intelligence-based molecular processing system provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figures 4A-4CThis is a schematic flowchart of the artificial intelligence-based molecular processing method provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating the evolution of molecular states provided in the embodiments of this application; Figure 6A This is a schematic diagram of the dynamic evolution provided in the embodiments of this application; Figure 6B This is a schematic diagram of the dynamic evolution provided in the embodiments of this application; Figure 7 This is a schematic diagram of the input and output of the decoding network provided in the embodiments of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0024] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0026] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0027] 1) Molecular Dynamics (MD): This is a set of molecular simulation methods. This method mainly relies on Newtonian mechanics to simulate the motion of molecular systems. It extracts samples from systems composed of different states of molecular systems to calculate the configuration integral of the system. Based on the result of the configuration integral, it further calculates the thermodynamic quantities and other macroscopic properties of the system.
[0028] 2) Multi-scale: This refers to the use of multiple different order of magnitudes in molecular dynamics simulation tasks. For example, each 100 steps can be used as a coarse-grained simulation, and a finer-grained simulation can be performed within each 100-step interval.
[0029] 3) Contrastive Learning: This is a discriminative representation learning framework (or method) based on the idea of contrast, mainly used for unsupervised representation learning. Contrastive learning learns the feature representation of a sample by comparing it with positive and negative samples in the feature space.
[0030] 4) Internal coordinates: These are coordinates that describe the motion inside a molecular system. They represent the position of the atomic nucleus through the connection of valence bonds and bond lengths r, bond angles α, and dihedral angles θ.
[0031] 5) Kabsch algorithm: used to solve for optimal rotations, and has important applications in molecular biology, especially in comparing the similarity of proteins.
[0032] 6) Brownian motion: refers to the continuous, random motion of particles suspended in a liquid or gas.
[0033] 7) Brown's Bridge: A continuous-time stochastic process B(t) whose probability distribution follows the conditional probability distribution of a Wiener process W(t) under the condition that B(0) = B(1) = 0. Brown's Bridge is a class of theorems that discuss the probability distribution and the limiting properties of sample functions of a series of stochastic processes.
[0034] 8) Wiener process: also known as Brownian motion process, is an important independent increment process. The Wiener process is a Markov process with a mean of 0 and a variance of 1.
[0035] 9) Von Mises distribution: refers to a continuous probability distribution model on a circle, also known as the circular normal distribution.
[0036] 10) Gaussian mixture distribution: This is a mixture model that combines multiple Gaussian models. It can be understood as a linear combination of multiple Gaussian distribution functions. Theoretically, Gaussian mixture models can fit any type of distribution.
[0037] In related technologies, solvers can be used to perform numerical calculations directly; see [link to relevant documentation]. Figure 1 , Figure 1This is a schematic diagram of molecular state evolution in related technologies. For macromolecules, direct simulation is extremely time-consuming. Considering that molecular state simulation tasks exhibit statistical regularities during dynamic changes, machine learning can be used to obtain these regularities and predict the entire dynamic evolution process, thus eliminating the need for costly numerical simulation calculations.
[0038] In molecular state simulation tasks based on molecular dynamics, the molecules under study are often not in ideal states that satisfy Lagrangian mechanics. For example, the molecules may be in a solvent (such as water molecules). This means that it is necessary to consider not only the state of the target molecule but also the influence of solvent molecules on it. First, related technologies do not consider the influence of solvent molecules on the target molecule. Second, the machine learning algorithms in related technologies do not effectively constrain the neural network models used in the simulation. For example, related technologies use long short-term memory (LSTM) network models, which allow simulation of molecular states based on historical states, but lack explicit stochastic processes to constrain the network, making it difficult for the neural network model to optimize to the optimal solution. Furthermore, the training of LSM network models cannot be parallelized, making it difficult to use large-scale GPU parallel training techniques. Finally, in implementing the embodiments of this application, the applicant found that while machine learning algorithms can predict some frequently occurring states, the prediction effect for other important but less frequent states is very poor.
[0039] To address the aforementioned issues, embodiments of this application provide an artificial intelligence-based molecular processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can simultaneously improve the accuracy and speed of molecular state simulation.
[0040] The AI-based molecular processing method provided in this application can be implemented by a terminal / server alone; or it can be implemented collaboratively by a terminal and a server. For example, the terminal can independently undertake the AI-based molecular processing method described below, or the terminal can send a prediction request for the molecular state of a target molecule within a first time period to the server. The server executes the AI-based molecular processing method according to the received prediction request for the molecular state of the target molecule within the first time period. The method performs state sampling processing on the random distribution of the first state corresponding to the first time and the second time, to obtain the first state at the first time and the second state at the second time. Multiple third times are determined from the first time period consisting of the first time and the second time. The first state and the second state are encoded respectively to obtain the first code of the first state and the second code of the second state. A random distribution of the second state at the third time is generated, and random state sampling processing is performed on the random distribution of the second state to obtain the third code of the third time. The third code of the third time and the third code of the historical third time are decoded to obtain the third state at the third time.
[0041] The electronic device for molecular processing provided in this application can be various types of terminal devices or servers. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0042] Taking servers as an example, such as server clusters deployed in the cloud, AI as a Service (AIaaS) is offered to users. The AIaaS platform breaks down several common AI services and provides them as independent or packaged services in the cloud. This service model is similar to an AI-themed marketplace, where all users can access and use one or more AI services provided by the AIaaS platform through application programming interfaces.
[0043] For example, one type of artificial intelligence cloud service could be a molecular processing service, where a cloud server encapsulates the molecular processing program provided in this application embodiment. A user invokes the molecular processing service in the cloud service through a terminal (running a client, such as a compound screening client), causing the cloud-deployed server to invoke the encapsulated molecular processing program. This program performs state sampling processing on a random distribution of a first state corresponding to a first time moment and a second time moment, obtaining a first state at the first time moment and a second state at the second time moment. Multiple third times are determined from a first time period consisting of the first and second times moments, and the first and second states are encoded respectively, obtaining a first code for the first state and a second code for the second state. A random distribution of the second state at the third time moment is generated, and random state sampling processing is performed on the random distribution of the second state to obtain a third code for the third time moment. The third code for the third time moment and the third codes of historical third times moments are decoded to obtain the third state at the third time moment.
[0044] See Figure 2 , Figure 2 This is a schematic diagram of the architecture of an artificial intelligence-based molecular processing system provided in an embodiment of this application. The terminal 400 is connected to the server 200 through a network 300, which can be a wide area network, a local area network, or a combination of both.
[0045] Terminal 400 (running a client, such as a molecular analysis client) can be used to obtain a prediction request for the molecular state of a target molecule within a first time period. For example, researchers input the target molecule through the input interface of terminal 400, which automatically generates a prediction request for the molecular state of the target molecule within the first time period. Terminal 400 sends the prediction request to server 200. Server 200 performs state sampling processing on the random distribution of the first state corresponding to the first time and the second time, to obtain the first state at the first time and the second state at the second time. Multiple third times are determined from the first time period consisting of the first time and the second time. The first state and the second state are encoded respectively to obtain the first code of the first state and the second code of the second state. A random distribution of the second state at the third time is generated, and random state sampling processing is performed on the random distribution of the second state to obtain the third code of the third time. The third code of the third time and the third code of the historical third time are decoded to obtain the third state at the third time. Server 200 returns the third state of the target molecule at each third time to terminal 400.
[0046] In some embodiments, a molecular processing plugin can be embedded in the client running on the terminal to implement an AI-based molecular processing method locally on the client. For example, after the terminal 400 receives a prediction request for the molecular state of a target molecule within a first time period, it calls the molecular processing plugin to implement an AI-based molecular processing method. This involves sampling the random distribution of the first state at corresponding first and second time moments to obtain the first state at the first time moment and the second state at the second time moment; determining multiple third time moments from the first time period consisting of the first and second time moments; encoding the first and second states respectively to obtain the first code for the first state and the second code for the second state; generating a random distribution of the second state at the third time moment; sampling the random distribution of the second state to obtain the third code for the third time moment; and decoding the third code for the third time moment and the third codes of historical third time moments to obtain the third state at the third time moment.
[0047] In some embodiments, after the terminal 400 obtains a prediction request for the molecular state of the target molecule within a first time period, it calls the molecular processing interface of the server 200 (which can be provided as a cloud service, i.e., a molecular processing service). The server 200 performs state sampling processing on the random distribution of the first state corresponding to the first time and the second time, to obtain the first state at the first time and the second state at the second time. Multiple third times are determined from the first time period consisting of the first time and the second time, and the first state and the second state are encoded respectively to obtain the first code of the first state and the second code of the second state. A random distribution of the second state at the third time is generated, and random state sampling processing is performed on the random distribution of the second state to obtain the third code of the third time. The third code of the third time and the third code of the historical third time are decoded to obtain the third state at the third time.
[0048] Researchers can use the predicted third state at each third time step to verify the discovery of drug synthesis related to the target molecule, thereby improving the research efficiency of drug synthesis pathways; or, researchers can use the predicted third state at each third time step to reveal the scientific laws related to the target molecule.
[0049] See Figure 3 , Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 3The terminal 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 3 The general labeled all buses as Bus System 440.
[0050] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0051] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0052] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0053] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0054] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0055] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with user interface 430 (e.g., a display screen, a speaker, etc.). The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0056] In some embodiments, the artificial intelligence-based molecular processing device provided in this application can be implemented in software. Figure 3 An AI-based molecular processing device 455, stored in memory 450, is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a first sampling module 4551, an encoding module 4552, a second sampling module 4553, and a decoding module 4554. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0057] As previously stated, the AI-based molecular processing method provided in this application can be implemented by various types of electronic devices. See also Figure 4A , Figure 4A This is a flowchart illustrating the artificial intelligence-based molecular processing method provided in the embodiments of this application, combined with... Figure 4A Steps 101 to 104 are described below.
[0058] In step 101, the first state of the target molecule is randomly distributed and sampled at corresponding first and second time points to obtain the first state of the target molecule at the first time point and the second state of the target molecule at the second time point.
[0059] In the embodiments of this application, both the first state and the second state refer to molecular states. In molecular state simulation tasks based on molecular dynamics, the target molecule is not in an ideal state that satisfies Lagrangian mechanics. For example, the target molecule may be in a solvent. In such a system, it is necessary to consider not only the state of the target molecule but also the influence of solvent molecules on the target molecule. These influences cause the state of the target molecule to be non-fixed at each moment, but to satisfy a stochastic process. That is, the target molecule exists in various possible states with a certain probability at each moment. This stochastic process is the random distribution of the first state.
[0060] In some embodiments, before sampling the first state of the target molecule at corresponding first and second time moments to obtain the first state of the target molecule at the first time moment and the second state of the target molecule at the second time moment, the first sample states of the target molecule at multiple first sample moments are obtained, wherein the time interval between two adjacent first sample moments is the same as the length of the first time period; distribution fitting is performed based on the first sample states at multiple first sample moments to obtain the random distribution of the first state of the target molecule. By determining the first random distribution of the first state that the coarse-grained time series follows through the embodiments of this application, the accuracy of subsequent fine-grained fitting can be improved.
[0061] As an example, the first-state random distribution is obtained by fitting the molecular states (first sample states) at multiple first sample times. The molecular states at multiple first sample times are simulated through original numerical calculations. The time interval between each first sample time is the same as the time interval between the first and second times. That is, the length of the time interval between two adjacent first sample times is the same as the length of the first time interval. Typically, the length of the first time interval is 100 steps, and each step is 10 femtoseconds. Therefore, the length of the first time interval is 1000 femtoseconds. This means that the simulated first-state random distribution can be used to describe the molecular states of the target molecule in a time series with an interval of 1000 femtoseconds. The length of the first time interval is greater than the set number of steps, thus obtaining the state distribution that the coarse-grained time series follows.
[0062] As an example, the first time step can be less than the second time step or greater than the second time step. For instance, when the first time step is less than the second time step, the molecular state at time step 0 is... The molecular state at the second time point T is ,by and Fine-grained simulations are performed, serving as the starting and ending points (steps 102 to 104). The process is then repeated according to user requirements; if there is a need to continue simulation, new samples are taken from the first-state random distribution followed by the coarse-grained time sequence. The last sample As a new And continue to target the new And new Perform fine-grained simulations and repeat this process until the user's needs are met.
[0063] In step 102, multiple third moments are determined from the first time period consisting of the first moment and the second moment, and the first state and the second state are encoded respectively to obtain the first code of the first state and the second code of the second state.
[0064] In some embodiments, determining multiple third moments from the first time period consisting of the first moment and the second moment in step 102 can be achieved through the following technical solutions: performing any of the following processes: equally dividing the first time period based on the target time interval to obtain multiple segmentation points, and using the moments corresponding to the multiple segmentation points as the third moments; or randomly dividing the first time period to obtain multiple segmentation points, and using the moments corresponding to the multiple segmentation points as the third moments. By cleverly combining multi-scale phenomena in molecular dynamics through the embodiments of this application, accurate predictions can be maintained in fine-grained simulations, and the model can be made not limited to certain molecular states.
[0065] As an example, the molecular state at the first time step and the molecular state at the second time step both follow a large-scale dynamic random process (the first state is randomly distributed). The change of the molecular state between the first time step and the second time step follows a finer-grained dynamic process. That is, the process from the first time step to the third time step A adjacent to the first time step is a fine-grained dynamic process, the process from the third time step A to the third time step B adjacent to the third time step A is a fine-grained dynamic process, and the process from the third time step D to the second time step is a fine-grained dynamic process, where the third time step D is the third time step adjacent to the second time step.
[0066] In some embodiments, the mapping process is implemented through an encoding network. In step 102, the first state and the second state are encoded to obtain the first encoding of the first state and the second encoding of the second state. This can be achieved by the following technical solution: the first state is mapped from a high-dimensional sequence data space to a low-dimensional latent space to obtain the first encoding of the first state; the second state is mapped from a high-dimensional sequence data space to a low-dimensional latent space to obtain the second encoding of the second state.
[0067] As an example, the encoding network can be a graph neural network, which includes N cascaded encoding layers. The above-mentioned mapping process from the high-dimensional sequence data space to the low-dimensional latent space to obtain the first encoding of the first state can be implemented through the following technical solution: obtaining the first state data of the first state, wherein the first state data includes the bond length data and bond angle data of the target molecule at the first time step; performing feature extraction processing on the first input of the nth encoding layer in the N cascaded encoding layers to obtain the nth feature extraction result at the first time step, and transmitting the nth feature extraction result at the first time step to the (n+1)th encoding layer for further feature extraction processing; wherein the value of N satisfies 2≤N, and n is an integer starting from 1 and increasing, and the value of n satisfies 1≤n≤N-1; when n is 1, the first input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the first time step, and when n is N-1, the output of the (n+1)th encoding layer is the (n+1)th feature extraction result at the first time step, which is the first encoding of the first state.
[0068] As an example, the encoding network can be a graph neural network, which includes N cascaded encoding layers. The above-mentioned mapping process from the high-dimensional sequence data space to the low-dimensional latent space to obtain the second encoding of the second state can be implemented through the following technical solution: obtaining the second state data of the second state, wherein the second state data includes the bond length data and bond angle data of the target molecule at the second time step; performing feature extraction processing on the second input of the nth encoding layer in the N cascaded encoding layers to obtain the nth feature extraction result at the second time step, and transmitting the nth feature extraction result at the second time step to the (n+1)th encoding layer for further feature extraction processing; wherein the value of N satisfies 2≤N, and n is an integer starting from 1 and increasing from 1, and the value of n satisfies 1≤n≤N-1; when n is 1, the second input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the second time step, and when n is N-1, the output of the (n+1)th encoding layer is the (n+1)th feature extraction result at the second time step, which is the second encoding of the second state.
[0069] In step 103, a second state random distribution corresponding to each third time moment is generated based on the first code, the second code, and the time identifier of each third time moment. Random state sampling processing is performed on each second state random distribution to obtain the third code of each third time moment. The time identifier represents the temporal order of multiple third time moments.
[0070] As an example, a time marker represents the temporal sequence of multiple third times. For instance, the first time is time zero, and the second time is the 1000th femtosecond. When the third time is the dividing point of the first time interval, one step size can be 10 femtoseconds. That is, the time interval between the first and third times, the time interval between two adjacent third times, and the time interval between the third and second times are all 10 femtoseconds. The time marker of the third time can be 1, 2, 3, ..., 99. That is, the time marker directly represents the temporal sequence of the third time within the first time interval. For instance, when the time marker of the third time is 1, it indicates that the third time is the next time adjacent to the first time. When the time marker of the third time is 99, it indicates that the third time is the previous time adjacent to the second time.
[0071] As an example, the time marker can also be the actual sampling time of the third time. The actual sampling time can also indirectly represent the time sequence of the third time within the first time period. For example, the actual sampling time of the third time A is 8 femtoseconds, the actual sampling time of the adjacent third time B is 15 femtoseconds, the actual sampling time of the adjacent third time C of the third time B is 20 femtoseconds, and so on. Then it can be clearly seen that the third time A precedes the third time B, and the third time B precedes the third time C.
[0072] In some embodiments, see Figure 4B , Figure 4B This is a flowchart illustrating the artificial intelligence-based molecular processing method provided in this application embodiment. In step 103, a second state random distribution corresponding to each third time moment is generated based on the first code, the second code, and the time identifier of each third time moment, which can be executed for each third time moment. Figure 4B Steps 1031 to 1034 shown are implemented.
[0073] In step 1031, the length of the first time period is obtained.
[0074] As an example, the first time point is time zero, the second time point is the 1000th femtosecond, and one step size is 10 femtoseconds. That is, the time interval between the first time point and the third time point, the time interval between two adjacent third time points, and the time interval between the third time point and the second time point are all 10 femtoseconds. The length of the first time interval is 1000 femtoseconds.
[0075] In step 1032, the target mean is determined based on the length of the first time period, the time identifier of the third time period, the first code, and the second code.
[0076] In some embodiments, determining the target mean in step 1032 based on the length of the first time period, the time identifier of the third time period, the first code, and the second code can be achieved through the following technical solution: determining a first ratio between the length of the first time period and the time identifier of the third time period; obtaining a first difference between the value and the first ratio; multiplying the first difference with the first code to obtain a first multiplication result; multiplying the first ratio with the second code to obtain a second multiplication result; and determining the sum of the first multiplication result and the second multiplication result as the target mean.
[0077] As an example, the time identifier of the third time moment can be 1, 2, 3, ..., 99. The time identifier represents the time sequence of the third time moment within the first time period. For example, when the time identifier of the third time moment is 1, it indicates that the third time moment is the next time moment adjacent to the first time moment. When the time identifier of the third time moment is 99, it indicates that the third time moment is the previous time moment adjacent to the second time moment. See formula (1): (1); in, Let t be the target mean, t be the time marker of the third time interval, and T be the length of the first time interval. It is the first ratio. It is the first difference. It is the first code. It is the second code. This is the result of the first multiplication. It is the result of the second multiplication.
[0078] In step 1033, the target variance is determined based on the length of the first time period and the time marker of the third time period.
[0079] The embodiments of this application can utilize random processes to constrain the encoding of the latent space, thereby enabling both the encoding and decoding networks to achieve high-accuracy encoding and decoding effects, thus accurately predicting the molecular state at the fine-grained third time step.
[0080] In some embodiments, determining the target variance in step 1033 based on the length of the first time period and the time marker of the third time period can be achieved by the following technical solution: determining a second difference between the time marker of the third time period and the length of the first time period; multiplying the time marker of the third time period and the second difference to obtain a third multiplication result; and determining the second ratio between the third multiplication result and the length of the first time period as the target variance.
[0081] For example, see formula (2): (2); in, Here, t is the objective variance, t is the time marker for the third time step, and T is the length of the first time interval. It is the second difference. It is the result of the third multiplication.
[0082] In step 1034, the random distribution of the target mean and the target variance constraint is determined as the second state random distribution corresponding to the third time step.
[0083] In step 104, the following processing is performed for each third time point: the third encoding of the third time point and the third encoding of the historical third time point are decoded to obtain the third state of the target molecule at the third time point, wherein the historical third time point is the third time point that was decoded earlier than the third time point among multiple third time points.
[0084] In some embodiments, the decoding process is implemented through a decoding network, which includes a mask self-attention layer and a forward propagation layer; see also Figure 4C , Figure 4C This is a flowchart illustrating the artificial intelligence-based molecular processing method provided in this application embodiment. In step 104, the third code at the third time step and the third code at the historical third time step are decoded to obtain the third state of the target molecule at the third time step. Figure 4C Steps 1041 to 1042 shown are implemented.
[0085] In step 1041, the third encoding at the third time step and the third encoding at the historical third time step are processed by a mask-based self-attention layer to obtain the mask-based self-attention processing result.
[0086] As an example, the Transformer model has a self-attention layer that performs self-attention processing. Masked self-attention is an improvement on self-attention processing, and its specific implementation is the same as that of masked self-attention in the GPT2 model.
[0087] In step 1042, the mask-based self-attention processing result is processed by a forward propagation layer to obtain the third state of the target molecule at the third time step.
[0088] In some embodiments, the multilayer sensing processing of the mask-based self-attention processing result in step 1042 to obtain the third state of the target molecule at the third time step can be achieved through the following technical solution: performing bond length mapping processing on the mask-based self-attention processing result to obtain multiple bond length distributions and distribution weights for each bond length distribution; performing weighted processing on the multiple bond length distributions based on the distribution weights of each bond length distribution to obtain a synthetic bond length distribution; performing bond angle mapping processing on the mask-based self-attention processing result to obtain multiple bond angle distributions and distribution weights for each bond angle distribution; performing weighted processing on the multiple bond angle distributions based on the distribution weights of each bond angle distribution to obtain a synthetic bond angle distribution; performing bond length sampling processing on the synthetic bond length distribution to obtain the bond length of the target molecule at the third time step, and performing bond angle sampling processing on the synthetic bond angle distribution to obtain the bond angle of the target molecule at the third time step; combining the bond length and bond angle of the target molecule at the third time step to form the third state of the target molecule at the third time step. The embodiments of this application can map molecular states from the encoded latent space to the source space, thereby accurately recovering the molecular states in the source space.
[0089] For example, see formula (3): (3); in, It is the synthetic bond length distribution. It is the k-th bond length distribution. It is the distribution weight of the k-th bond length distribution. It is the number of bond length distributions. It is the mean of the k-th bond length distribution. It is the variance of the k-th bond length distribution.
[0090] For example, see formula (4): (4); in, It is a synthetic bond angle distribution. It is the distribution weight of the k-th bond angle distribution. It is the number of bond angles. It is the scaling factor for the k-th bond angle distribution. It is the phase of the cosine function of the k-th bond angle distribution.
[0091] In some embodiments, before encoding the first state and the second state to obtain the first code for the first state and the second code for the second state, positive sample combinations and at least one negative sample combination corresponding to the positive sample combination are obtained; wherein, the positive sample combination includes a first sample time, a second sample time, and a third sample time, the third sample time being randomly determined from the first sample time period consisting of the first sample time and the second sample time; the negative sample combination includes a first sample time, a second sample time, and a fourth sample time, the fourth sample time being within the second sample time period; and a first prediction code corresponding to the first sample time, a second prediction code corresponding to the second sample time, a third prediction code corresponding to the third sample time, and a fourth prediction code corresponding to the fourth sample time are determined by an encoding network. The code is as follows: Based on the first sample time, the second sample time, and the third sample time, the first predicted code corresponding to the first sample time, the second predicted code corresponding to the second sample time, and the third predicted code corresponding to the third sample time are used to determine the first distance of the corresponding positive sample combination; for each negative sample combination, based on the first sample time, the second sample time, and the fourth sample time of the negative sample combination, the first predicted code corresponding to the first sample time, the second predicted code corresponding to the second sample time, and the fourth predicted code corresponding to the fourth sample time are used to determine the second distance of the corresponding negative sample combination; based on the first distance of the corresponding positive sample combination and the second distance of each negative sample combination, the loss of the corresponding coding network is determined; the parameter change value of the coding network when the loss reaches the minimum value is determined, and the parameters of the coding network are updated based on the parameter change value.
[0092] As an example, we first collect a positive sample combination and at least one negative sample as a sample set. For each sample molecule, we obtain the sample molecule's molecular state at multiple sample times. The sample molecule state is obtained through simulation using the original molecular dynamics method. The positive sample combination is collected as follows: a sample is randomly selected during the first sample time interval, which consists of the first sample time and the second sample time. (At the third sample time), a positive sample combination is formed. The negative sample combinations are collected as follows: sample times are randomly collected in time periods outside the first sample time period. (At the fourth sample time), a negative sample combination is formed. Multiple sample times were randomly collected during time periods outside the first sample time period. (At the fourth sample time), thus allowing for separate comparisons with... The negative sample combination is formed when the difference between the molecular state of the sample at the fourth sample time and the molecular state of the sample at the third sample time is greater than the difference threshold.
[0093] As an example, the first sample state of the sample molecule at the first sample time, the second sample state at the second sample time, the third sample state at the third sample time, and the fourth sample state at the fourth sample time are encoded respectively to obtain the first prediction code corresponding to the first sample time, the second prediction code corresponding to the second sample time, and the fourth prediction code corresponding to the fourth sample time.
[0094] As an example, the sample molecule is the same as the target molecule, or the sample molecule is a molecule whose attribute similarity to the target molecule is greater than the attribute similarity threshold.
[0095] In some embodiments, the determination of the first distance for a corresponding positive sample combination based on the first sample time, the second sample time, and the third sample time, the first predicted code corresponding to the first sample time, the second predicted code corresponding to the second sample time, and the third predicted code corresponding to the third sample time, can be achieved through the following technical solution: determining a first sample ratio between the third sample time and the total number of sample times within the first sample time period; determining a first sample difference between the first sample ratio and the first sample ratio, and multiplying the first sample difference by the first predicted code to obtain a first sample product; multiplying the first sample ratio by the second predicted code to obtain a second sample product, and summing the first sample product and the second sample product to obtain a first sample sum result; and using the distance between the third predicted code and the first sample sum result as the first distance for the corresponding positive sample combination. The embodiments of this application can improve the training effect of the encoding network and enhance its representational ability.
[0096] For example, see formula (5): = (5); in, It is the first distance corresponding to the positive sample combination. This is the first sample time. It is the third sample time. This is the second sample time. It is the processing logic of the encoding network. It is a parameter. It is the third predicted code corresponding to the third sample time. It is the first predicted code corresponding to the first sample time. This is the second predicted code corresponding to the second sample time, where t is the third sample time. It is the length of the first sample time period.
[0097] In some embodiments, determining the loss of the corresponding encoding network based on the first distance for each positive sample combination and the second distance for each negative sample combination can be achieved through the following technical solution: obtaining the standardized distance positively correlated with each second distance; summing the multiple standardized distances to obtain the sum of the second samples; and obtaining the loss of the corresponding encoding network that is positively correlated with the first distance and negatively correlated with the sum of the second samples. The embodiments of this application can improve the training effect of the encoding network and enhance its representational ability.
[0098] For example, see formula (6): (6); in, It is the first distance corresponding to the positive sample combination. It is the second distance corresponding to the combination of negative samples. It is the standardized distance that is positively correlated with the second distance. It is the result of adding the second sample. This corresponds to the loss of the coding network.
[0099] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0100] A terminal (running a client, such as a molecular analysis client) can be used to obtain prediction requests for the molecular state of a target molecule within a first time period. For example, researchers input the target molecule through the terminal's input interface, which automatically generates a prediction request for the molecular state of the target molecule within the first time period. The terminal sends the prediction request to the server, which performs state sampling processing on the random distribution of the first state corresponding to the first and second time moments to obtain the first state at the first time moment and the second state at the second time moment. Multiple third time moments are determined from the first time period consisting of the first and second time moments, and the first and second states are encoded respectively to obtain the first code for the first state and the second code for the second state. A random distribution of the second state at the third time moment is generated, and random state sampling processing is performed on the random distribution of the second state to obtain the third code for the third time moment. The third code for the third time moment and the third codes of the historical third time moments are decoded to obtain the third state at the third time moment. Researchers can use the predicted third state at each third time moment to verify the discovery of the synthesis of drugs related to the target molecule, thereby improving the research efficiency of drug synthesis pathways; or, researchers can use the predicted third state at each third time moment to reveal scientific laws related to the target molecule.
[0101] Molecular state simulation based on molecular dynamics is of paramount importance in computational chemistry and pharmaceutical fields. Related technologies require high-performance computing to accelerate algorithms, and with the advancement of science and technology and the widespread application of molecular dynamics, the demand for molecular state simulation based on molecular dynamics is increasing. Therefore, using deep learning techniques to directly learn the underlying molecular state change patterns based on molecular dynamics data in large databases is particularly important. Specifically, molecular state simulation based on molecular dynamics can serve as an effective validation tool for drug discovery, improving the research efficiency of new drug synthesis pathways; it can reveal scientific laws and provide new scientific knowledge; and it can provide more accurate molecular state predictions than human experts. Direct scientific calculations require many human assumptions and a large amount of prior knowledge; applying machine learning to solve this problem greatly reduces the amount of prior knowledge required, significantly improving the efficiency of new drug development.
[0102] In some embodiments, the molecular state simulation task provided in this application more realistically considers the dynamics of molecule-solvent binding. Unlike molecular state simulations based on molecular dynamics in a vacuum, the binding process with a solvent involves a non-conservative system (not following Hamiltonian mechanics), meaning that at each moment, the molecular state is a random variable following a certain random distribution. In this application, the molecular state is represented by its internal coordinates. Specifically, for each atom in the molecule, it is represented in Euclidean space as three-dimensional coordinates. To consider invariance, this application uses internal coordinates to standardize the molecular state; that is, the molecule's state x at each moment is determined by bond lengths and bond angles. The conversion between three-dimensional coordinates and internal coordinates can be directly calculated using Kabsch transformation. That is, for each moment, the molecular state is composed of the combination of bond lengths and bond angles between individual atoms, and the specific combination occurs with a certain probability.
[0103] See Figure 1 In related technologies, molecular dynamics-based molecular state simulation directly performs numerical calculations by assuming the force fields in which they exist, that is, simulating each evolutionary process. Because there is an integral relationship between the simulation of the force field and the simulation of the state, the computational cost of direct numerical calculation is extremely high. See also Figure 5 , Figure 5This is a schematic diagram illustrating the evolution of molecular states provided in an embodiment of this application. In this embodiment, the molecular state is first mapped to a low-dimensional latent space, where a dynamic process is explicitly assumed. Finally, the next state of the molecule is predicted in the latent space and then mapped back to the original high-dimensional space. Specifically, during the dynamic evolution in the latent space, multi-scale phenomena common in molecular dynamics can be utilized to decompose the dynamic evolution into two steps. See also... Figure 6A , Figure 6A This is a schematic diagram of dynamic evolution provided in an embodiment of this application. It samples from the molecular state distribution that follows a coarse-grained temporal sequence to obtain the molecular states corresponding to two times 0 and T, respectively. and Coarse-grained time series are sequences with relatively long time steps (steps larger than a set threshold). For example, a sequence with a step size of 100 represents a sequence where two adjacent time points are 100 steps apart, where the step size can be 10 femtoseconds. Then, see... Figure 6B , Figure 6B This is a schematic diagram of dynamic evolution provided in the embodiments of this application, to and Fine-grained simulations are performed, using these as the starting and ending points. The process is then repeated according to user requirements; if further simulation is needed, new samples are taken from the molecular state distribution followed by the coarse-grained time sequence. The last sample As this step And continue to target the new And new Perform fine-grained simulations and repeat this process until the user's needs are met.
[0104] In some embodiments, Figure 5 The process of mapping molecular states to a low-dimensional latent space is achieved through an encoder network. The encoder provides a nonlinear mapping from the original input space (high-dimensional space) to the latent space. The training objective of the Encoder is to obtain a representation that maps high-dimensional sequence data to a low-dimensional latent space, following a certain stochastic process. In molecular dynamics-based molecular state simulation tasks considering solvents, the stochastic process is Brownian motion, which is mathematically characterized by the Wiener process. To more accurately simulate long-term molecular state changes, this embodiment uses a Brownian bridge process for modeling. (Corresponding to any starting point) and the end point The probability density of the Brownian bridge process is given by formula (7): (7); Where N(x, y) refers to a Gaussian distribution with mean x and variance y. It is the mean. It is variance. It is the state distribution at time t between 0 and T.
[0105] Using the Brownian bridge process instead of the Brownian process allows for the design of multi-scale models, i.e., starting points. and the end point Following large-scale dynamic processes, in and They follow a more fine-grained dynamic process. One step to It's large-scale. , , ..., It is a more granular dynamic process.
[0106] For the training set We hope they satisfy a Brownian bridge process in the latent space. This can be implemented using contrastive learning, where the positive sample combination is defined as: in A sample is randomly collected from the middle. To form a positive sample combination The definition of a negative sample combination is: A sample is randomly collected during a time period other than the specified time. Multiple negative sample combinations The loss function for contrastive learning is defined as shown in formulas (8) and (9): (8); = (9); In some embodiments, Figure 5 The process of predicting the next state of a molecule in the latent space is achieved through a decoding network. Molecular dynamics-based decoding networks require training on relatively long sequences; therefore, the GPT2 model can be used as the decoding network, allowing for parallel training of the decoding process. The GPT2 model includes a masked self-attention layer and a multilayer perceptron. In the masked self-attention layer, attention calculations only involve tokens preceding the current token, not tokens following it. The GPT2 model consists of 12 GPT2 layers, each with a dimension of 768. See also... Figure 7 , Figure 7 This is a schematic diagram of the input and output of the decoding network provided in this application embodiment, showing the output of each token. Based on the previous token (The third code of the third historical moment) and generated by the Brownian bridge process (Third encoding) Calculated.
[0107] In some embodiments, Figure 5 The process of mapping molecular states to a low-dimensional latent space is achieved through an encoder network. Figure 5 The process of predicting the next state of a molecule in the latent space is achieved through a decoding network. The overall reasoning process is described in the following description: Repeat the following process: Sampling from the state distribution that the coarse-grained temporal sequence follows and .
[0108] repeat
[0109] and Obtained by Encoder and According to formula (7), the random distribution corresponding to t is obtained, and then the obtained random distribution is sampled to obtain... .
[0110] and Obtained by Decoder .
[0111] In some embodiments, molecular states include bonds and key angle ,key Follows a mixture Gaussian distribution, bond angle It follows a mixed von Mises distribution, see formulas (10) and (11): (10); (11); in, It is the output of the decoding network. It is a mixture Gaussian distribution. Is a key The function, where, and These are the mean and variance of a Gaussian distribution. It is the weight of the k-th Gaussian distribution. It is a mixed von Mises distribution. It is a key corner The function, It is the weight of the k-th function. It is a scaling factor. It is the phase of the cosine function.
[0112] In some embodiments, training data is obtained through numerical simulation. For example, to simulate a 100-nanosecond molecular dynamics process, numerical simulation tools can be used to simulate molecular state changes within 10 picoseconds, and then sampling is performed every 10 femtoseconds to obtain sequences with a step size of 1000. This process is repeated 10,000 times to obtain 10,000 sequences with a step size of 1000 for training. Simulations from 10 picoseconds to 10 nanoseconds serve as test data, i.e., simulations are performed using a well-trained molecular dynamics model (encoder network and decoder network). After passing the test, simulations from 10 nanoseconds to 100 nanoseconds can be performed. For multi-scale partitioning, every 100 steps can be used as a coarse-grained simulation, followed by finer-grained simulations at intervals of 100 steps.
[0113] The parallel design of the training process allows for data parallelism, as GPT2 itself can perform data parallelism during training. Therefore, sequences with different initial points can be placed on different GPUs for data parallel training.
[0114] The molecular processing method provided in this application can perform molecular state simulation based on molecular dynamics more efficiently and accurately. Using GPT2 as the decoding network allows for efficient parallel design, faster training with more data, and a more powerful expressive capability for the decoding network. This application introduces stochastic processes to explicitly model the dynamic evolution of molecular states in the latent space, enabling better and faster optimization of both the encoding and decoding models. Furthermore, this application cleverly incorporates multi-scale phenomena from molecular dynamics, maintaining accurate predictions at fine-grained simulation levels while ensuring the model is not limited to certain molecular states.
[0115] This application utilizes explicit dynamic constraints to apply different constraints to different systems. For other systems, different distribution functions can be used for constraints. The encoder is modeled as a graph neural network that can generalize to different systems, so that data from simulations of different molecules under the same force field can be used to train a larger model, making the parameters of this large model applicable to more molecules. The hyperparameters of the model can be adjusted, such as the number of layers in the decoding network, or even the decoding network can be replaced with other models, etc.
[0116] The following continues to describe the exemplary structure of the artificial intelligence-based molecular processing device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 3As shown, the software modules stored in the AI-based molecular processing device 455 in the memory 450 may include: a first sampling module 4551, used to perform state sampling processing on the random distribution of the first state of the target molecule corresponding to the first time moment and the second time moment, to obtain the first state of the target molecule at the first time moment and the second state of the target molecule at the second time moment; an encoding module 4552, used to determine multiple third times from the first time moment and the second time moment, and to encode the first state and the second state respectively, to obtain the first code of the first state and the second code of the second state; a second sampling module 4553, used to generate a random distribution of the second state corresponding to each third time moment according to the first code, the second code and the time identifier of each third time moment, and to perform random state sampling processing on each random distribution of the second state, to obtain the third code of each third time moment, wherein the time identifier represents the temporal order of multiple third times; and a decoding module 4554, used to perform the following processing for each third time moment: to decode the third code of the third time moment and the third code of the historical third time moment, to obtain the third state of the target molecule at the third time moment, wherein the historical third time moment is the third time moment that was decoded earlier than the third time moment among multiple third times.
[0117] In some embodiments, before performing state sampling processing on the random distribution of the first state of the target molecule at corresponding first and second time moments to obtain the first state of the target molecule at the first time moment and the second state of the target molecule at the second time moment, the first sampling module 4551 is further configured to: obtain the first sample state of the target molecule at multiple first sample moments, wherein the time interval between two adjacent first sample moments is the same as the length of the first time period; and perform distribution fitting processing based on the first sample states at multiple first sample moments to obtain the random distribution of the first state of the target molecule.
[0118] In some embodiments, the encoding module 4552 is further configured to: perform any one of the following processes: equally divide the first time period based on the target time interval to obtain multiple segmentation points, and use the times corresponding to the multiple segmentation points as the third time; randomly divide the first time period to obtain multiple segmentation points, and use the times corresponding to the multiple segmentation points as the third time.
[0119] In some embodiments, the encoding process is implemented through an encoding network, which includes N cascaded encoding layers. The encoding module 4552 is further configured to: acquire first state data of a first state, wherein the first state data includes bond length data and bond angle data of the target molecule at a first moment; perform feature extraction processing on the first input of the nth encoding layer through the nth encoding layer in the N cascaded encoding layers to obtain the nth feature extraction result at the first moment, and transmit the nth feature extraction result at the first moment to the (n+1)th encoding layer for further feature extraction processing; wherein the value of N satisfies 2≤N, and n is an integer starting from 1 and increasing in value, and the value of n satisfies 1≤n≤N-1; when n is 1, the first input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the first moment, and when n is N-1, the output of the (n+1)th encoding layer is... The (n+1)th feature extraction result at the first time step is the first code of the first state; the second state data of the second state is obtained, wherein the second state data includes the bond length data and bond angle data of the target molecule at the second time step; through the nth coding layer in N cascaded coding layers, the second input of the nth coding layer is processed to obtain the nth feature extraction result at the second time step, and the nth feature extraction result at the second time step is transmitted to the (n+1)th coding layer for further feature extraction processing; wherein the value of N satisfies 2≤N, and n is an integer starting from 1 and increasing, and the value of n satisfies 1≤n≤N-1; when n is 1, the second input of the nth coding layer is the bond length data and bond angle data of the target molecule at the second time step, and when n is N-1, the output of the (n+1)th coding layer is the (n+1)th feature extraction result at the second time step, which is the second code of the second state.
[0120] In some embodiments, the second sampling module 4553 is further configured to: obtain the length of the first time period; perform the following processing for each third time moment: determine the target mean based on the length of the first time period, the time identifier of the third time moment, the first code, and the second code, wherein the length of the first time period is; determine the target variance based on the length of the first time period and the time identifier of the third time moment; and determine the random distribution of the target mean and the target variance constraint as the second state random distribution corresponding to the third time moment.
[0121] In some embodiments, the second sampling module 4553 is further configured to: determine a first ratio between the length of the first time period and the time identifier of the third time period; obtain a first difference between the value and the first ratio; multiply the first difference with the first code to obtain a first multiplication result; multiply the first ratio with the second code to obtain a second multiplication result; and determine the sum of the first multiplication result and the second multiplication result as the target mean.
[0122] In some embodiments, the second sampling module 4553 is further configured to: determine a second difference between the time marker of the third time moment and the length of the first time period; multiply the time marker of the third time moment and the second difference to obtain a third multiplication result; and determine a second ratio between the third multiplication result and the length of the first time period as the target variance.
[0123] In some embodiments, the decoding process is implemented through a decoding network, which includes a mask self-attention layer and a forward propagation layer; the decoding module 4554 is further configured to: perform mask-based self-attention processing on the third encoding at the third time step and the third encoding at the historical third time step through the mask self-attention layer to obtain a mask-based self-attention processing result; and perform multilayer sensing processing on the mask-based self-attention processing result through the forward propagation layer to obtain the third state of the target molecule at the third time step.
[0124] In some embodiments, the decoding module 4554 is further configured to: perform bond length mapping processing on the mask-based self-attention processing result to obtain multiple bond length distributions and distribution weights for each bond length distribution; perform weighted processing on the multiple bond length distributions based on the distribution weights of each bond length distribution to obtain a synthetic bond length distribution; perform bond angle mapping processing on the mask-based self-attention processing result to obtain multiple bond angle distributions and distribution weights for each bond angle distribution; perform weighted processing on the multiple bond angle distributions based on the distribution weights of each bond angle distribution to obtain a synthetic bond angle distribution; perform bond length sampling processing on the synthetic bond length distribution to obtain the bond length of the target molecule at the third time step; and perform bond angle sampling processing on the synthetic bond angle distribution to obtain the bond angle of the target molecule at the third time step.
[0125] In some embodiments, before encoding the first state and the second state respectively to obtain the first code of the first state and the second code of the second state, the encoding module 4552 is further configured to: obtain positive sample combinations of sample molecules and at least one negative sample combination corresponding to the positive sample combination; wherein, the positive sample combination includes a first sample time, a second sample time, and a third sample time, the third sample time being randomly determined from a first sample time period consisting of the first sample time and the second sample time; the negative sample combination includes a first sample time, a second sample time, and a fourth sample time, the fourth sample time being within the second sample time period; and determine, through the encoding network, a first prediction code corresponding to the first sample time, a second prediction code corresponding to the second sample time, a third prediction code corresponding to the third sample time, and a code corresponding to the fourth sample time. The fourth predictive code is used; based on the first sample time, the second sample time, and the third sample time, the first predictive code corresponding to the first sample time, the second predictive code corresponding to the second sample time, and the third predictive code corresponding to the third sample time are used to determine the first distance of the corresponding positive sample combination; for each negative sample combination, based on the first sample time, the second sample time, and the fourth sample time of the negative sample combination, the first predictive code corresponding to the first sample time, the second predictive code corresponding to the second sample time, and the fourth predictive code corresponding to the fourth sample time are used to determine the second distance of the corresponding negative sample combination; based on the first distance of the corresponding positive sample combination and the second distance of each negative sample combination, the loss of the corresponding encoding network is determined; the parameter change value of the encoding network when the loss reaches the minimum value is determined, and the parameters of the encoding network are updated based on the parameter change value.
[0126] In some embodiments, the encoding module 4552 is further configured to: determine a first sample ratio between the length of the third sample time and the length of the first sample time period; determine a first sample difference between the first sample ratio and the first sample ratio, and multiply the first sample difference with the first predictive code to obtain a first sample product; multiply the first sample ratio with the second predictive code to obtain a second sample product, and sum the first sample product and the second sample product to obtain a first sample sum result; and use the distance between the third predictive code and the first sample sum result as the first distance of the corresponding positive sample combination.
[0127] In some embodiments, the encoding module 4552 is further configured to: obtain the standardized distance that is positively correlated with each second distance; sum the multiple standardized distances to obtain the sum of the second samples; and obtain the loss of the corresponding encoding network that is positively correlated with the first distance and negatively correlated with the sum of the second samples.
[0128] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. An electronic device's processor reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the artificial intelligence-based molecular processing method described above in this application.
[0129] This application provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by a processor, the processor will execute the artificial intelligence-based molecular processing method provided in this application embodiment. For example, ... Figures 4A-4C The illustrated method is a molecular processing method based on artificial intelligence.
[0130] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0131] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0132] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0133] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0134] In summary, by sampling the first state of the target molecule at the first and second time points according to the random distribution of the first state, we obtain the first state of the target molecule at the first time point and the second state at the second time point, which is equivalent to coarse-grained sampling. Then, we determine multiple third time points from the first time period consisting of the first and second time points, which is equivalent to fine-grained division of time. We encode the first and second states respectively to obtain the first code of the first state and the second code of the second state, which is equivalent to mapping the molecular state from the state space to the coding latent space. Based on the first code, the second code, and the time identifier of the third time point, we generate the random distribution of the second state at the third time point, thereby introducing the random process of the random distribution of the second state to constrain the fine-grained molecular state. Subsequently, we perform random state sampling on the random distribution of the second state to obtain the third code of the third time point, and then decode the third code to obtain the third state of the target molecule at the third time point. This restores the molecular state from the coding latent space to the state space, thereby improving the calculation speed and accuracy of the molecular state.
[0135] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A molecular processing method based on artificial intelligence, characterized in that, The method includes: The first state of the target molecule is randomly distributed and sampled at the corresponding first and second time points to obtain the first state of the target molecule at the first time point and the second state of the target molecule at the second time point. Multiple third moments are determined from the first time period consisting of the first moment and the second moment, and the first state and the second state are encoded respectively to obtain the first code of the first state and the second code of the second state. Based on the first code, the second code, and the time identifier of each third time moment, a second state random distribution corresponding to each third time moment is generated, and random state sampling processing is performed on each second state random distribution to obtain a third code for each third time moment, wherein the time identifier represents the temporal order of the plurality of third times moments; For each of the third moments, the following processing is performed: the third encoding of the third moment and the third encoding of the historical third moment are decoded to obtain the third state of the target molecule at the third moment, wherein the historical third moment is the third moment among the plurality of third moments that was decoded before the third moment.
2. The method according to claim 1, characterized in that, The determination of multiple third moments from the first time period consisting of the first moment and the second moment includes: Perform any of the following processes: The first time period is equally divided based on the target time interval to obtain multiple segmentation points, and the time corresponding to the multiple segmentation points is taken as the third time. The first time period is randomly divided to obtain multiple segmentation points, and the time corresponding to the multiple segmentation points is taken as the third time.
3. The method according to claim 1, characterized in that, The encoding process is implemented through an encoding network, which includes N cascaded encoding layers. The encoding process, which encodes the first state and the second state respectively to obtain a first code for the first state and a second code for the second state, includes: Obtain first state data of the first state, wherein the first state data includes bond length data and bond angle data of the target molecule at the first moment; The first input of the nth coding layer in the N cascaded coding layers is processed by feature extraction to obtain the nth feature extraction result at the first time step. The nth feature extraction result at the first time step is then transmitted to the (n+1)th coding layer for further feature extraction processing. Wherein, the value of N is 2≤N, and n is an integer starting from 1 and increasing in value, and the value of n is 1≤n≤N-1; when n is 1, the first input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the first time; when n is N-1, the (n+1)th feature extraction result of the (n+1)th time at the first time is the first code of the first state. Acquire second state data of the second state, wherein the second state data includes bond length data and bond angle data of the target molecule at the second time moment; The second input of the nth coding layer in the N cascaded coding layers is processed by feature extraction to obtain the nth feature extraction result at the second time step. The nth feature extraction result at the second time step is then transmitted to the (n+1)th coding layer for further feature extraction processing. Wherein, the value of N is 2≤N, and n is an integer starting from 1 and increasing, and the value of n is 1≤n≤N-1; when n is 1, the second input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the second time step; when n is N-1, the (n+1)th feature extraction result at the second time step output by the (n+1)th encoding layer is the second encoding of the second state.
4. The method according to claim 1, characterized in that, The step of generating a second state random distribution corresponding to each third time step based on the first code, the second code, and the time identifier of each third time step includes: Get the length of the first time period; For each of the aforementioned third time points, the following processing is performed: The target mean is determined based on the length of the first time period, the time identifier of the third time period, the first code, and the second code. The target variance is determined based on the length of the first time period and the time identifier of the third time period; The random distribution of the target mean and the target variance constraint is determined as the second state random distribution corresponding to the third time step.
5. The method according to claim 4, characterized in that, The step of determining the target mean based on the length of the first time period, the time identifier of the third time period, the first code, and the second code includes: Determine a first ratio between the length of the first time period and the time marker of the third time period; Obtain the first difference between the value one and the first ratio; The first difference is multiplied by the first code to obtain the first multiplication result; The first ratio is multiplied by the second code to obtain the second multiplication result; The sum of the first multiplication result and the second multiplication result is determined as the target mean.
6. The method according to claim 4, characterized in that, The determination of the target variance based on the length of the first time period and the time identifier of the third time period includes: Determine the second difference between the time marker of the third moment and the length of the first time period; The time identifier of the third time point is multiplied by the second difference to obtain the third multiplication result; The second ratio between the third multiplication result and the length of the first time period is determined as the target variance.
7. The method according to claim 1, characterized in that, The decoding process is implemented through a decoding network, which includes a mask self-attention layer and a forward propagation layer. The decoding process of the third code at the third time point and the third code at the historical third time point to obtain the third state of the target molecule at the third time point includes: The mask-based self-attention layer is used to perform mask-based self-attention processing on the third code at the third time step and the third code at the historical third time step to obtain the mask-based self-attention processing result. The forward propagation layer performs multi-layer perceptual processing on the mask-based self-attention processing result to obtain the third state of the target molecule at the third time step.
8. The method according to claim 7, characterized in that, The step of performing multi-layer perceptual processing on the mask-based self-attention processing result to obtain the third state of the target molecule at the third time step includes: The mask-based self-attention processing results are subjected to bond length mapping to obtain multiple bond length distributions and the distribution weight of each bond length distribution; The multiple bond length distributions are weighted based on the distribution weight of each bond length distribution to obtain a composite bond length distribution; The mask-based self-attention processing results are subjected to bond corner mapping processing to obtain multiple bond corner distributions and the distribution weight of each bond corner distribution; The multiple bond angle distributions are weighted based on the distribution weight of each bond angle distribution to obtain a composite bond angle distribution; Bond length sampling processing is performed on the synthetic bond length distribution to obtain the bond length of the target molecule at the third time step, and bond angle sampling processing is performed on the synthetic bond angle distribution to obtain the bond angle of the target molecule at the third time step; The bond lengths and bond angles of the target molecule at the third time point constitute the third state of the target molecule at the third time point.
9. The method according to claim 1, characterized in that, The encoding process is implemented through an encoding network. Before encoding the first state and the second state respectively to obtain the first code of the first state and the second code of the second state, the method further includes: Obtain positive sample combinations of sample molecules and at least one negative sample combination corresponding to the positive sample combination; The positive sample combination includes a first sample time, a second sample time, and a third sample time, wherein the third sample time is randomly determined from a first sample time period consisting of the first sample time and the second sample time; the negative sample combination includes the first sample time, the second sample time, and a fourth sample time, wherein the fourth sample time is within the second sample time period. The coding network determines a first prediction code corresponding to the first sample time, a second prediction code corresponding to the second sample time, a third prediction code corresponding to the third sample time, and a fourth prediction code corresponding to the fourth sample time. Based on the first sample time, the second sample time, and the third sample time, the first prediction code corresponding to the first sample time, the second prediction code corresponding to the second sample time, and the third prediction code corresponding to the third sample time, a first distance corresponding to the positive sample combination is determined; For each negative sample combination, based on the first sample time, the second sample time, and the fourth sample time of the negative sample combination, the first prediction code corresponding to the first sample time, the second prediction code corresponding to the second sample time, and the fourth prediction code corresponding to the fourth sample time, a second distance corresponding to the negative sample combination is determined. The loss of the encoding network is determined based on the first distance corresponding to the positive sample combination and the second distance corresponding to each negative sample combination. Determine the parameter change value of the encoding network when the loss reaches its minimum value, and update the parameters of the encoding network based on the parameter change value.
10. The method according to claim 9, characterized in that, The step of determining the first distance corresponding to the positive sample combination based on the first sample time, the second sample time, the third sample time, the first prediction code corresponding to the first sample time, the second prediction code corresponding to the second sample time, and the third prediction code corresponding to the third sample time includes: Determine the first sample ratio of the length of the third sample time period to the length of the first sample time period; Determine a first sample difference between a given ratio and the first sample, and multiply the first sample difference with the first prediction code to obtain a first sample product; The first sample ratio is multiplied by the second prediction code to obtain the second sample product, and the first sample product and the second sample product are summed to obtain the first sample sum result. The distance between the third predicted code and the sum of the first sample is taken as the first distance for the corresponding positive sample combination.
11. The method according to claim 9, characterized in that, The step of determining the loss of the encoding network based on the first distance corresponding to each positive sample combination and the second distance corresponding to each negative sample combination includes: Obtain the standardized distance that is positively correlated with each of the second distances; The standardized distances are summed to obtain the sum of the second samples; Obtain the loss of the corresponding encoding network that is positively correlated with the first distance and negatively correlated with the sum of the second samples.
12. A molecular processing device based on artificial intelligence, characterized in that, The device includes: The first sampling module is used to perform state sampling processing on the random distribution of the first state of the target molecule at the corresponding first time and second time, so as to obtain the first state of the target molecule at the first time and the second state of the target molecule at the second time. The encoding module is used to determine a plurality of third moments from a first time period consisting of the first moment and the second moment, and to encode the first state and the second state respectively to obtain a first code for the first state and a second code for the second state. The second sampling module is used to generate a second state random distribution corresponding to each third time moment based on the first code, the second code, and the time identifier of each third time moment, and to perform random state sampling processing on each second state random distribution to obtain a third code for each third time moment, wherein the time identifier represents the temporal order of the plurality of third times moments; The decoding module is configured to perform the following processing for each of the third moments: decode the third encoding of the third moment and the third encoding of the historical third moment to obtain the third state of the target molecule at the third moment, wherein the historical third moment is the third moment among the plurality of third moments that was decoded before the third moment.
13. The apparatus according to claim 12, characterized in that, The first sampling module is also used for: Perform any of the following processes: The first time period is equally divided based on the target time interval to obtain multiple segmentation points, and the time corresponding to the multiple segmentation points is taken as the third time. The first time period is randomly divided to obtain multiple segmentation points, and the time corresponding to the multiple segmentation points is taken as the third time.
14. The apparatus according to claim 12, characterized in that, The encoding process is implemented through an encoding network, which includes N cascaded encoding layers. The encoding module is further used for: Obtain first state data of the first state, wherein the first state data includes bond length data and bond angle data of the target molecule at the first moment; The first input of the nth coding layer in the N cascaded coding layers is processed by feature extraction to obtain the nth feature extraction result at the first time step. The nth feature extraction result at the first time step is then transmitted to the (n+1)th coding layer for further feature extraction processing. Wherein, the value of N is 2≤N, and n is an integer starting from 1 and increasing in value, and the value of n is 1≤n≤N-1; when n is 1, the first input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the first time; when n is N-1, the (n+1)th feature extraction result of the (n+1)th time at the first time is the first code of the first state. Acquire second state data of the second state, wherein the second state data includes bond length data and bond angle data of the target molecule at the second time moment; The second input of the nth coding layer in the N cascaded coding layers is processed by feature extraction to obtain the nth feature extraction result at the second time step. The nth feature extraction result at the second time step is then transmitted to the (n+1)th coding layer for further feature extraction processing. Wherein, the value of N is 2≤N, and n is an integer starting from 1 and increasing, and the value of n is 1≤n≤N-1; when n is 1, the second input of the nth encoding layer is the bond length data and bond angle data of the target molecule at the second time step; when n is N-1, the (n+1)th feature extraction result at the second time step output by the (n+1)th encoding layer is the second encoding of the second state.
15. The apparatus according to claim 12, characterized in that, The second sampling module is also used for: Get the length of the first time period; For each of the aforementioned third time points, the following processing is performed: The target mean is determined based on the length of the first time period, the time identifier of the third time period, the first code, and the second code. The target variance is determined based on the length of the first time period and the time identifier of the third time period; The random distribution of the target mean and the target variance constraint is determined as the second state random distribution corresponding to the third time step.
16. The apparatus according to claim 15, characterized in that, The second sampling module is also used for: Determine a first ratio between the length of the first time period and the time marker of the third time period; Obtain the first difference between the value one and the first ratio; The first difference is multiplied by the first code to obtain the first multiplication result; The first ratio is multiplied by the second code to obtain the second multiplication result; The sum of the first multiplication result and the second multiplication result is determined as the target mean.
17. The apparatus according to claim 15, characterized in that, The second sampling module is also used for: Determine the second difference between the time marker of the third moment and the length of the first time period; The time identifier of the third time point is multiplied by the second difference to obtain the third multiplication result; The second ratio between the third multiplication result and the length of the first time period is determined as the target variance.
18. The apparatus according to claim 12, characterized in that, The decoding process is implemented through a decoding network, which includes a mask self-attention layer and a forward propagation layer. The decoding module is also used for: The mask-based self-attention layer is used to perform mask-based self-attention processing on the third code at the third time step and the third code at the historical third time step to obtain the mask-based self-attention processing result. The forward propagation layer performs multi-layer perceptual processing on the mask-based self-attention processing result to obtain the third state of the target molecule at the third time step.
19. The apparatus according to claim 18, characterized in that, The decoding module is also used for: The mask-based self-attention processing results are subjected to bond length mapping to obtain multiple bond length distributions and the distribution weight of each bond length distribution; The multiple bond length distributions are weighted based on the distribution weight of each bond length distribution to obtain a composite bond length distribution; The mask-based self-attention processing results are subjected to bond corner mapping processing to obtain multiple bond corner distributions and the distribution weight of each bond corner distribution; The multiple bond angle distributions are weighted based on the distribution weight of each bond angle distribution to obtain a composite bond angle distribution; Bond length sampling processing is performed on the synthetic bond length distribution to obtain the bond length of the target molecule at the third time step, and bond angle sampling processing is performed on the synthetic bond angle distribution to obtain the bond angle of the target molecule at the third time step; The bond lengths and bond angles of the target molecule at the third time point constitute the third state of the target molecule at the third time point.
20. The apparatus according to claim 12, characterized in that, The encoding process is implemented through an encoding network. Before encoding the first state and the second state respectively to obtain the first code of the first state and the second code of the second state, the encoding module is further configured to: Obtain positive sample combinations of sample molecules and at least one negative sample combination corresponding to the positive sample combination; The positive sample combination includes a first sample time, a second sample time, and a third sample time, wherein the third sample time is randomly determined from a first sample time period consisting of the first sample time and the second sample time; the negative sample combination includes the first sample time, the second sample time, and a fourth sample time, wherein the fourth sample time is within the second sample time period. The coding network determines a first prediction code corresponding to the first sample time, a second prediction code corresponding to the second sample time, a third prediction code corresponding to the third sample time, and a fourth prediction code corresponding to the fourth sample time. Based on the first sample time, the second sample time, and the third sample time, the first prediction code corresponding to the first sample time, the second prediction code corresponding to the second sample time, and the third prediction code corresponding to the third sample time, a first distance corresponding to the positive sample combination is determined; For each negative sample combination, based on the first sample time, the second sample time, and the fourth sample time of the negative sample combination, the first prediction code corresponding to the first sample time, the second prediction code corresponding to the second sample time, and the fourth prediction code corresponding to the fourth sample time, a second distance corresponding to the negative sample combination is determined. The loss of the encoding network is determined based on the first distance corresponding to the positive sample combination and the second distance corresponding to each negative sample combination. Determine the parameter change value of the encoding network when the loss reaches its minimum value, and update the parameters of the encoding network based on the parameter change value.
21. The apparatus according to claim 20, characterized in that, The encoding module is also used for: Determine the first sample ratio of the length of the third sample time period to the length of the first sample time period; Determine a first sample difference between a given ratio and the first sample, and multiply the first sample difference with the first prediction code to obtain a first sample product; The first sample ratio is multiplied by the second prediction code to obtain the second sample product, and the first sample product and the second sample product are summed to obtain the first sample sum result. The distance between the third predicted code and the sum of the first sample is taken as the first distance for the corresponding positive sample combination.
22. The apparatus according to claim 20, characterized in that, The encoding module is also used for: Obtain the standardized distance that is positively correlated with each of the second distances; The standardized distances are summed to obtain the sum of the second samples; Obtain the loss of the corresponding encoding network that is positively correlated with the first distance and negatively correlated with the sum of the second samples.
23. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; A processor, when executing computer-executable instructions stored in the memory, implements the artificial intelligence-based molecular processing method according to any one of claims 1 to 11.
24. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, they implement the artificial intelligence-based molecular processing method according to any one of claims 1 to 11.
25. A computer program product comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer-executable instructions are executed by a processor, they implement the artificial intelligence-based molecular processing method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Molecular feature determination method, related device and equipment
CN114664391A
An assembly and a method suitable for identifying a code sequence of a biomolecule
EP0732584A2