Chip architecture simulation system for molecular dynamics
By designing a chip architecture simulation system for molecular dynamics, the problem of insufficient optimization of molecular dynamics specific computing modes in the research of heterogeneous accelerators is solved, and the efficient simulation of the DeePMD algorithm and the support of the flow computing architecture is realized, improving the accuracy and efficiency of molecular dynamics simulation.
Patent Information
- Application Number
- CN202510216967.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In the research of heterogeneous accelerators for the field of molecular dynamics, existing simulators lack in in-depth optimization of specific computational patterns of molecular dynamics, lack of dedicated simulation frameworks for DeePMD algorithms, and limited support for flow computing architectures.
A chip architecture simulation system for molecular dynamics is designed, which includes a programmable IO module, an address mapping module, a data transmission module and a heterogeneous accelerator. The heterogeneous accelerator uses a flow computing architecture to simulate the stress calculation process of the DeePMD model on the atoms, and dynamically optimizes the parallelism degree through the parallel degree scheduling submodule to minimize the stress calculation time.
It realizes a dedicated simulation framework for accelerators for DeePMD algorithm, improves support for the flow computing architecture, realizes continuous flow processing of data, calculates the stress of atoms in real time, accurately obtains the stress conditions of each atom in the molecular system, provides accurate data support for molecular dynamics simulation, and optimizes the accelerator performance in specific calculation modes of molecular dynamics.
Smart Images

Figure CN120032731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of molecular dynamics calculations, in particular to the field of computing chip architecture simulations, and more particularly to a chip architecture simulation system for molecular dynamics. Background Art
[0002] Molecular dynamics (MD) is one of the important research directions in the field of modern scientific computing. MD simulation obtains the dynamic and thermodynamic properties of the system by calculating the time evolution of the position and velocity of atoms in a microscopic system. In recent years, machine learning has played an increasingly important role in the field of high-performance computing, among which the Neural Network Molecular Dynamics (NNMD) algorithm has become the mainstream method in this field. NNMD significantly improves the computational efficiency while ensuring the computational accuracy by performing machine learning training on first-principles data. In particular, the performance of the Deep Potential (DP) model is particularly outstanding, and it can simultaneously achieve the same accuracy as the Ab Initio Molecular Dynamics (AIMD) and the computational efficiency of the Empirical Force Field (EFF) method.
[0003] However, the current acceleration solutions based on the Deep Potential Molecular Dynamics (Deep Potential Molecular Dynamics) model mainly rely on general-purpose high-performance computing devices (such as GPU, Summit supercomputer, Sunway Taihu Light, etc.), and these general-purpose architectures have obvious bottlenecks in computing scalability. Therefore, the development of dedicated hardware accelerators specifically for the DeePMD model has become a key way to solve the current challenges. This transition from a general-purpose architecture to a dedicated architecture requires deep customization of the computing unit, storage hierarchy, and on-chip interconnection to break through the limitations of existing simulation efficiency.
[0004] In terms of hardware simulation tools, Gem5 (GEMS and M5) and SystemC are currently the most representative simulation platforms. These tools have been widely recognized in the field of processor architecture research, especially in the cycle-accurate simulation of CPUs and GPUs. However, when it comes to heterogeneous accelerator research in the MD field, existing simulators still have the following shortcomings:
[0005] There is a lack of in-depth optimization for specific computational modes of molecular dynamics; there is no specialized simulation framework for the characteristics of the DeePMD algorithm; and there is limited support for streaming computing architectures.
[0006] Among them, the streaming computing architecture is a computing model with data flow as the core. Its characteristics are that data flows between processing units in a pipeline manner, and the computing tasks are broken down into multiple stages, each of which is processed by a dedicated hardware unit. This architecture is particularly suitable for handling large-scale parallel calculations of atomic interactions in molecular dynamics simulations through pipeline processing of data flow. The streaming computing architecture is also commonly referred to as a pipeline by reducing data handling and memory access delays.
[0007] Existing simulators (such as Gem5 and SystemC) are mainly designed for general computing architectures and lack optimization for molecular dynamics-specific computing modes, resulting in limited support for streaming computing architectures. The reasons for the limited support for streaming computing architectures are as follows:
[0008] Lack of data flow optimization: Existing simulators are not optimized for the data flow characteristics in molecular dynamics and cannot efficiently simulate the force calculation process;
[0009] Insufficient task decomposition and scheduling: The streaming computing architecture requires computing tasks to be decomposed into multiple stages and dynamically scheduled, but existing simulators lack a customized force calculation framework;
[0010] Insufficient support for parallelism: Streaming computing architectures require high-parallel hardware support, but existing simulators have limitations in parallel modeling and scheduling.
[0011] Based on the above analysis, the existing simulators have the following problems:
[0012] 1. It is mainly designed for general computing architecture, resulting in the current lack of optimization for specific molecular dynamics computing modes; 2. There is a lack of a dedicated simulation framework for the DeePMD algorithm, which makes it impossible to efficiently simulate the force calculation process, resulting in insufficient support for the characteristics of the DeePMD algorithm; 3. Existing simulators have limitations in task decomposition, scheduling, and parallelism modeling, resulting in limited support for streaming computing architectures by existing simulators.
[0013] It should be noted that this background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily the prior art. The relevant information is submitted and disclosed together with the present invention solution. If there is no evidence that the relevant information has been disclosed before the application date of the present invention, the relevant information shall not be regarded as the prior art. Summary of the invention
[0014] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a chip architecture simulation system for molecular dynamics.
[0015] The objective of the present invention is achieved through the following technical solutions:
[0016] According to a first aspect of the present invention, a chip architecture simulation system for molecular dynamics is provided, the system is used to simulate the force calculation process of an accelerator for a DeePMD model, the system includes a programmable IO module, an address mapping module, a data transmission module and a heterogeneous accelerator; wherein the programmable IO module is used to access and control the heterogeneous accelerator through a host; the address mapping module is used to establish a mapping relationship between a virtual address of a host-side memory and a physical address of a heterogeneous accelerator; the data transmission module is used to perform data transmission between the host-side memory and the heterogeneous accelerator according to the mapping relationship, including transmitting atomic information involved in the calculation in molecular dynamics; the heterogeneous accelerator includes: a computing logic submodule, which is used to simulate the force calculation process of the DeePMD model on atoms according to the atomic information, calculate the forces on the atoms, and count the force calculation time; a resource modeling submodule, which is used to evaluate the resource occupancy rate of the computing logic submodule during the force calculation process; a parallelism scheduling submodule, which is used to dynamically optimize the parallelism of the computing logic submodule under the constraint that the resource occupancy rate is less than a preset threshold, and the optimization goal is to minimize the force calculation time.
[0017] In some embodiments of the present invention, the computing logic submodule includes: an atomic management unit, which is used to manage atomic information and task scheduling of the atomic force calculation process, wherein the atomic information includes multiple atoms, the initial position and initial velocity of each atom; a general processing unit including multiple computing units, which is used to simulate multiple reasoning stages of the atomic force calculation process according to the atomic information using a streaming computing architecture through multiple computing units, obtain the atomic force, and count the force calculation time, wherein the multiple reasoning stages are obtained based on the decomposition of the atomic force calculation process of the DeePMD model.
[0018] In some embodiments of the present invention, the multiple computing units in the total processing unit are connected in sequence, and a buffer component is included between each two adjacent computing units, which is used to cache data and perform data transmission between the two computing units; wherein the multiple computing units adopt a streaming computing architecture including: simulating an inference stage through each computing unit, obtaining the inference result of the stage, and transmitting the inference result to the next target computing unit adjacent to the computing unit through the buffer component.
[0019] In some embodiments of the present invention, the buffer component includes: a data packet queue for caching data packets, a status flag for indicating the busy or idle state of the target computing unit, and a data packet counter for counting the number of data packets; wherein, the buffer component is configured to: determine whether the target computing unit is idle according to the status flag, and if so, schedule the data packet queue to transfer the cached data packets to the target computing unit according to the first-in-first-out principle and update the number of data packets, otherwise, stop the data packet transfer.
[0020] In some embodiments of the present invention, the parallelism scheduling sub-module is configured to: under the constraint that the resource occupancy rate is less than a preset threshold, dynamically optimize the parallelism of each computing unit in the computing logic sub-module, and the optimization goal is to minimize the force calculation time; wherein, the system adjusts the number of buffer components between every two adjacent computing units or the capacity of the cached data in the buffer component according to the optimized parallelism of each computing unit.
[0021] In some embodiments of the present invention, the system is configured to use a heterogeneous accelerator to calculate the movement trajectories of each atom in the following manner: calculate the force on each atom at each moment within a predetermined time, the force on each atom at the initial moment is calculated based on its initial position and initial velocity, and the force on each atom at each subsequent moment is calculated based on its velocity and position updated at the previous moment; update the velocity and position of each atom at the next moment based on the force on each atom at each moment within the predetermined time; obtain the movement trajectories of each atom within the predetermined time based on the positions of each atom at all moments within the predetermined time.
[0022] In some embodiments of the present invention, the multiple computing units include: a neighbor atom filtering unit for screening valid atom pairs according to the positions of each atom, extracting the initial features of the valid atom pairs to obtain an initial matrix, and a valid atom pair means an atom pair in which two atoms are adjacent and meet a preset screening condition; an embedding vector calculation unit for calculating the embedding vector of each atom pair and obtaining an embedding matrix based on the embedding vectors of all atom pairs; a description matrix calculation unit for calculating an atomic environment local description matrix based on the initial matrix and the embedding matrix; a neural network prediction unit for predicting the potential energy of each atom based on the description matrix; and a force calculation unit for calculating the force on each atom according to the potential energy of each atom.
[0023] In some embodiments of the present invention, the atom management unit includes: a velocity calculation unit for calculating the acceleration of an atom at each moment based on the force on the atom at each moment; a displacement calculation unit for calculating the displacement of the atom from each moment to the next moment based on the acceleration of the atom at each moment and the velocity updated at the previous moment; and a position update unit for updating the position of the atom at the next moment based on the displacement of the atom from each moment to the next moment.
[0024] In some embodiments of the present invention, the data transmission module is configured to: receive a data access request initiated by a heterogeneous accelerator to a memory, process the request, and based on a data feedback mechanism, receive data returned by the memory according to the request and transmit it to the heterogeneous accelerator.
[0025] In some embodiments of the present invention, the computing process of the heterogeneous accelerator includes:
[0026] Check whether the data involved in the calculation is ready, including checking whether there are any data packets to be processed in the heterogeneous accelerator; according to the tasks scheduled by the atomic management unit, use the computing unit of the general processing unit to complete the corresponding reasoning stage in the atomic force calculation process and obtain the reasoning result; cache the reasoning result in the buffer component for transmission to the next target computing unit; schedule the next task of the force calculation process based on the event-driven mechanism; determine whether the scheduling of the next task is completed. If not, repeat the above process. Otherwise, output the calculation result.
[0027] Compared with the prior art, the advantages of the present invention are:
[0028] First, the system of the present invention implements a dedicated simulation framework for the accelerator for the DeePMD algorithm, providing support for the subsequent efficient simulation of the force calculation process. Secondly, the computing logic submodule of the heterogeneous accelerator adopts a streaming computing architecture to simulate the force calculation process of the DeePMD model on atoms. The streaming computing architecture enables the present invention to achieve continuous flow processing of data in terms of force calculation, realize real-time calculation of atomic forces, accurately obtain the force conditions of each atom in the molecular system, and provide accurate data support for molecular dynamics simulation. Finally, by dynamically optimizing the parallelism of the computing logic submodule with minimizing the force calculation time as the optimization goal, the optimization of the dedicated accelerator under the specific calculation mode of molecular dynamics of the present invention is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:
[0030] Figure 1 A schematic diagram of the structure of a chip architecture simulation system for molecular dynamics according to an embodiment of the present invention;
[0031] Figure 2 A schematic diagram of the data interaction principle of a system according to an embodiment of the present invention;
[0032] Figure 3 A schematic diagram of the structure of a computing logic submodule according to an embodiment of the present invention;
[0033] Figure 4A schematic diagram of a data transmission structure principle of a general processing unit according to an embodiment of the present invention;
[0034] Figure 5 A schematic diagram of the principle of the startup process of a heterogeneous accelerator according to an embodiment of the present invention;
[0035] Figure 6 Schematic diagram of the computing process principle of a heterogeneous accelerator according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0037] As mentioned in the background technology section, existing simulators have the following problems: 1. They are mainly designed for general computing architectures, resulting in the current lack of optimization for specific molecular dynamics computing modes; 2. There is a lack of a dedicated simulation framework for the DeePMD algorithm, and the force calculation process cannot be efficiently simulated, resulting in insufficient support for the characteristics of the DeePMD algorithm; 3. Existing simulators have limitations in task decomposition, scheduling, and parallelism modeling, resulting in limited support for streaming computing architectures in existing simulators.
[0038] In response to the above problems, the inventors analyzed and found that the simulation acceleration work of the existing general computing architecture using a combination of traditional CPU and GPU has reached a bottleneck and is difficult to improve further, and there is a severe lack of chip architecture simulation for molecular dynamics. In order to simulate longer trajectories in a short time, the molecular dynamics simulation algorithm itself is constantly improving. Although there are many algorithms that can simulate molecular dynamics motion, the DeePMD model takes into account both accuracy and speed. Therefore, the inventors turned to heterogeneous chip design and mapped it to proprietary hardware for calculation based on the DeePMD model. However, the chip design process is complicated, and the inventors further used simulators to pre-verify and evaluate heterogeneous chip architectures, so as to provide an important theoretical basis and tool support for the design and optimization of molecular dynamics-specific chips.
[0039] Based on the above analysis, the inventor proposes a chip architecture simulation system for molecular dynamics. First, the system includes a programmable IO module, an address mapping module, a data transmission module and a heterogeneous accelerator. The communication between the host and the heterogeneous accelerator is realized through the above three modules to control the heterogeneous accelerator to simulate the calculation process of the accelerator for the DeePMD model, realize the dedicated simulation framework of the accelerator for the DeePMD algorithm, and provide support for the subsequent efficient simulation of the force calculation process. Secondly, the heterogeneous accelerator includes a computing logic submodule, a resource modeling submodule and a parallelism scheduling submodule. The computing logic submodule can support the streaming computing architecture and realize the simulation of the DeePMD model for the atomic force calculation process. The streaming computing architecture enables the present invention to realize continuous flow processing of data in terms of force calculation, avoiding the delay and resource waste caused by the batch processing of data, that is, to calculate the force of atoms in real time, accurately obtain the force conditions of each atom in the molecular system, and provide accurate data support for molecular dynamics simulation. Finally, the resource modeling submodule is used to evaluate the resource occupancy rate of the computing logic submodule during the force calculation process, and the parallelism scheduling submodule is used to dynamically optimize the parallelism of the computing logic submodule with the optimization goal of minimizing the force calculation time under the constraint that the resource occupancy rate is less than a preset threshold, thereby realizing the optimization of the dedicated accelerator of the present invention under the specific computing mode of molecular dynamics.
[0040] According to an embodiment of the present invention, the simulation system architecture is implemented based on Gem5 extension, and Gem5 itself is open source and can be freely extended. For example, the address mapping module and the data transmission module are both implemented based on Gem5 extension. It should be understood that this is only for illustration, and the simulation system architecture can also be implemented based on SystemC extension, which is not limited by the present invention. The following description is based on Gem5 extension.
[0041] According to one embodiment of the present invention, see Figure 1, which is a schematic diagram of the chip architecture simulation system structure for molecular dynamics. The system is used to simulate the force calculation process of the accelerator for the DeePMD model. The system includes a programmable IO module, an address mapping module, a data transmission module and a heterogeneous accelerator, and the heterogeneous accelerator includes a computing logic submodule, a parallelism scheduling submodule and a resource modeling submodule; wherein the programmable IO module is used to access and control the heterogeneous accelerator through the host; the address mapping module is used to establish a mapping relationship between the virtual address of the host-side memory and the physical address of the heterogeneous accelerator; the data transmission module is used to perform data transmission between the host-side memory and the heterogeneous accelerator according to the mapping relationship, including the transmission of atomic information involved in the calculation of molecular dynamics; the computing logic submodule is used to simulate the force calculation process of the DeePMD model on the atoms according to the atomic information, calculate the atomic force, and count the force calculation time; the resource modeling submodule is used to evaluate the resource occupancy rate of the computing logic submodule during the force calculation process; the parallelism scheduling submodule is used to dynamically optimize the parallelism of the computing logic submodule under the constraint that the resource occupancy rate is less than the preset threshold, and the optimization goal is to minimize the force calculation time.
[0042] In order to better understand the present invention, the system simulation, the interaction between the various modules of the system, and the structural functions are described in detail below in conjunction with specific embodiments.
[0043] 1. System Simulation
[0044] According to one embodiment of the present invention, the system simulates the scientific calculation of molecular dynamics. Schematically, given a molecular system, the future motion trajectory of the molecular system can be simulated, which is usually applied to the fields of materials science, chemistry, protein, medicine, etc. The molecular system includes multiple atoms and information about each atom, and the information about each atom includes size, initial position coordinates and initial velocity. The scientific calculation of molecular dynamics includes: describing the interaction force between atoms according to the information of each atom in the molecular system, that is, calculating the force on the atoms.
[0045] According to one embodiment of the present invention, the force calculation of atoms in MD simulation: first calculate the interaction force between atoms, wherein the force is usually calculated at the atomic level, even if the atoms belong to the same molecule, it is necessary to calculate the interaction between them, such as, for copper (Cu) metal, there are only Cu atoms, so only the force of the Cu-Cu atomic pair needs to be calculated, and the water molecule needs to calculate the forces of HO and HH respectively; then calculate the overall force of the atom, the overall force of the atom is the sum of the forces between itself and other atoms. Among them, in order to better simulate the microscopic motion state, the simulation interval is usually at the femtosecond (fs) level, and the simulation calculation is continuously iterated at each femtosecond interval to achieve the motion trajectory simulation within a predetermined time period. Among them, 1fs= s, if you want to simulate the motion trajectory of the molecular system in 1 s, you need to perform it at intervals of every femtosecond Iteration simulation calculation.
[0046] According to one embodiment of the present invention, the system is configured to simulate the motion trajectory of each atom in a molecular system using a heterogeneous accelerator in the following manner: calculate the force on each atom at each moment within a predetermined time, the force on each atom at the initial moment is calculated based on its initial position and initial velocity, and the force on each atom at each subsequent moment is calculated based on its velocity and position updated at the previous moment; based on the force on each atom at each moment within the predetermined time, update the velocity and position of each atom at the next moment; obtain the motion trajectory of each atom within the predetermined time based on the position of each atom at all moments within the predetermined time.
[0047] 2. Data Interaction in the System
[0048] According to an embodiment of the present invention, a programmable IO (PIO) module, an address mapping module and a data transmission module are used to implement efficient data interaction between a host (CPU), a heterogeneous accelerator and a host-side memory.
[0049] According to one embodiment of the present invention, see Figure 2 , which is a schematic diagram of the data interaction principle of the system. In the figure, the interaction includes ① user program reading and writing data: that is, realizing data interaction between the host and the heterogeneous accelerator through the programmable IO module to access and control the heterogeneous accelerator; ② memory data access request: when the heterogeneous accelerator receives the startup control sent by the host, it initiates a data access request to the memory through the data transmission module; ③ memory data return: the data transmission module receives the data access request initiated by the heterogeneous accelerator to the memory, processes the request, and based on the data return mechanism, receives the data returned by the memory according to the request, and transmits the returned data to the heterogeneous accelerator. The following is an explanation of the interaction of the above three parts:
[0050] 1) User program reads and writes data
[0051] According to one embodiment of the present invention, user programs read and write data: This is achieved through the interaction between the host and the heterogeneous accelerator. Among them, the data interaction between the host and the heterogeneous accelerator: the PIO mechanism is adopted through the programmable IO module to realize direct communication between the host and the heterogeneous accelerator, so as to realize the function of access control of the heterogeneous accelerator. Among them, the host includes control registers and data registers, and the control registers are used to set the working mode, startup parameters and other control data of the heterogeneous accelerator; the data registers are used to transmit small-scale real-time data. The host can access the control registers and data registers through the operation of user program reading and writing data, thereby realizing the configuration management of the heterogeneous accelerator in the user space.
[0052] Schematically, the operation of reading and writing data through the user program of the host corresponds to the operation of the data structure inside the heterogeneous accelerator; therefore, although the reading and writing start from the registers of the external host, what is actually changed is the control data information inside the heterogeneous accelerator, including the metadata required for controlling the internal working mode and calculation of the heterogeneous accelerator, such as the initialization startup data.
[0053] 2) Memory data access request
[0054] According to one embodiment of the present invention, memory data access request: is implemented through the interaction between memory and heterogeneous accelerator. That is, when the heterogeneous accelerator needs to access memory data, it initiates a data access request to the memory through the data transmission module. The process of data access request includes: address conversion: converting the address request on the heterogeneous accelerator side into the host-side memory address; access arbitration: managing and coordinating multiple concurrent data access requests; request distribution: correctly routing the data access request to the target address space of the host-side memory.
[0055] The data transmission module includes multiple function interfaces, including:
[0056] 1. recvTimingReq(Pkt), used to receive data acquisition request;
[0057] 2. recvTimingResp(Pkt), used to receive data response request;
[0058] 3. handRequest(Pkt), used to process data acquisition requests;
[0059] 4. readPI, used for data read request;
[0060] 5. writePI, used for data write request;
[0061] 6. accessMemory, used to respond to data request events.
[0062] Among them, the recvTimingReq(Pkt), recvTimingResp(Pkt) and handRequest(Pkt) interfaces are inherited from Gem5 itself, and the readPI, writePI and accessMemory interfaces are extended implementations.
[0063] 3) Memory data transmission
[0064] According to one embodiment of the present invention, memory data feedback is achieved through the interaction between memory and heterogeneous accelerator. After the heterogeneous accelerator initiates a data access request, the memory data is transmitted to the heterogeneous accelerator through the memory data feedback mechanism. To ensure that the heterogeneous accelerator obtains the initialization data required for calculation, data feedback is implemented by expanding the function interface memCallback in the data transmission module, and triggering the calculation event callback on the heterogeneous accelerator side to form a complete heterogeneous accelerator operation chain. memCallback is based on the Gem5 expansion design. The data feedback mechanism includes: 1. Data caching: setting a data buffer on the heterogeneous accelerator side; 2. Callback processing: the heterogeneous accelerator receives and processes the data returned by the memory; 3. Data distribution: the heterogeneous accelerator distributes the acquired data to the computing logic sub-module therein.
[0065] The technical solution of the embodiment of the second part above can at least achieve the following beneficial technical effects: a flexible mechanism is implemented for user space to directly control heterogeneous accelerators, while ensuring efficient data access for large-scale molecular dynamics, providing a basic guarantee for the overall simulation performance of the system.
[0066] 3. Heterogeneous Accelerator Structure Principle
[0067] According to an embodiment of the present invention, the heterogeneous accelerator includes a computing logic submodule, a resource modeling submodule and a parallelism scheduling submodule. The three submodules of the heterogeneous accelerator are described below respectively:
[0068] 1) Computational logic submodule
[0069] According to one embodiment of the present invention, the computational logic submodule adopts an AM-PE heterogeneous architecture design to achieve efficient decoupling of atom management and force field calculation in molecular dynamics simulation. Figure 3, which is a schematic diagram of the structure of the computing logic submodule. The computing logic submodule in the figure includes a heterogeneous core. It should be understood that this is only for illustration, and the computing logic submodule can also be provided with multiple heterogeneous cores. Each heterogeneous core includes an atom management unit (Atom Manager, referred to as AM) and a total processing unit (Processing Element, referred to as PE). The atom management unit includes a velocity calculation unit, a displacement calculation unit and a position update unit. The velocity calculation unit is used to calculate the acceleration of the atom, the displacement calculation unit calculates the displacement of the atom according to the acceleration, and the position update unit updates the position of the atom according to the displacement. The total processing unit includes multiple computing units (Domain-Specific Accelerator, referred to as DSA), which are respectively recorded as computing unit 1, computing unit 2, ..., computing unit n-1, and computing unit n.
[0070] Among them, the atomic management unit is used to manage atomic information and task scheduling of the atomic force calculation process. The atomic information includes multiple atoms, the initial position and initial velocity of each atom; the total processing unit is used to simulate multiple reasoning stages of the atomic force calculation process according to the atomic information using a streaming computing architecture through multiple computing units, obtain the atomic force, and count the force calculation time. Among them, the multiple reasoning stages are obtained based on the decomposition of the DeePMD model for the atomic force calculation process.
[0071] According to one embodiment of the present invention, according to the strategy of decomposing and mapping the reasoning stage of the DeePMD model to hardware, multiple reasoning stages are obtained. Among them, the entire force calculation process of the total processing unit PE corresponds to the calculation process of the DeePMD model, and each DSA in the total processing unit PE is equivalent to a reasoning stage in the DeePMD model. Therefore, the number and function of the calculation unit can be set according to the DeePMD model. Schematically, if the multiple reasoning stages are 7, they are the neighbor atom filtering stage (Filter), the embedding vector calculation stage (Embedding), the description matrix calculation stage (Descriptor), the fully connected neural network calculation stage (FittingNet), the reverse description matrix calculation stage (DescriptorGrad), the reverse embedding vector calculation stage (EmbeddingGrad), and the force calculation stage (Force). Among them, the atomic pair screening and initial feature calculation are completed by Filter; Embedding generates the atomic pair feature vector; Descriptor constructs the local descriptor of the atomic environment; FittingNet predicts the potential energy of the atomic pair based on the neural network and calculates the interatomic force through Force and in combination with DescriptorGrad and EmbeddingGrad to obtain the atomic force. Therefore, the total processing unit includes at least 7 computing units, which are used to simulate the computing process of the above-mentioned stages respectively.
[0072] According to one embodiment of the present invention, the total processing unit schedules each computing unit to complete the calculation of the corresponding reasoning stage based on the event-driven mechanism of Gem5, and the event-driven mechanism is used to achieve precise synchronization between computing units, ensuring the timing accuracy of data processing. Now, in combination with a specific example, for example, the simulated molecular system is a copper (Cu) atomic system, and the scale of the overall Cu atomic system is very large, and it will be distributed to multiple heterogeneous cores for calculation. It is necessary to take the center Box as the origin in three-dimensional space and expand outward in three-dimensional space to form 27 Box areas of 3×3×3 (including the center Box itself), so as to decompose the entire atomic system into multiple Boxes, and each heterogeneous core calculates the 108 atoms of the corresponding Box. Below, taking a single heterogeneous core and taking the heterogeneous core including the above 7 computing units as an example, the reasoning process of the single heterogeneous core completing the force calculation of 108 local Cu atoms is explained based on the event-driven mechanism. The reasoning process is as follows:
[0073] 1. Computational unit corresponding to the Filter stage: Calculate the distance between the two atoms in each atom pair among all atom pairs, perform distance judgment to screen valid atom pairs, extract the initial features of valid atom pairs, and obtain the initial matrix R. First, use the preset screening conditions to select valid neighbor atoms belonging to each atom in the 108 atoms to obtain its neighbor atom table. The preset screening condition is that the distance between the two atoms is less than or equal to the cutoff radius (Rcut). Taking Rcut as 8 as an example, only when the distance between two atoms (i and j, respectively) is <= 8, atom j is regarded as the valid neighbor atom of i, and the atom forms a valid atom pair with each atom in its neighbor atom table.
[0074] Among them, when constructing the neighbor atom table, the atoms in the 26 neighbor Box areas around each center Box are considered as possible neighbors. For each atom i in the center Box, check each atom j among all the atoms in the 27 Box areas, and calculate whether the distance between atom i and atom j is <= Rcut. If so, add them to the neighbor atom table. Therefore, experience shows that the number of effective neighbor atoms of a single atom usually does not exceed 512, that is, the number of atoms in each neighbor atom table does not exceed 512. Therefore, in a single computing kernel, the final neighbor atom table can be represented by the initial matrix R with a matrix of dimension 108*512*4, where 4 represents the three-dimensional space (x-axis, y-axis, z-axis) and the actual distance s between the atom and each atom in its neighbor atom table.
[0075] 2. Computational unit corresponding to the Embedding stage: For each atom pair (i, j), the embedding vector is calculated, for example, using a fifth-order polynomial calculation to obtain an embedding matrix G of 512*128 dimensions.
[0076] 3. The calculation unit corresponding to the Descriptor stage: calculates the overall description matrix of 108 local atoms. The calculation process is to multiply the embedding matrix G and the initial matrix R to obtain a 108*2048 atomic environment local description matrix D.
[0077] 4. The computing unit corresponding to the FittingNet stage: The description matrix D of 108*2048 is used as input. The energy E of the molecular system is obtained by forward calculation through a three-layer fully connected neural network. E includes the potential energy of each atom. The derivative D' of the description matrix D is obtained by reverse differentiation of E, and its dimension is 108*2048.
[0078] 5. The calculation unit corresponding to the DescriptorGrad stage: derive D' to obtain the derivative R' of the initial matrix R and the derivative G' of the embedded matrix G.
[0079] 6. The calculation unit corresponding to the EmbeddingGrad stage: From R' and G', the derivative s' of the actual distance s between the atom and each atom in the neighboring atom table is obtained;
[0080] 7. The calculation unit corresponding to the Force stage: calculates the force, and obtains the force of the current valid atomic pair (i, j) in three-dimensional space (x-axis, y-axis, z-axis) from s'. The overall force of atom i is the sum of the forces of all corresponding valid atomic pairs.
[0081] According to one embodiment of the present invention, on the basis of the above embodiment, Force, DescriptorGrad and EmbeddingGrad can be re-divided into one reasoning stage, then the DeepPMD model is divided into 5 reasoning stages, and the corresponding multiple calculation units include 5, namely: a neighbor atom filtering unit, which is used to screen valid atom pairs and extract the initial features of valid atom pairs according to the positions of each atom to obtain an initial matrix, wherein the valid atom pairs represent atom pairs in which two atoms are adjacent and meet the preset screening conditions; an embedding vector calculation unit, which is used to calculate the embedding vector of each pair of atom pairs, and obtain the embedding matrix based on the embedding vectors of all atom pairs; a description matrix calculation unit, which is used to calculate the local description matrix of the atomic environment based on the initial matrix and the embedding matrix; a neural network prediction unit, which is used to predict the potential energy of each atom based on the description matrix; a force calculation unit, which is used to calculate the force of each atom pair formed by each atom and each neighbor atom according to the potential energy of each atom, wherein the force of the atom is the sum of the forces of all valid atom pairs corresponding to it.
[0082] The technical solution of the above embodiment can at least achieve the following beneficial technical effects: the general processing unit adopts a streaming computing architecture (also called a pipeline architecture), and simulates multiple reasoning stages through multiple customized computing units DSA to achieve efficient atomic force calculation. The currently implemented DSA is for the specific module reasoning stage of molecular dynamics calculation (such as Filter, Embedding, etc.), but the framework of the general processing unit can be flexibly extended to various heterogeneous chip architecture designs. That is, the design of the AM-PE heterogeneous architecture of the present invention not only improves the efficiency of calculations in the specific field of molecular dynamics, but also provides an effective solution for the rapid design verification of general heterogeneous computing architectures.
[0083] According to one embodiment of the present invention, in the calculation logic submodule, the general processing unit PE calculates the force of the atom at each moment according to the position of the atom at each moment, and the atom management unit includes: a speed calculation unit, which is used to calculate the acceleration of the atom at each moment based on the force of the atom at each moment; a displacement calculation unit, which is used to calculate the displacement of the atom from each moment to the next moment based on the acceleration of the atom at each moment and the speed updated at the previous moment; a position update unit, which is used to update the position of the atom at the next moment based on the displacement of the atom from each moment to the next moment. The present invention uses this Newtonian mechanics calculation method to calculate the force of the atom at the current moment, update the atomic position coordinates at the next moment according to the current force, and then calculate the force at the next moment, so as to repeat this cycle to achieve the simulation of the motion trajectory of the atom and obtain the microscopic motion trend of the molecular system.
[0084] According to an embodiment of the present invention, the statistical force calculation time is: the time required to calculate the force on the atom at each moment.
[0085] According to one embodiment of the present invention, the total processing unit realizes efficient data transmission between computing units based on the custom ModuleQueue mechanism of Gem5, that is, the computing units adopt a modular FIFO architecture (i.e., ModuleFIFO architecture) for data transmission, and support the transmission of different types of data packets. Multiple computing units in the total processing unit are connected in sequence, and a buffer component is included between each two adjacent computing units, which is used to cache data and perform data transmission between the two computing units; wherein, the multiple computing units adopt a streaming computing architecture including: simulating an inference stage through each computing unit, obtaining the inference result of the stage, and transmitting the inference result to the next target computing unit adjacent to the computing unit through the buffer component.
[0086] Schematically, see Figure 4 , which is a schematic diagram of the data transmission structure principle of the total processing unit. In the figure, input data to the total processing unit, after the computing unit 1 of the total processing unit performs the first reasoning stage processing, the reasoning result is transmitted to the buffer component indicated by its arrow, and then transmitted to the target computing unit 2 through the buffer component. After the computing unit 2 further performs the second reasoning stage processing, the reasoning result is transmitted to the buffer component indicated by its arrow..., and then processed by computing unit i,..., computing unit j,..., computing unit k, computing unit in turn, to obtain the calculation result. And a buffer component is included between each two adjacent computing units.
[0087] According to one embodiment of the present invention, the buffer component includes: a data packet queue (pkgs) for caching data packets, a status flag (status) for indicating the busy or idle state of the target computing unit, and a data packet counter (pkg_idx) for counting the number of data packets. The buffer component is configured to: determine whether the target computing unit is idle according to the status flag, and if so, schedule the data packet queue to transmit its cached data packets to the target computing unit using the first-in-first-out principle and update the number of data packets, otherwise, stop the data packet transmission. The number of buffer components or the capacity of cached data in the buffer components can be configured according to actual application requirements, such as setting multiple buffer components or setting multiple data packet queues.
[0088] The technical solution of the embodiment of the data transmission of the above-mentioned total processing unit can at least achieve the following beneficial technical effects: first, the ModuleFIFO architecture supports the transmission of different types of data packets, ensuring the versatility and scalability of the architecture; second, the introduction of the state management mechanism (i.e., status) avoids data congestion and improves the throughput of the system; finally, the configurable buffer component facilitates the optimization of resource allocation according to actual application requirements. The data transmission architecture not only realizes efficient collaboration between computing units DSA, but also provides a reliable experimental platform for performance evaluation and optimization of heterogeneous accelerators.
[0089] 2) Resource modeling submodule
[0090] According to one embodiment of the present invention, the resource modeling submodule is used to quantitatively evaluate the resource usage of each computing unit in the computing logic submodule, and provide a basis for the optimization of the computing logic submodule. For example, by using an automated tool (such as AuToFF) to evaluate the resource usage of the computing unit during molecular dynamics simulation, or by using a performance analysis tool (such as Gprof, Valgrind, Intel VTune) to perform performance analysis on the computing unit, the resource usage of each computing unit can be accurately measured.
[0091] 3) Parallelism scheduling submodule
[0092] According to an embodiment of the present invention, the parallelism scheduling sub-module is configured to: under the constraint that the resource occupancy rate is less than a preset threshold, dynamically optimize the parallelism of each computing unit in the computing logic sub-module, and the optimization objective is to minimize the force calculation time; wherein, the system adjusts the number of buffer components between every two adjacent computing units or the capacity of the cached data in the buffer components according to the optimized parallelism of each computing unit. For example, adjusting the number or width of the data packet queues in the buffer components to achieve capacity adjustment. The technical solution of this embodiment can at least achieve the following beneficial technical effects: realizing dynamic adjustment and optimization of the parallelism of each computing unit in the computing logic sub-module, and minimizing the force calculation time. The shorter the force calculation time, the faster the simulation speed, and the longer the motion trajectory that the present invention can simulate. In addition, the present invention can maximize the utilization rate of the hardware resources of each computing unit in the computing logic sub-module and achieve the best utilization rate of the heterogeneous accelerator.
[0093] According to an embodiment of the present invention, if the resource occupancy rate of computing unit 4 is relatively high, the number of its hardware components on the FPGA can only be 1; while for its adjacent computing units 3 and 5 before and after, according to the bandwidth limitation between computing units 3 and 5, the parallelism can be set to 2 to 8 (that is, both computing unit 1 and computing unit 3 can be set with multiple); the resource occupancy of computing unit 2 is extremely low, and the parallelism can be set to 4 to 16. Currently, the hardware chip is a single pipeline, and usually only 1 buffer component needs to be set. Among them, if computing units 1 and 2 need to process two threads in parallel (that is, the parallelism is 2), then 2 buffer components need to be set between computing units 1 and 2 or 2 data packet queues need to be set in one buffer component to support the simulation of the multi-thread pipeline.
[0094] The technical solution of the above-mentioned third part of the embodiment can at least achieve the following beneficial technical effects: Since in the actual hardware design, the heterogeneous accelerator itself is implemented using hardware-level code (Register-Transfer Level, abbreviated as RTL), where RTL implementation means converting the functional description of the design into a hardware description language at the register transfer level (such as Verilog or VHDL), thereby generating a specific hardware circuit. The present invention can provide a reliable platform for the architecture verification, optimization, and performance evaluation of the heterogeneous accelerator through the simulated heterogeneous accelerator in the simulation system, so as to optimize before generating the RTL code of the heterogeneous accelerator, improve the efficiency of the hardware design, and reduce the resource waste in the hardware implementation process.
[0095] IV. Principle of the overall execution process of the system
[0096] According to an embodiment of the present invention, the principle of the overall execution process of the system sequentially includes the startup process of the heterogeneous accelerator and the calculation process of the heterogeneous accelerator.
[0097] According to one embodiment of the present invention, see Figure 5 , which is a schematic diagram of the principle of the startup process of the heterogeneous accelerator, the startup process includes the following steps:
[0098] Step a01: Check whether the heterogeneous accelerator has received the start signal of the control register to complete the initialization. If the initialization is not completed, wait until the initialization is completed; if it is completed, proceed to the next step. This step can be completed by reading the status register of the heterogeneous accelerator on the host side to ensure that the heterogeneous accelerator has the basic conditions for normal operation.
[0099] Step a02: Detect whether the heterogeneous accelerator initiates a data access request to the host memory. In this step, the system monitors the memory access control signal of the heterogeneous accelerator to determine whether it is necessary to obtain the data required for calculation from the host memory. If there is a memory access request, the system will process the request through the address mapping module; if there is no memory access request, proceed to the next step.
[0100] Step a03: Execute a conventional register read and write operation. In this step, the host side controls the heterogeneous accelerator to initiate a data access request to the host side memory through a register read and write control signal.
[0101] Step a04: Determine whether data transmission and callback function execution are completed. The system confirms the completion of this step by checking the data transmission status and callback function execution status. Ensure that the heterogeneous accelerator successfully obtains the required computing data and completes the corresponding callback function processing.
[0102] Step a05: Determine whether the data involved in the calculation is ready. This step checks whether the data obtained from the memory has been correctly loaded into the local storage of the heterogeneous accelerator and the data preprocessing has been completed to ensure the correctness of the calculation.
[0103] Step a06: Verify whether the data involved in the calculation in the heterogeneous accelerator is ready. If it is ready, start the calculation process of the heterogeneous accelerator. The heterogeneous accelerator starts to execute the calculation task, including various calculation operations in the molecular dynamics simulation.
[0104] According to one embodiment of the present invention, see Figure 6 , which is a schematic diagram of the computing process principle of the heterogeneous accelerator, the computing process includes:
[0105] Step b01: Check whether the data involved in the calculation is ready, including checking whether there are data packets to be processed in the heterogeneous accelerator. This step is completed by checking the status of the ModuleFIFO buffer to ensure that there are data packets to be processed. If the data is not ready, the system will wait; if the data is ready, it will proceed to the next step.
[0106] Step b02: According to the tasks scheduled by the atomic management unit, use the computing units of the heterogeneous accelerator to complete the corresponding inference stage during the force calculation process of the atoms, and obtain the inference results. This step calls specific DSA computing units in the PE (such as neighbor atom filtering unit, embedding vector calculation unit, description matrix calculation unit, etc.) to complete the specific computing tasks related to atomic force calculation. The system processes data sequentially according to the preset computing process and in accordance with the deep pipeline architecture.
[0107] Step b03: Cache the inference results into the buffer component for transmission to the next target computing unit. This step uses the ModuleQueue mechanism to transfer the output results of the current computing unit to the next-level computing unit. The system checks the status flag of the target buffer component to ensure that data transmission does not cause buffer overflow.
[0108] Step b04: Based on the event-driven mechanism, schedule the next task in the force calculation process. This step is based on the event-driven mechanism of Gem5 to create and schedule corresponding computing events for the next computing cycle, ensuring that the computing unit can operate continuously and efficiently and make full use of the advantages of the pipeline architecture.
[0109] Step b05: Determine whether the scheduling of the next task is completed. If not, repeat the above process; otherwise, output the calculation results. The system checks the execution status of the current task to confirm whether data needs to be further processed. If the scheduling is not completed, return to step b01 to continue processing; otherwise, output the calculation results and set the accelerator operating status to exit.
[0110] To verify the beneficial effects of the present invention, the inventor conducted the following experiments:
[0111] 1): According to the heterogeneous accelerator designed in the simulation system of the present invention, simulate the movement trajectory of a molecular system containing only copper (Cu) atoms. The single-step iteration period can reach 2359 cycles; at the same time, the existing Vivado hardware simulation period is 2380 cycles, and the error between the present invention and the Vivado hardware simulation is only -0.88%, reaching the level of accurate simulation.
[0112] 2): Design two computing processes for the heterogeneous accelerator according to the simulation system of the present invention. The single-step simulation duration of the serial computing process is 87.41 microseconds at 2Ghz, and an acceleration of 11.4 times can be achieved; the single-step simulation duration of the parallel computing process is 1.19 microseconds at 2Ghz, and an acceleration of 420.2 times can be achieved.
[0113] Based on the above experimental results, it shows that: The heterogeneous accelerator designed in the simulation system of the present invention not only quickly simulates the molecular dynamics calculation process but also can achieve accurate simulation, and it also has a good auxiliary effect on the RTL code design of the heterogeneous accelerator.
[0114] It should be noted that although the above describes the various steps in a specific order, it does not mean that the various steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.
[0115] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0116] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a protruding structure in a groove on which instructions are stored, and any suitable combination thereof.
[0117] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A chip architecture simulation system for molecular dynamics, which is used to simulate the force calculation process of an accelerator for a DeePMD model. The system includes a programmable IO module, an address mapping module, a data transmission module, and a heterogeneous accelerator; in, Programmable IO modules for accessing and controlling heterogeneous accelerators through the host; An address mapping module is used to establish a mapping relationship between a virtual address of a host-side memory and a physical address of a heterogeneous accelerator side; A data transmission module is used to perform data transmission between the host memory and the heterogeneous accelerator according to the mapping relationship, including the transmission of atomic information involved in the calculation of molecular dynamics; Heterogeneous accelerators include: The calculation logic submodule is used to simulate the force calculation process of the DeePMD model on the atom based on the atomic information and adopt the streaming computing architecture to calculate the atomic force and count the force calculation time; The resource modeling submodule is used to evaluate the resource usage of the computational logic submodule during the force calculation process; The parallelism scheduling submodule is used to dynamically optimize the parallelism of the computing logic submodule under the constraint that the resource occupancy rate is less than a preset threshold, and the optimization goal is to minimize the load calculation time.
2. The system according to claim 1, characterized in that The computing logic submodule includes: Atom management unit, used to manage atomic information and task scheduling of atomic force calculation process. Atom information includes multiple atoms, initial position and initial velocity of each atom; A total processing unit including multiple computing units is used to adopt a streaming computing architecture through multiple computing units to simulate multiple reasoning stages of the force calculation process of atoms according to atomic information, obtain the forces of atoms, and count the force calculation time, wherein the multiple reasoning stages are obtained based on decomposing the DeePMD model's atomic force calculation process.
3. The system according to claim 2, characterized in that The plurality of computing units in the general processing unit are connected in sequence, and a buffer component is included between each two adjacent computing units, the component being used for caching data and performing data transmission between the two computing units; Among them, the method of using a streaming computing architecture for multiple computing units includes: simulating an inference stage through each computing unit, obtaining the inference result of the stage, and transmitting the inference result to the next target computing unit adjacent to the computing unit through a buffer component.
4. The system according to claim 3, characterized in that The buffer component includes: a data packet queue for caching data packets, a status flag for indicating the busy or idle status of the target computing unit, and a data packet counter for counting the number of data packets; The buffer component is configured to determine whether the target computing unit is idle according to the status flag. If so, the packet queue is scheduled to transmit the cached packets to the target computing unit using the first-in-first-out principle and update the number of packets. Otherwise, the packet transmission is stopped.
5. The system according to claim 3 or 4, characterized in that: The parallelism scheduling submodule is configured to: dynamically optimize the parallelism of each computing unit in the computing logic submodule under the constraint that the resource occupancy rate is less than a preset threshold, and the optimization goal is to minimize the load calculation time; The system adjusts the number of buffer components between every two adjacent computing units or the capacity of cached data in the buffer components according to the optimized parallelism of each computing unit.
6. The system according to claim 2, characterized in that The system is configured to calculate the motion trajectory of each atom using a heterogeneous accelerator in the following manner: Calculate the force on each atom at each moment within a predetermined time. The force on each atom at the initial moment is calculated based on its initial position and initial velocity. The force on each atom at each subsequent moment is calculated based on its velocity and position updated at the previous moment. Based on the force applied to each atom at each moment within a predetermined time, the velocity and position of each atom at the next moment are updated; The movement trajectory of each atom within the predetermined time is obtained based on the position of each atom at all times within the predetermined time.
7. The system according to claim 2 or 6, characterized in that: The plurality of computing units include: A neighbor atom filtering unit is used to screen valid atom pairs according to the positions of each atom, extract the initial features of the valid atom pairs, and obtain an initial matrix. The valid atom pairs represent the atom pairs in which two atoms are adjacent and meet the preset screening conditions. An embedding vector calculation unit, used to calculate the embedding vector of each atom pair, and obtain an embedding matrix based on the embedding vectors of all atom pairs; A description matrix calculation unit, used for calculating the atomic environment local description matrix based on the initial matrix and the embedded matrix; A neural network prediction unit, used for predicting the potential energy of each atom based on the description matrix; The force calculation unit is used to calculate the force on each atom according to the potential energy of each atom.
8. The system according to claim 6, characterized in that The atomic management unit includes: A velocity calculation unit, used for calculating the acceleration of the atom at each moment based on the force applied to the atom at each moment; A displacement calculation unit, used to calculate the displacement of the atom from each moment to the next moment based on the acceleration of the atom at each moment and the velocity updated at the previous moment; The position updating unit is used to update the position of the atom at the next moment based on the displacement of the atom from each moment to the next moment.
9. The system according to claim 2, characterized in that The data transmission module is configured to: receive a data access request initiated by the heterogeneous accelerator to the memory, process the request, and based on a data return mechanism, receive data returned by the memory according to the request and transmit it to the heterogeneous accelerator.
10. The system according to claim 3, characterized in that The computing process of the heterogeneous accelerator includes: Check whether the data involved in the calculation is ready, including checking whether there are data packets to be processed in the heterogeneous accelerator; According to the tasks scheduled by the atomic management unit, the calculation unit of the general processing unit is used to complete the corresponding reasoning stage in the atomic force calculation process to obtain the reasoning result; Cache the inference results in the buffer component for transmission to the next target computing unit; Based on the event-driven mechanism, schedule the next task of the force calculation process; Determine whether the scheduling of the next task is completed. If not, repeat the above process. Otherwise, output the calculation result.
Citation Information
Patent Citations
Parallelization acceleration method for a molecular dynamic simulation model
CN109871553A
Neural network reasoning acceleration method based on heterogeneous platform
CN114742225A
GPU parallel computing method, device and equipment for Prony analysis and medium
CN116339981A
Molecular dynamics simulation method, and method and device for training simulation model
CN116467921A
Chip simulation method and device for molecular dynamics calculation, equipment and storage medium
CN119004825A