Atomic Tersoft potential simulation method and system for many-core architecture processor

By using thread replicas to replace atomic locks, vectorized update stress and mixed precision optimization methods on multi-core architecture processors, the problem of low computational efficiency of Tersoff potential is solved, and efficient molecular dynamics simulation performance is achieved.

CN120388630AActive Publication Date: 2025-07-29SHANDONG UNIV

Patent Information

Application Number
CN202510883907.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

In the molecular dynamics simulation of the prior art, Tersoff potential calculation efficiency is low and cannot meet the high-performance computing needs. Especially in multi-core architecture processors, it is difficult to make full use of computing power, and the atomic lock operation is time-consuming, single-thread efficiency is low, vectorization processing is difficult, and computing resources are seriously wasted.

Method used

The atomic Tersoff potential simulation method of the multi-core architecture processor is adopted. By opening thread copy instead of atomic locks, preprocessing the binding of atomic tags to the adjacency table, initializing the array to store short neighbor atoms, vectorizing the stress, combining with mixing accuracy optimization, improving computing efficiency.

Benefits of technology

The Tersoff potential computing performance has been significantly improved, meeting the efficiency requirements of molecular dynamics simulation, and the optimized algorithm has been improved by more than 4.5 times on multi-core architecture processors, maintaining efficient computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388630A_ABST
    Figure CN120388630A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of molecular dynamics, and provides an atomic Tersoft potential simulation method and system of a many-core architecture processor in order to solve the problem that the efficiency requirement of molecular dynamics simulation cannot be met at present. The atomic Tersoft potential simulation method of the many-core architecture processor comprises the following steps: developing a copy of each thread to replace an atomic lock, pre-processing a label of an i atom corresponding to each thread, and reconstructing and binding with an adjacency list; initializing an array with the size of N * N; traversing N i atoms, screening short neighbor atoms of the i atoms, and filling the initialized array; transposing the filled array so as to carry out Tersoft potential gravity calculation of the N i atoms; and after the stress calculation of the short neighbors of the N i atoms is completed, vectorizing and updating the stress of the current N i atoms until the stress calculation of all i atoms is completed, reducing the stress of a copy of each thread, ending the calculation, and releasing all copy spaces, so that the calculation performance of Tersoft potential is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of molecular dynamics, and particularly relates to an atomic Tersoff potential simulation method and system for a many-core architecture processor. Background Technique

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] The demand for the efficiency of molecular dynamics simulation (MD) is increasing continuously. Taking the LAMMPS (Large-scale Atomic / Molecular Massively Parallel Simulator) molecular dynamics simulation software as an example, it can simulate the Tersoff potential of semiconductors and the related potential fields of soft materials and mesoscopic particles, etc. According to the principle formula of the Tersoff potential, the calculation of the Tersoff potential function is very complex, and in practice, the Tersoff potential simulation process is also very time-consuming: on a top desktop processor, for the Tersoff potential of a system with 512,000 atoms, only 0.4 ns can be simulated per day.

[0004] Although LAMMPS uses an adjacency list to eliminate unnecessary interactions between distant atoms, for some complex interactions with many transcendental mathematical functions such as the Tersoff potential, the calculation is still relatively intensive, and thus still cannot meet the demand for the efficiency of molecular dynamics simulation. Summary of the Invention

[0005] In order to solve the technical problems existing in the above background technique, the present invention provides an atomic Tersoff potential simulation method and system for a many-core architecture processor, which deeply analyzes and optimizes the atomic Tersoff potential simulation according to the hardware characteristics of the many-core architecture processor, can significantly improve the calculation performance of the Tersoff potential in molecular dynamics simulation software, and provides a more efficient calculation platform for molecular dynamics simulation in the field of materials science.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions: The first aspect of the present invention provides an atomic Tersoff potential simulation method for a many-core architecture processor.

[0007] An atomic Tersoff potential simulation method for a many-core architecture processor includes: Opening copies of each thread to replace the atomic lock, preprocessing the tags of the i atoms corresponding to each thread and binding them to the reconstruction of the adjacency list; Initializing an array of size N*N to store N short neighbor atoms of N i atoms; Traverse N i atoms, use the instruction set to filter out the short neighbor atoms of the i atoms from the adjacency list, and fill the initialized array; Transpose the filled array. The j-th row of the transposed array represents the j-th short neighbor of the N i atoms for calculating the Tersoff potential gravitational force of the N i atoms; After the force calculation of the short neighbors of the N i atoms is completed, vectorize and update the forces of the current N i atoms until the force calculation of all i atoms is completed, reduce the forces of the copies of each thread and end the calculation, and release all copy spaces; where N is the number of 32-bit integer type elements that can be processed by a vector unit in the many-core architecture processor; j is a positive integer not less than 1 and not greater than N.

[0008] The second aspect of the present invention provides an atomic Tersoff potential simulation system for a many-core architecture processor.

[0009] An atomic Tersoff potential simulation system for a many-core architecture processor includes: A label processing module, which is used to create copies of each thread to replace the atomic lock, preprocess the labels of the i atoms corresponding to each thread and bind them to the reconstruction of the adjacency list; An array initialization module, which is used to initialize an array of size N*N to store the N short neighbor atoms of the N i atoms; An array filling module, which is used to traverse N i atoms, use the instruction set to filter out the short neighbor atoms of the i atoms from the adjacency list, and fill the initialized array; A gravitational force calculation module, which is used to transpose the filled array. The j-th row of the transposed array represents the j-th short neighbor of the N i atoms for calculating the Tersoff potential gravitational force of the N i atoms; A copy reduction module, which is used to vectorize and update the forces of the current N i atoms after the force calculation of the short neighbors of the N i atoms is completed until the force calculation of all i atoms is completed, reduce the forces of the copies of each thread and end the calculation, and release all copy spaces; where N is the number of 32-bit integer type elements that can be processed by a vector unit in the many-core architecture processor; j is a positive integer not less than 1 and not greater than N.

[0010] The third aspect of the present invention provides a computer program product.

[0011] A computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps in the atomic Tersoff potential simulation method of the many-core architecture processor are implemented.

[0012] The fourth aspect of the present invention provides a many-core architecture processor.

[0013] A many-core architecture processor includes a memory, a processor body, and a computer program stored in the memory and executable on the processor body. When the processor body executes the program, the steps of the atomic Tersoff potential simulation method for the many-core architecture processor as described above are implemented.

[0014] Compared with the prior art, the present invention has the following beneficial effects: The present invention opens up copies of each thread to replace atomic locks, combines the hardware multi-threading characteristics of the many-core architecture processor, pre-processes the atomic labels corresponding to each thread and reconstructs and binds them with the adjacency list, and performs Tersoff potential gravity calculations on multiple atoms in each thread. Finally, after the force calculations of all atoms with the same label are completed, the force of the copies of each thread is reduced and the calculation is terminated, freeing up all copy space. This significantly improves the calculation performance of the Tersoff potential in the molecular dynamics simulation software and meets the efficiency requirements of molecular dynamics simulation.

[0015] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0017] Figure 1 This is a flow chart of a method for simulating atomic Tersoff potential of a many-core architecture processor according to an embodiment of the present invention; Figure 2 1 is a schematic diagram of an atomic Tersoff potential simulation process of a many-core architecture processor according to an embodiment of the present invention; Figure 3 Schematic diagram of atomic vectorized calculation according to an embodiment of the present invention; Figure 4 It is the short neighbor screening of an embodiment of the present invention; Figure 5 This is an example of the svhistcnt function in an embodiment of the present invention; Figure 6 1 is a schematic structural diagram of an atomic Tersoff potential simulation system for a many-core architecture processor according to an embodiment of the present invention; Figure 7 is the algorithm speedup ratio after optimization of the 4.096M atomic system in the embodiment of the present invention; Figure 8 is the multi-process speedup ratio of the algorithm after optimization of different atomic number systems in the embodiment of the present invention; Figure 9It is the strong scalability of the optimized algorithm of the 4.096M atomic system in the embodiment of the present invention. Detailed implementation manners

[0018] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0019] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0020] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0021] Term explanation: Tersoff potential function: For a system with n atoms, the potential function can be written as: (1); (2); Among them, represents the potential energy of the system of n atoms; and are the repulsive term and attractive term of the potential function, is the smoothing function, is the attractive term function caused by the action of other atoms, and its form is as follows: (3); (4); (5); (6); It involves the calculation of the three-body potential, and the formula is as follows: (7); (8); Among them is the vector and the included angle of, is a function of the bond angle, is The value at the minimum, among the parameters used in equations (1) to (8), , and used for the calculation of the two-body potential is determined by the types of atoms and . The one used for the calculation of the three-body potential is determined by the types of atoms , and , while and are both used in the calculations of the two-body and three-body potentials and are generally determined by the type of atoms.

[0022] In the calculation of the Tersoff potential function, generally the two-body potential is calculated first. During the calculation of the two-body potential, atoms that are close enough and require the calculation of the three-body potential are found. Then, the three-body potential is calculated among these atoms. Through the above analysis, the following main problems exist in the original algorithm: The OPENMP version built into the original algorithm is relatively simple and cannot fully utilize the powerful computing power of many-core architecture processors. At the same time, in the OPENMP version, the update of the atomic force is usually done by adding atomic locks to avoid write conflicts, and atomic locks are relatively time-consuming operations with low efficiency.

[0023] The original algorithm can only process the calculation of one i atom and its adjacent atoms each time for a single thread, lacking an effective vectorized version, wasting computing resources, and having low single-thread efficiency.

[0024] It is difficult to vectorize the original algorithm: 1) A large number of gather and scatter operations are required to obtain parameters, which is very time-consuming, and extensive use will make the algorithm inefficient. 2) In the algorithm, short neighbor screening of atoms is first performed, that is, atoms that are close enough and require the calculation of the three-body potential are found. The number of short neighbors of an atom is usually 4 or 5. During the calculation of the three-body potential, if vectorized operations are performed longitudinally on multiple short neighbors of an atom, there will be a situation where the vector bits are not fully filled during loop processing, resulting in low efficiency. If vectorized operations are performed horizontally on the same short neighbor of multiple atoms, the screening operation will be very difficult. 3) In a vectorized operation, there is a situation where the x-th short neighbor of atom a and atom b is both c, making it difficult to update the force vectorially.

[0025] The original algorithm calculates the two-body potential first and then the three-body potential, and the same data item is repeatedly calculated in different functions. For example, in the calculations of the repulsive term and the attractive term, the smoothing function is calculated, resulting in a waste of data and computing resources.

[0026] In the case where the adjacency list remains unchanged, the original algorithm recalculates whether the atoms in the adjacency list participate in the calculation in each iteration. There are a large number of judgment statements, which are not only inconvenient for vectorization but also waste resources due to repeated calculations.

[0027] Based on the original algorithm, the present invention creates copies for multi-threading, and optimizes the single-threaded model by adjusting the calculation order of repulsive and attractive forces, fusion function, vectorizing the calculation of multiple atoms, and adopting mixed precision to further improve the calculation efficiency, so as to complete the optimization of the algorithm on the many-core architecture processor. After analysis, this optimization can be divided into two directions: multi-threaded version optimization and single-threaded version vectorization optimization.

[0028] Figure 1 The flowchart of the atomic Tersoff potential simulation method for a many-core architecture processor according to an embodiment of the present invention is given. The atomic Tersoff potential simulation method for a many-core architecture processor according to an embodiment of the present invention specifically includes the following steps S101 to S105.

[0029] The following combines Figure 1 and Figure 2 to detail the specific implementation process of steps S101 to S105.

[0030] S101: Create copies for each thread to replace the atomic lock, preprocess the labels of the i-atoms corresponding to each thread and bind them to the reconstruction of the adjacency list.

[0031] In this embodiment, rewrite the multi-threaded version and replace the use of the atomic lock by using copies: among them, associate the release and reallocation of the copy space with the reconstruction of the adjacency list.

[0032] Use a multi-threaded parallel strategy at the i-atom loop level. Each thread calculates the force on the i-atoms with a dynamic quantity to ensure load balance among threads. Each thread calculates the force on the corresponding atom and updates the copy.

[0033] For example, for global variables such as eng_vdwl, eng_coul, virial, eatom, vatom, f in the pair class, use memory->create to create copies related to the number of threads, and point eng_vdwl_thr, eng_coul_thr, virial_thr, eatom_thr, vatom_thr, f_thr to the start address of the copy space, and associate the release and reallocation of the copy space with the reconstruction of the adjacency list. Only when neighbor->ago is equal to 0, the copy space will be released and reallocated.

[0034] At the level of the i-atom loop, use omp (multi-threaded parallel strategy). Each thread performs calculations on multiple i-atoms. At the same time, adopt the dynamic strategy to ensure load balance among threads. Each thread calculates the forces on the corresponding atoms and updates the copies.

[0035] After the i-atom loop ends, merge all the copies. Specifically, it is to reduce the copies of eatom_thr, vatom_thr, and f_thr of different i-atoms in multiple threads, and serially reduce the copies of eng_vdwl_thr, eng_coul_thr, and virial_thr.

[0036] S102: Initialize an array of size N*N to store the N short neighbor atoms of N i-atoms.

[0037] S103: Traverse the N i-atoms, and use the instruction set to filter out the short neighbor atoms of the i-atoms from the adjacency list and fill the initialized array.

[0038] S104: Transpose the filled array. The j-th row of the transposed array represents the j-th short neighbor of the N i-atoms to perform the Tersoff potential gravitational calculation for the N i-atoms.

[0039] S105: After the force calculation for the short neighbors of the N i-atoms is completed, vectorize and update the forces of the current N i-atoms until the force calculation for all i-atoms is completed. Reduce the forces of the copies of each thread and end the calculation, and release all the copy spaces; where N is the number of 32-bit integer type elements that can be processed by a vector unit in a many-core architecture processor; j is a positive integer not less than 1 and not greater than N. It should be noted here that steps S102 to S104 are all implementation processes in a single thread.

[0040] After analysis, there are three layers of loops in calculating the tersoff potential. However, the vectorization efficiency is not high in the second and third layers. Therefore, it is decided to use vectorization in the outermost layer of the loop. The process includes vectorized neighbor screening and three-body potential calculation. Denote N as the number of atoms that can be processed by a vectorization unit. Before the calculation, traverse the neighbors of the i-atoms, label the atoms participating in the repulsive function calculation, and associate the label array with the reconstruction of the adjacency list. Only when the neighbor->ago in the reconstruction of the adjacency list is equal to 0 will the label array be processed again.

[0041] The process of calculating the Tersoff potential gravitational force of N i-atoms in a single thread is as follows: Step a1: Initialize an array A of size N*N to store the N short neighbor atoms of N i-atoms.

[0042] Step a2: Traverse the N i atoms, and use the instruction set to effectively screen the short neighbor atoms of the i atoms, and fill the array A. The element (i, j) in this array represents the j-th short neighbor of the i-th atom. In Figure 3 Among them, the black box represents a vector calculation unit, blue represents the i atom, green represents the neighbor of the i atom, and orange represents the short neighbor of the i atom.

[0043] Step a3: Transpose the array. After transposition, the first row of the array represents the first short neighbor of the N i atoms, and so on. Subsequently, the Tersoff potential of multiple i atoms is calculated.

[0044] Step a4: Load the N processed i atoms into the vector unit as i_vec, and load the j-th short neighbor of the N i atoms into the vector unit as j_vec, with j initialized to 1.

[0045] Step a5: Load the k-th short neighbor of the N i atoms into the vector unit as k_vec, where k is not equal to j. Calculate zeta_vec through i_vec, j_vec, and k_vec, increment k, loop through all k_vec, with k < N, and accumulate the zeta_vec calculated in each round.

[0046] Step a6: Determine whether the short neighbor atoms in the current j_vec participate in the repulsive calculation through the label array. The atoms participating in the repulsive function calculation calculate an additional part of the force in the fused function force_zeta_and_repulsive, and perform an accumulation or subtraction operation on the forces of the i_vec and j_vec atoms.

[0047] Step a7: Load the k-th short neighbor of the N i atoms into the vector unit as k_vec, where k is not equal to j. Calculate attractive_vec through i_vec, j_vec, and k_vec, and vectorially update the forces of the k_vec short neighbor atoms in a loop, and perform an accumulation or subtraction operation on the forces of the j_vec and i_vec atoms.

[0048] Step a8: Vectorially update the forces of the j_vec short neighbor atoms in a loop, and sequentially process the remaining short neighbor atoms in a loop, increment j, with j < N.

[0049] Step a9: Vectorially update the forces of the i_vec short neighbor atoms.

[0050] Step a10: End the calculation of the current N i atoms, obtain the other N i atoms, and repeat steps a1 to a9 until the force calculations and updates of all i atoms are completed.

[0051] In some specific implementation processes, in combination with Figure 4, the process of screening out the short neighbor atoms of atom i from the adjacency list using the instruction set is as follows: Step b1. Initialize the array shortcnt[N] = {0}, which is used to store the number of short neighbors corresponding to N atom i.

[0052] Step b2. Load the N neighbor atoms j of atom i into a vector unit svint64_t. inum represents the order of the currently processed atom i, and inum < N.

[0053] Step b3. Vectorize the calculation of the coordinate interpolation delx, dely, delz between neighbor atom j and atom i, calculate the distance rsq, and compare it with the cutoff radius cutoff of atom i to generate a predicate mask svbool_t. The position corresponding to atom j with rsq < cutoff in the predicate mask is 1.

[0054] Step b4. Use the svcompact function to load the elements at the positions where the svbool_t is 1 in the svint64_t into a new vector svint64_t.

[0055] Step b5. Use the svcntp function to count the number of 1s in the predicate svbool_t to obtain the number numcompact of short neighbor atoms among the currently processed N atoms j.

[0056] Step b6. Use svst to store the short neighbor vector into the short neighbor array A[inum] + shortcnt[inum], and shortcnt[inum] += numcompact.

[0057] Step b7. Return to step b2 to process the other neighbor atoms of atom i. After all are processed, traverse the next atom i, and inum++.

[0058] In the original algorithm, after screening out the short neighbors, the two-body potential repulsive force is calculated after conditional judgment. The conditional judgment is relatively complex and not conducive to vectorized acceleration. Moreover, the content of the conditional judgment is related to the atomic attributes of the current adjacency list. When the adjacency list remains unchanged, the conditional judgment results for the same atoms remain the same. The judgment of the present invention is bound to the reconstruction of the adjacency list. Only when the adjacency list is rebuilt equal to 0 each time, it is judged whether the atomic attributes are the same; atoms with the same attributes are marked with tags, and only the atoms with the marked tags are calculated subsequently.

[0059] In the original algorithm, after screening out the short neighbors, the two-body potential repulsive force is calculated after conditional judgment, that is, the Repulsive function is called at the atom j level (the second loop), and the three-body potential gravitational force is calculated among the short neighbors, that is, a loop for atom k is carried out at the atom j level (the third loop).

[0060] Combining formulas (2), (6), and (7) mentioned in the background, it can be found that the calculations of , and all require the use of terms. Moreover, all atoms in the short neighbors participate in the gravitational calculation, and some atoms participate in the repulsive calculation, which can be controlled by svbool for the participating part.

[0061] At the same time, it can be found that the calculations of the Repulsive function and the calculation of function are both at the j-atom level. Therefore, it is considered that during the calculation of , the data is recorded. When calling the force_zeta function to calculate , the Repulsive function is added, that is, the original function force_zeta is fused with Repulsive, and some data is reused to reduce calculations and storage.

[0062] At the same time, for the discontinuous data used multiple times during the vectorization process, it is obtained by gather for the first time, and the obtained data is stored in an array. Subsequently, it is directly loaded from the array instead of using gather.

[0063] In the specific implementation process, during the calculation of the Tersoff potential gravity of N i-atoms, the forces on the N i-atoms are updated by vectorized loop. Among them, a set function is used to generate a new numerical vector to count the number of times an atom appears in the current vector.

[0064] During the update of the atomic force, if the same atom appears twice or three times in a vector calculation unit, it is incorrect to directly accumulate the corresponding force vector to the original force. It is corrected by using a loop update method: For example, as Figure 5 shown, the svhistcnt function is used to generate a new numerical vector to count the number of times an atom appears in the current vector. For an atom that appears once in the original vector, the corresponding position value in the new vector is 1. For an atom that appears twice, the value corresponding to the second appearance of the atom in the new vector is 2. For an atom that appears three times, the value corresponding to the third appearance of the atom in the new vector is 3.

[0065] Initialize count to 1, and use the svcmpne function to generate a svbool_t predicate. The positions corresponding to the count number of occurrences in the numerical vector of this predicate are 1. The svbool_t controls the atoms updated in the current round. count is incremented by 1 in each loop until the number of elements with the count number of occurrences in the numerical vector is 0 as counted by the svcntp function, and the loop ends.

[0066] In the original algorithm, double - type data is used for calculating the atomic force and the overall energy of the system. It is planned to mix precisions on the basis of vectorization. Since the atomic force data will be iteratively used, in the embodiments of the present invention, double precision is used for the cumulative force calculation, and float precision is used for the direct cumulative calculation of the overall energy.

[0067] The following gives the steps of mixed - precision calculation as follows: Step c1: Initialize the tag array.

[0068] Step c2: Change N = N * 2 in the vectorization step, initialize the array A[N * 2][N * 2], and traverse N * 2 i - atoms.

[0069] Step c3: According to the vectorization step, until the calculation of delx, dely, and delz is completed using double precision, i.e., svfloat64_t, in the short - range neighbor screening.

[0070] Step c4: Use the svcvt function to convert the calculated data to float precision, i.e., svfloat32_t, and all subsequent calculations use svfloat32_t.

[0071] Step c5: Until the force accumulation, use the combination functions svzip1, svzip2, and svcvt to convert svfloat32_t back to svfloat64_t, and use svunpklo_b and svunpkhi_b to split the predicate vector to reduce the atomic iterative force error.

[0072] Using mixed precision in this way can perform calculations for N * 2 atoms at a time, greatly improving the calculation efficiency.

[0073] Table 1 Test results (except for the number of atoms, all other data units are ns / day);

[0074] Table 2 Multi - process ns / day of the optimized algorithm with different numbers of atoms;

[0075] Table 3 Multi - process Timesteps / s of the optimized algorithm with different numbers of atoms;

[0076] According to the experiments shown in Tables 1 - 3, for the optimized single - thread algorithm running on a processor for a system with 32,000 atoms, the simulation efficiency is about 4.85 times that of the original algorithm. For a system running about 4M atoms, the simulation efficiency is about 4.83 times that of the original algorithm. For a system running about 4M atoms, the simulation efficiency is about 3.5 times that of the original algorithm, and the optimization effect is obvious. The speedup ratio of the optimized algorithm for the 4.096M - atom system and the multi - process speedup ratio of the optimized algorithm for systems with different numbers of atoms are as Figure 7 and Figure 8 shown.

[0077] According to equation (9), the parallel efficiency of the strong scaling of the optimized algorithm is calculated as Figure 9 shown. The data shows that when the number of processes increases to 128, the system efficiency is still above 65%, and a relatively high simulation speed can be achieved.

[0078] (9); wherein, represents the parallel efficiency of strong scaling; is the running time of N processes; is the running time of a single process.

[0079] Figure 6 is a schematic structural diagram of an atomic Tersoff potential simulation system for a many - core architecture processor. According to Figure 6 , the atomic Tersoff potential simulation system for a many - core architecture processor specifically includes the following modules: A label processing module 601, which is used to create copies of each thread to replace the atomic lock, pre - process the labels of the i - atoms corresponding to each thread, and bind them to the reconstruction of the adjacency list; An array initialization module 602, which is used to initialize an array of size N * N to store the N short - neighbor atoms of N i - atoms; An array filling module 603, which is used to traverse N i - atoms, use the instruction set to filter out the short - neighbor atoms of the i - atoms from the adjacency list, and fill the initialized array; A gravitational force calculation module 604, which is used to transpose the filled array. The j - th row of the transposed array represents the j - th short - neighbor of N i - atoms to perform the Tersoff potential gravitational force calculation for N i - atoms; A copy reduction module 605, which is used to, after the force calculation of the short - neighbors of N i - atoms is completed, vectorially update the forces of the current N i - atoms until the force calculation of all i - atoms is completed, reduce the forces of the copies of each thread and end the calculation, and release all copy spaces; where N is the number of 32 - bit integer - type elements that can be processed by a vector unit in a many - core architecture processor; j is a positive integer not less than 1 and not greater than N.

[0080] It should be noted here that the specific implementation processes in the label processing module 601, the array initialization module 602, the array filling module 603, the gravitational force calculation module 604, and the replica reduction module 605 in the atomic Tersoff potential simulation system of the many-core architecture processor correspond one by one to the steps in the above-mentioned atomic Tersoff potential simulation method of the many-core architecture processor, and their specific implementation processes are the same, so they will not be elaborated here.

[0081] In one or more embodiments, a many-core architecture processor is provided, including a memory, a processor body, and a computer program stored on the memory and executable on the processor body. When the processor body executes the program, it implements the steps in the atomic Tersoff potential simulation method of the many-core architecture processor as described above. Figure 1 shown.

[0082] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.

[0083] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0084] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0085] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An atomic Tersoff potential simulation method for a many-core architecture processor, characterized in that, including: Create copies of each thread to replace the atomic lock, preprocess the tags of the i-atoms corresponding to each thread and bind them to the reconstruction of the adjacency list; Initialize an array of size N*N to store the N short neighbor atoms of N i-atoms; Traverse the N i-atoms, use the instruction set to filter out the short neighbor atoms of the i-atoms from the adjacency list, and fill the initialized array; Transpose the filled array. The j-th row of the transposed array represents the j-th short neighbor of the N i-atoms to perform the Tersoff potential gravitational calculation for the N i-atoms; When the force calculation of the short neighbors of the N i-atoms is completed, vectorize and update the forces of the current N i-atoms until the force calculation of all i-atoms is completed, reduce the forces of the copies of each thread and end the calculation, and release all copy spaces; where N is the number of 32-bit integer type elements that can be processed by a vector unit in a many-core architecture processor; j is a positive integer not less than 1 and not greater than N.

2. The atomic Tersoff potential simulation method for the many-core architecture processor according to claim 1, wherein When the reconstruction of the adjacency list is equal to 0 each time, determine whether the atoms have the same attributes; mark the tags of the atoms with the same attributes, and only calculate the atoms with the marked tags subsequently.

3. The atomic Tersoff potential simulation method for the many-core architecture processor according to claim 1, wherein During the Tersoff potential gravitational calculation of the N i-atoms, vectorize and loop to update the forces of the N i-atoms.

4. The atomic Tersoff potential simulation method for the many-core architecture processor according to claim 3, characterized in that Use a set function to generate a new numerical vector to count the number of times an atom appears in the current vector.

5. The atomic Tersoff potential simulation method for the many-core architecture processor according to claim 1, wherein, Use double precision when accumulating force calculations and float precision when directly accumulating the overall energy calculation.

6. The atomic Tersoff potential simulation method for the many-core architecture processor according to claim 1, wherein Associate the release and reallocation of the copy space with the reconstruction of the adjacency list.

7. The atomic Tersoff potential simulation method for the many-core architecture processor according to claim 1, wherein Use a multi-threaded parallel strategy at the i-atom loop level. Each thread performs the force calculation with a dynamic number of i-atoms to ensure thread load balancing. Each thread calculates the force of the corresponding atoms and updates the copies.

8. An atomic Tersoff potential simulation system for a many-core architecture processor, characterized in that, including: A tag processing module, which is used to create copies of each thread to replace the atomic lock, preprocess the tags of the i-atoms corresponding to each thread and bind them to the reconstruction of the adjacency list; An array initialization module, which is used to initialize an array of size N*N to store the N short neighbor atoms of N i-atoms; An array filling module, which is used to traverse the N i-atoms, use the instruction set to filter out the short neighbor atoms of the i-atoms from the adjacency list, and fill the initialized array; A gravitational calculation module, which is used to transpose the filled array. The j-th row of the transposed array represents the j-th short neighbor of the N i-atoms to perform the Tersoff potential gravitational calculation for the N i-atoms; A copy reduction module, which is used to, when the force calculation of the short neighbors of the N i-atoms is completed, vectorize and update the forces of the current N i-atoms until the force calculation of all i-atoms is completed, reduce the forces of the copies of each thread and end the calculation, and release all copy spaces; where N is the number of 32-bit integer type elements that can be processed by a vector unit in a many-core architecture processor; j is a positive integer not less than 1 and not greater than N.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps in the atomic Tersoff potential simulation method of the many-core architecture processor described in any one of claims 1-7 are implemented.

10. A many-core architecture processor, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the atomic Tersoff potential simulation method of the many-core architecture processor described in any one of claims 1-7.

Citation Information

Patent Citations

  • Molecular dynamics simulation acting force analysis method and device and computer equipment

    CN118692576A

  • Method of simulating low molecule in macromolecule

    JP2011123874A

Cited By

  • Molecular dynamics simulation method and system for many-core architecture processor

    CN121306292A