A machine learning potential energy surface model construction method

By employing data-driven feature engineering and a multi-layered nested architecture for machine learning models, the problems of insufficient fitting accuracy and computational speed of existing machine learning potential surface models are solved, achieving broad applicability and efficient computation for multiphase and periodic/aperiodic systems.

CN117273169BActive Publication Date: 2026-02-10TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311242223.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2026-02-10
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

Existing machine learning potential surface models are insufficient in terms of fitting accuracy and computational speed, and are difficult to apply to multiphase systems and periodic/aperiodic systems.

Method used

A data-driven approach is used to construct a machine learning potential surface model. Through feature engineering, the model learns autonomously and iteratively, avoiding human bias. It supports multiphase and periodic/non-periodic systems and is trained using a multi-layered nested architecture machine learning model.

Benefits of technology

It improves the model's fitting accuracy and computation speed, is applicable to multiphase and periodic/aperiodic systems, shortens computation time, and enhances computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117273169B_ABST
    Figure CN117273169B_ABST
Patent Text Reader

Abstract

A machine learning potential energy surface model construction method comprises the following steps: collecting data for training a model, the data comprising data of one or more systems, the data of each system comprising atomic coordinates of the system and corresponding properties of the system; performing data inspection and data storage on the collected data; constructing a machine learning potential energy surface model, comprising: establishing a feature engineering; obtaining a final data set for training the model through the feature engineering; constructing a machine learning model; and training the machine learning model using the final data set to obtain the machine learning potential energy surface model. By using a data-driven method, the present application avoids manual feature extraction, and the model autonomously iterates learning from the data set, thereby avoiding the introduction of human bias. The model of the present application has very strong scalability, can be scaled to larger systems, and is not limited by the properties of the target system, and simultaneously supports multi-phase systems and periodic / non-periodic systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a potential energy surface model. In particular, it relates to a method for constructing a machine learning potential energy surface model. Background Technology

[0002] A potential energy surface (PES) is a theoretical model used in chemistry, physics, and materials science to study reaction kinetics. Its basic idea is to treat the state points of a reaction system (including parameters such as temperature, pressure, and substance concentration) as the values ​​of a multidimensional potential energy function, and to describe the changes of this potential energy function between different state points in a certain way, thereby deriving the kinetic process of the reaction.

[0003] Machine learning potential energy surface model (ML-PES model) refers to a machine learning model used to predict the properties of molecular systems. It learns the potential energy surface using machine learning algorithms, thus allowing for the prediction of the properties of the target structure. The advantage of this approach is that it eliminates the need to solve the Schrödinger equation from first-principles calculations, significantly accelerating the computation of structural properties and allowing for simulations on larger time and spatial scales with limited computational resources. Common ML-PES models include neural network models, support vector machine models, and decision tree models. They typically require large amounts of data for training and necessitate tuning of the model's hyperparameters to ensure a good fit to the data.

[0004] ML-PES models have a wide range of applications, including chemistry, physics, and materials science. For example, in chemistry, ML-PES models can be used to simulate chemical reactions, helping researchers find better catalysts. In physics and materials science, ML-PES models can be used to study and predict the physical and chemical properties of materials, helping materials scientists design new materials.

[0005] In practical applications, ML-PES models may also encounter some challenges. For example, building an ML-PES model requires a large amount of data for training, which places higher demands on both the quality and quantity of the data. Data derived from first-principles calculations consumes a significant amount of computing power. The hyperparameters of the model need to be adjusted to ensure that the model can fit the data well. Complex machine learning algorithms are required to improve the model's accuracy. The trade-off between computational accuracy and computational efficiency also exists in ML-PES models.

[0006] In its future development, the ML-PES model may continue to be inspired by computing power, machine learning techniques, and other fields, thereby improving its accuracy and efficiency. Furthermore, the ML-PES model may find wider application in new fields, providing us with more assistance. Its emergence offers a new method for studying chemistry and materials-related problems, and provides new ideas and methods for our research in related fields.

[0007] Numerous reports have been published on machine learning potential surface models. For example, Behler, in Chem. Rev. 2021, 121, 16, 10037–10072, categorizes them into four generations based on the types of interactions they can describe: the first generation is only applicable to low-dimensional systems; the second generation is based on environment-dependent atomic energy; the third generation includes long-range electrostatic interactions using local charges; and the fourth generation considers global charge distributions, including non-local charge transfer. In terms of final fitting accuracy, existing machine learning potential surface models can generally achieve the same accuracy as the DFT calculation, such as energy fitting accuracy of 1E0-1E2 meV / atom and force accuracy of [missing information]. Meanwhile, the computation time reaches the level of 1E-3 for typical DFT calculations, usually completing calculations that would otherwise take hours in milliseconds to seconds. However, common potential energy surface fitting schemes still have limitations in establishing potential energy surface models. While numerous feature engineering and model construction methods have been developed this century, their computational accuracy, speed, and application scope remain unsatisfactory, and there is still significant room for improvement in simulating multiphase systems. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a method for constructing a machine learning potential energy surface model that can improve the fitting accuracy and calculation speed of the machine learning potential energy surface model and has wide applicability to various systems.

[0009] The technical solution adopted in this invention is: a method for constructing a machine learning potential energy surface model, comprising the following steps:

[0010] 1) Collect data for training the model

[0011] The data includes data from more than one system, and the data for each system includes: the atomic coordinates of the system and the corresponding properties of the system; wherein,

[0012] The system properties described include one or both of the following: overall system properties and atomic-level properties; wherein,

[0013] The aforementioned atom-by-atom properties include one or more of the following: atomic force, atomic magnetism, atomic charge, and atomic ionization energy;

[0014] The overall properties of the system include one or more of the following: energy, cell stress, adsorption energy, and D-band centers;

[0015] 2) Perform data verification and data storage on the collected data;

[0016] 3) Construct a machine learning potential surface model, including:

[0017] (3.1) Establish feature engineering;

[0018] (3.2) Through feature engineering, the final dataset used to train the model is obtained;

[0019] (3.3) Constructing machine learning models;

[0020] (3.4) The machine learning model is trained using the final dataset to obtain the machine learning potential surface model.

[0021] This invention presents a machine learning potential surface model construction method that, by employing a data-driven approach, avoids manual feature extraction. The model learns autonomously and iteratively from the dataset, thus preventing the introduction of human bias. The extracted feature engineering precisely satisfies the physicochemical symmetry of the system, maintaining output invariance during translation, rotation, and permutations. Furthermore, since it requires only distance and a small number of transformation calculations, the computational cost of feature engineering is significantly less than various reported methods involving distance, angle, dihedral angle, or graphs. In addition, the model of this invention exhibits strong scalability, capable of scaling to larger systems without being limited by the properties of the target system, and supports both multiphase and periodic / aperiodic systems. In actual fitting tests, the fitting accuracy on the tested dataset also outperforms other reported potential surface models. Attached Figure Description

[0022] Figure 1 This is a flowchart of a machine learning potential surface model construction method according to the present invention;

[0023] Figure 2 This is a flowchart of the feature engineering construction method in this invention;

[0024] Figure 3a This is the overall structure diagram of the machine learning model in this invention;

[0025] Figure 3b This is a structural diagram of the single-atom model in this invention;

[0026] Figure 3c This is a structural diagram of the element subnetwork in this invention. Detailed Implementation

[0027] The following describes in detail a method for constructing a machine learning potential surface model according to the present invention, with reference to embodiments and accompanying drawings.

[0028] like Figure 1 As shown, a method for constructing a machine learning potential surface model according to the present invention includes the following steps:

[0029] 1) Collect data for training the model

[0030] The data includes data from more than one system, and the data for each system includes: the atomic coordinates of the system and the corresponding properties of the system, wherein:

[0031] The system properties described include one or both of the following: overall system properties and atomic-level properties.

[0032] Atom-by-atom properties include one or more of the following: atomic force, atomic magnetism, atomic charge, and atomic ionization energy;

[0033] The overall properties of the system include one or more of the following: energy, cell stress, adsorption energy, and D-band centers;

[0034] 2) Perform data verification and data storage on the collected data.

[0035] (2.1) Data verification includes two stages: data cleaning and data validation.

[0036] (2.1.1) The data cleaning mentioned includes removing invalid data, correcting erroneous data, and converting data formats.

[0037] Remove invalid data: This removes invalid data such as atomic coordinates or corresponding properties that are empty;

[0038] Convert data format: Convert the format used for data storage;

[0039] The data verification described in (2.1.2) is to verify the chemical structure and physicochemical properties, and to check the rationality and consistency of the data. This includes using general physicochemical rules to perform preliminary verification of the data and to exclude data that is obviously outside the normal range.

[0040] For example, one can check whether the chemical structure conforms to chemical rules and whether there are unreasonable bond lengths and bond angles; one can also check whether the physicochemical properties are within a reasonable range and whether they are consistent with known data. In addition, it is necessary to verify the calculation results of indirect properties such as adsorption energy, and calculations that do not converge need to be recalculated.

[0041] (2.2) Data storage

[0042] Data storage requires organizing, saving, and managing the acquired data. Data organization involves saving the data as designated data files or databases according to the system type. These data files or databases include: xyz files, database files, SQL databases, or MongoDB.

[0043] The data files or databases mentioned are built locally or in the cloud, and the data is stored in a location accessible during training;

[0044] Data management should consider data deduplication and data backup to ensure the security and validity of existing data.

[0045] 3) Constructing a machine learning potential surface model

[0046] (3.1) Establishing feature engineering

[0047] (3.1.1) Periodic treatment of the system

[0048] The periodicity mentioned includes one-dimensional periodicity, two-dimensional periodicity, and three-dimensional periodicity;

[0049] For a periodic system, the number of times the original unit cell is replicated is calculated based on the given maximum cutoff radius, and the original unit cell is replicated in the periodic direction, with the number of replications being the calculated number.

[0050] The formula for calculating the number of times the original unit cell is replicated is as follows:

[0051] Let the unit cell parameter be denoted as LatticeVector, which is a 3x3 matrix, such as [[100,0,0],[0,100,0],[0,0,100]].

[0052] Let the maximum cutoff radius be MaxRc, then the number of replications in the three directions of the unit cell is as follows:

[0053] The number of replications in the first direction of the unit cell is:

[0054] ceil(MaxRc / abs(LatticeVector[0]·(LatticeVector[1]×LatticeVector[2]) / |LatticeVector[1]×LatticeVector[2]|))

[0055] The number of replications in the second direction of the unit cell is:

[0056] ceil(MaxRc / abs(LatticeVector[1]·(LatticeVector[0]×LatticeVector[2]) / |LatticeVector[0]×LatticeVector[2]|))

[0057] The number of replications in the third direction of the unit cell is:

[0058] ceil(MaxRc / abs(LatticeVector[2]·(LatticeVector[0]×LatticeVector[1]) / |LatticeVector[0]×LatticeVector[1]|))

[0059] ceil rounds up, abs takes the absolute value, · is the dot product, and × is the cross product.

[0060] In particular, for periodic systems, the original unit cell is replicated at least once on each side of the periodic boundary.

[0061] Spheres are epitaxially formed at all sites of the original unit cell with the maximum cutoff radius. If the original unit cell already replicated in the periodic direction cannot cover the sphere, an integer number of original unit cells are replicated bidirectionally in the periodic direction until the sphere is completely covered.

[0062] The maximum cutoff radius is specified manually or calculated based on the system properties. The calculation involves randomly selecting n systems from the dataset, calculating the inter-atomic distances between each pair of atoms in the original unit cell of the system, sorting the calculated inter-atomic distances from smallest to largest, dividing the total number of inter-atomic distances by the set number of cutoff radii to obtain a sampling interval, sampling the sorted list according to the sampling interval to obtain a cutoff radius list, and using the maximum distance value in the cutoff radius list as the maximum cutoff radius.

[0063] (3.1.2) Describe the atomic environment of each atom within the original unit cell of the system.

[0064] (a) Identify the atomic environment of each atom.

[0065] The number of atomic environments for each atom is the same as the number of cutoff radii in the cutoff radius list. Each atomic environment in each atom refers to the atomic environment within a sphere formed with that atom as the center and a cutoff radius in the cutoff radius list as the radius.

[0066] (b) Construct the base point set of each atom in each atomic environment;

[0067] Within each atomic environment, based on the predetermined number of atoms constituting the polyhedron, all atoms within that environment are combined to form n or more polyhedra, where n ≥ 1. These n or more polyhedra are then sorted, and the top m polyhedra, where m ≤ n, form a set of m base points. The sorting rules are as follows:

[0068] First, sort the polyhedra from largest to smallest by volume or surface area. When two polyhedra have the same volume or surface area, sort them from largest to smallest by volume or surface area of ​​their inscribed or circumscribed spheres.

[0069] (c) Sort the atoms

[0070] (c1) Calculate the distance between each atom in each atomic environment and the atoms in each base point group (the calculated distances can be non-linearly transformed or not) and sort them (from largest to smallest or from smallest to largest) and save them;

[0071] (c2) Sort the atoms of the base point group: Calculate the distance from each atom in the base point group to the central atom in the atomic environment of that atom, and sort them (from largest to smallest or from smallest to largest); when the distances are the same, the atoms with the same distances are sorted according to the distance from the atom to the centroid, center or circumcenter in the system. If they are still the same, they are sorted according to the angle formed by the vector formed by the atom and the central atom to the principal axis, axis of symmetry or the maximum extension direction in the system. The maximum extension direction refers to the direction of the vector formed by the two atoms with the maximum distance between all pairs of atoms in the original unit cell of the system, and the vectors are saved.

[0072] (c3) If the system contains multiple elements, the atomic elements in the atomic environment are sorted: the element dimension is established in the constructed tensor, and the distance from each atom in the atomic environment to the atom in the base group is placed in the corresponding tensor index position according to the sorting of the element types of the atoms in the atomic environment; the sorting of element types is either in the order of the periodic table or in a set order.

[0073] (3.1.3) The description of the atomic environment is aggregated into the description of a single system.

[0074] (a) Connect the atomic environments of all atoms in the system in sequence, or construct a high-dimensional tensor by overlapping each atomic environment in a new dimension, or sum or multiply them to form a description of the atomic environment of each atom with multiple cutoff radii.

[0075] (b) The descriptions of the atomic environments with multiple cutoff radii for each atom are sequentially linked into a list according to the order of the atoms in the system, or each atomic environment is overlapped in a new dimension to construct a high-dimensional tensor, or summed, or multiplied to form the feature engineering output of the entire system.

[0076] (3.2) Input the atomic coordinates in the data stored in step 2) into the feature engineering to obtain feature engineering data. Perform data processing on the feature engineering data, including standardization, sparsification or densification, to obtain new feature engineering data.

[0077] The new feature engineering data in each system and the system-corresponding properties corresponding to the new feature engineering data after processing in step 2) are combined to form a new dataset for each system. Then, the new datasets of two or more systems are combined to form a final dataset.

[0078] The final dataset is grouped and preprocessed, including splitting into training, validation, and test sets, and then batch-processed to obtain a dataset that can be directly input into a machine learning model.

[0079] (3.3) Constructing a machine learning model

[0080] The machine learning model is a model with a multi-layered nested architecture, consisting of multiple single-atom models, and each single-atom model is composed of multiple element sub-networks.

[0081] Each of the aforementioned element subnetworks consists of a backbone network and a plurality of detection branches connected to the backbone network, each detection branch being a neural network.

[0082] The number of detection branches is determined by the number of corresponding properties in the system, that is, the number of detection branches is greater than or equal to the number of corresponding properties in the system.

[0083] The number of element subnetworks is determined by the number of elements in the system, that is, the number of element subnetworks is greater than or equal to the total number of elements in the system.

[0084] The single-atom model is determined by the number of atoms in the system with the largest number in the final dataset, that is, the number of single-atom models is greater than or equal to the total number of atoms in the largest system.

[0085] (3.4) The machine learning model is trained using the final dataset to obtain the machine learning potential surface model.

[0086] The machine learning model is trained and validated using the training and validation sets in the final dataset obtained in step (3.2), and tested using the test set in the final dataset, thus obtaining the machine learning potential surface model.

[0087] The following is an example:

[0088] Example 1

[0089] Embodiment 1 of the present invention provides a feature engineering construction method, such as... Figure 2 The following are included:

[0090] Based on the existing dataset, the cell is expanded according to its periodicity to cover all atomic epitaxial spheres with the maximum cutoff radius. The maximum cutoff radius is specified as...

[0091] Sampling 10% of the distances between all atoms in the system, a sequence from the maximum to the maximum cutoff radius is constructed, and the 20th, 40th, 60th, and 80th percentiles are used as parameters for the cutoff radius.

[0092] For each system in the dataset, perform the following operations:

[0093] The interatomic spacing of all atomic pairs in the system is calculated by calculating the nearest neighbor atomic table of the central unit cell based on the selected cutoff radius, with the distance taken as the Euclidean distance;

[0094] Based on the nearest neighbor atom table construction combination, the number of atoms in the base point group is 4, and the first 4 base point groups with the largest volume of the triangular pyramid constructed from the nearest neighbor atoms are selected.

[0095] For each central atom, read the distance from each nearest neighbor atom to each base point group, group them by element, connect them to form a high-dimensional tensor, and align them.

[0096] Continue constructing the distance tensor by pressing the next cutoff radius, and pad all the cutoff radius results with nan values ​​to the same length for all atoms, as the characteristic engineering output of the system.

[0097] All system calculation results are aligned with nan values ​​and used as the feature engineering results for this dataset.

[0098] Example 2

[0099] Embodiment 2 of the present invention provides a method for constructing a machine learning model, such as Figure 3a , Figure 3b , Figure 3c The following are included:

[0100] The backbone network is constructed, consisting of three convolutional layers and corresponding activation layers. After reorganization, the detection branch is connected, consisting of two convolutional layers and corresponding activation layers. The above constitutes the sub-network of elements.

[0101] Multiple equivalent element subnetworks are constructed based on the number of elements involved in the dataset, and the parameters are not shared between the subnetworks.

[0102] For multi-element systems, the output of the element sub-network is connected to the element identification layer, so that it outputs the corresponding element sub-network output according to the type of input element, thereby constructing a single-atom model;

[0103] The single-atom model is called multiple times based on the maximum number of atoms in the dataset, and the results are input into the atom number adaptation layer. The outputs of the single-atom model are then integrated as the final fitting output of the system.

[0104] Matters not covered in this invention are common knowledge.

[0105] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.

[0106] While the invention has been fully described through various embodiments, and while these embodiments have been described in considerable detail, the applicant does not intend to limit the scope of the appended claims or in any way to such details. Additional advantages and modifications will readily become apparent to those skilled in the art. Therefore, the invention is not, in its broader aspects, limited to the specific details, representative apparatuses and methods, and the illustrative examples shown and described. Consequently, deviations from such details may be made without departing from the spirit or scope of the overall inventive concept.

Claims

1. A method for constructing a machine learning potential surface model, characterized in that, Includes the following steps: 1) Collect data for training the model The data includes data from more than one system, and the data for each system includes: the atomic coordinates of the system and the corresponding properties of the system; wherein, The system properties described include one or both of the following: overall system properties and atomic-level properties; wherein, The aforementioned atom-by-atom properties include one or more of the following: atomic force, atomic magnetism, atomic charge, and atomic ionization energy; The overall properties of the system include one or more of the following: energy, cell stress, adsorption energy, and D-band centers; 2) Perform data verification and data storage on the collected data; 3) Construct a machine learning potential surface model, including: (3.1) Establish feature engineering; including: (3.1.1) The system is subjected to periodic processing, wherein the periodicity includes: one-dimensional periodicity, two-dimensional periodicity, and three-dimensional periodicity; the periodic processing of the system includes: For a periodic system, the number of times the original unit cell is replicated is calculated based on the given maximum cutoff radius, and the original unit cell is replicated in the periodic direction. The number of replications is the calculated number. The process of replicating the original unit cell in a periodic direction involves epitaxially forming a sphere at all points of the original unit cell with the maximum cutoff radius. If the original unit cell replicated in the periodic direction cannot cover the sphere, then an integer number of original unit cells are replicated bidirectionally in the periodic direction until the sphere is completely covered. The maximum cutoff radius is specified manually or calculated based on the system properties. The calculation involves randomly selecting n systems from the dataset, calculating the inter-atomic distances between any two atoms in the original unit cell of the system, sorting the calculated inter-atomic distances from smallest to largest, dividing the total number of inter-atomic distances by the set number of cutoff radii to obtain a sampling interval, sampling the sorted list according to the sampling interval to obtain a cutoff radius list, and taking the maximum distance value in the cutoff radius list as the maximum cutoff radius. For a periodic system, the original unit cell is replicated at least once on each side of the periodic boundary; (3.1.2) Describe the atomic environment of each atom within the original unit cell of the system; including: (a) Identify the atomic environment of each atom; (b) Construct the base point set of each atom in each atomic environment; (c) Sort the atoms; (3.1.3) The description of the atomic environment is aggregated into a description of a single system; including: (a) Connect the atomic environments of all atoms in the system in sequence, or construct a high-dimensional tensor by overlapping each atomic environment in a new dimension, or sum or multiply them to form a description of the atomic environment of each atom with multiple cutoff radii. (b) The descriptions of the atomic environments with multiple cutoff radii for each atom are sequentially linked into a list according to the order of the atoms in the system, or each atomic environment is overlapped in a new dimension to construct a high-dimensional tensor, or summed, or multiplied to form the feature engineering output of the entire system. (3.2) Through feature engineering, the final dataset used to train the model is obtained; (3.3) Constructing machine learning models; (3.4) The machine learning model is trained using the final dataset to obtain the machine learning potential surface model.

2. The method for constructing a machine learning potential surface model according to claim 1, characterized in that, The data verification described in step 2) includes two stages: data cleaning and data validation. The data cleaning includes removing invalid data, correcting erroneous data, and converting data formats; wherein, The removal of invalid data refers to removing invalid data where atomic coordinates or corresponding system properties are empty; the conversion of data format refers to converting the format used for data storage. The data verification mentioned above involves verifying the chemical structure and physicochemical properties, checking the rationality and consistency of the data, including using general physicochemical rules to perform preliminary verification of the data and excluding data that is obviously outside the normal range.

3. The method for constructing a machine learning potential surface model according to claim 1, characterized in that, The data storage described in step 2) involves organizing, saving, and managing the obtained data. The data is organized by saving it as a set data file or database according to the system type. The data file or database includes: xyz file, database file, SQL database, or MongoDB. The data files or databases mentioned are built locally or in the cloud, and the data is stored in a location accessible during training; Data storage should consider data deduplication and data backup to ensure the security and validity of existing data.

4. The method for constructing a machine learning potential surface model according to claim 1, characterized in that, The formula for calculating the number of times the original unit cell is replicated, as described in (3.1.1), is as follows: Let the unit cell parameter be LatticeVector, which is a 3x3 matrix; let the maximum cutoff radius be MaxRc. Then the number of replications in the three directions of the unit cell is as follows. The number of replications in the first direction of the unit cell is: ceil(MaxRc / abs(LatticeVector[0]·(LatticeVector[1]×LatticeVector[2]) / |LatticeVector[1]×LatticeVector[2]|)); The number of replications in the second direction of the unit cell is: ceil(MaxRc / abs(LatticeVector[1]·(LatticeVector[0]×LatticeVector[2]) / |LatticeVector[0]×LatticeVector[2]|)); The number of replications in the third direction of the unit cell is: ceil(MaxRc / abs(LatticeVector[2]·(LatticeVector[0]×LatticeVector[1]) / |LatticeVector[0]×LatticeVector[1]|)); ceil rounds up, abs takes the absolute value, · is the dot product, and × is the cross product.

5. The method for constructing a machine learning potential surface model according to claim 1, characterized in that, The description of the atomic environment of each atom within the original unit cell in (3.1.2) specifically includes: (a) Identify the atomic environment of each atom. The number of atomic environments for each atom is the same as the number of cutoff radii in the cutoff radius list. Each atomic environment in each atom refers to the atomic environment within a sphere formed with that atom as the center and a cutoff radius in the cutoff radius list as the radius. (b) Construct the base point set of each atom in each atomic environment; Within each atomic environment, based on the predetermined number of atoms constituting the polyhedron, all atoms within that environment are combined to form n or more polyhedra, where n ≥ 1. These n or more polyhedra are then sorted, and the top m polyhedra, where m ≤ n, form a set of m base points. The sorting rules are as follows: First, sort the polyhedra from largest to smallest by volume or surface area. When two polyhedra have the same volume or surface area, sort them from largest to smallest by volume or surface area of ​​their inscribed or circumscribed spheres. (c) Sort the atoms (c1) Calculate the distance between each atom in each atomic environment and the atoms in each base point group, sort them, and save them; (c2) Sort the atoms of the base point group: Calculate the distance from the atom in each base point group to the central atom in the atomic environment of the atom, and sort them; when the distances are the same, the atoms with the same distances are sorted according to the distance from the atom to the centroid, center or circumcenter in the system. If they are still the same, they are sorted according to the angle formed by the vector formed by the atom and the central atom to the principal axis, axis of symmetry or the maximum extension direction in the system. The maximum extension direction refers to the direction of the vector formed by the two atoms with the maximum distance between all pairs of atoms in the original unit cell of the system, and the vector is saved. (c3) If the system contains multiple elements, the atomic elements in the atomic environment are sorted: the element dimension is established in the constructed tensor, and the distance from each atom in the atomic environment to the atom in the base point group is placed in the corresponding tensor index position according to the sorting of the element types of the atoms in the atomic environment; the sorting of element types is either in the order of the periodic table or in a set order.

6. The method for constructing a machine learning potential surface model according to claim 1, characterized in that, Step 3) The final dataset for training the model is obtained through feature engineering as described in step (3.2). This involves inputting the atomic coordinates of the data stored in step 2) into the feature engineering to obtain feature engineering data. The feature engineering data is then processed, including standardization, sparsification, or densification, to obtain new feature engineering data. The new feature engineering data in each system and the system-corresponding properties corresponding to the new feature engineering data after processing in step 2) are combined to form a new dataset for each system. Then, the new datasets of two or more systems are combined to form a final dataset. The final dataset is grouped and preprocessed, including splitting into training, validation, and test sets, and then batch-processed to obtain a dataset that can be directly input into a machine learning model.

7. The method for constructing a machine learning potential surface model according to claim 1, characterized in that, Step 3) The machine learning model described in step (3.3) is a model with a multi-layered nested architecture, which is composed of a number of single-atom models, and each single-atom model is composed of a number of element sub-networks; Each of the aforementioned element subnetworks consists of a backbone network and a plurality of detection branches connected to the backbone network, each detection branch being a neural network. The number of detection branches is determined by the number of corresponding properties in the system, that is, the number of detection branches is greater than or equal to the number of corresponding properties in the system. The number of element subnetworks is determined by the number of elements in the system, that is, the number of element subnetworks is greater than or equal to the total number of elements in the system; Step 3) The construction of the machine learning potential energy surface model described in step (3.3) involves using the training set and validation set in the final dataset to train and validate the machine learning model, and using the test set in the final dataset to test the trained and validated machine learning model, ultimately obtaining the machine learning potential energy surface model.

Citation Information

Patent Citations

  • Matter structure description method applicable to machine learning potential energy surface construction

    CN108536998A

  • Biomacromolecule system database construction method and system based on transfer learning

    CN115171794A