Method for preparing machine learning potential crystal structure training data set in multi-source mode

The machine learning potential crystal structure training data set is prepared by multi-source method, combining Rose equations, experimental data, traditional potential functions and first-principle calculations, which solves the problem of high-cost and high-time training data generation, and realizes efficient and widely covered data set generation, which is suitable for accurate simulation of various material systems.

CN120375976APending Publication Date: 2025-07-25HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510331830.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the construction of training data sets for machine learning potential functions relies on first-principles calculations, which leads to high computational costs and long time, making it difficult to achieve standardization and automation, especially in a multivariate hybrid system, which is difficult to fully capture system characteristics.

Method used

The training data set is prepared by a multi-source method, including the Rose equation to generate data of a specified structure type, the experimental data that changes in pressure with volume based on the experimental data, the traditional potential function and first-principle calculation generate data of representative configurations, and the training data set is expanded and improved through first-principle expansion and improvement.

Benefits of technology

It significantly reduces the computational cost of training data generation and improves data generation efficiency. The generated data set covers a wide range of potential energy surfaces, can accurately describe atomic behavior under extreme conditions such as high temperature and high pressure, and is suitable for a variety of material systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375976A_ABST
    Figure CN120375976A_ABST
Patent Text Reader

Abstract

The invention relates to the field of material science, in particular to a method for preparing a machine learning potential crystal structure training data set in a multi-source mode, and the method comprises the steps: S1, generating first-class training data of a specified structure type of a crystal through employing a Rose equation; s2, based on the experimental data, generating second-class training data of the pressure of the crystal structure changing along with the volume; s3, a representative configuration of the crystal is generated by using a traditional potential function, and single-point energy calculation is performed through a first principle to obtain third-class training data; and S4, based on the training data obtained in the steps S1, S2 and S3, further adopting a first principle to expand and perfect the training data, and finally generating a crystal structure data set. The training data is prepared in a multi-source mode, the Rose equation, the experimental data and the traditional potential function are combined with the first principle calculation, and compared with a method completely depending on the first principle calculation, the calculation cost of training data generation is greatly reduced, and meanwhile the data generation efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of materials science, and particularly to a method for preparing a training data set of a machine learning potential crystal structure in a multi-source manner. Background Art

[0002] In the multi-scale simulation of materials science, computer simulation at the atomic scale is not only a key means to deeply understand the internal microscopic mechanism of materials, but also provides a necessary data basis for the construction of meso-scale and macro-continuum models, thus playing an important role in materials science research. In this field, molecular dynamics simulation has become a core atomic-scale calculation method because it can track the movement trajectories of atoms or molecules and reveal the dynamic behavior laws at the microscopic level of matter. However, the accuracy of molecular dynamics simulation highly depends on the accurate description of the interatomic interaction potential function, which directly determines the reliability of the simulation results.

[0003] Although traditional potential function models are constructed based on physical meanings and are calibrated and optimized by combining limited experimental data and first-principles calculation results, and are widely used in materials simulation, their limitations gradually emerge when dealing with many-body interactions and complex material systems. Especially when dealing with multi-component mixed systems, a single potential function model often has difficulty in comprehensively capturing the system characteristics. In addition, the optimization process of potential parameters depends on the professional experience of developers and the understanding of atomic simulation, making the development of potential functions highly personalized and difficult to achieve standardization and automation, thus limiting their application in materials simulation.

[0004] Compared with traditional potential functions, machine learning potential functions, with their complex mathematical models and high-dimensional parameter spaces, can learn and construct models that can effectively predict the interatomic interactions in unknown systems by training a large amount of known data (such as atomic positions, system energies, atomic forces, and box stresses, etc.). Such models can capture more complex interaction details, thus more accurately simulating the properties and behaviors of materials, and their importance in materials science is becoming increasingly prominent.

[0005] However, when constructing a machine learning potential function, the preparation of the training data set is crucial. As the basis for machine learning algorithms to learn the interatomic interaction laws, the comprehensiveness of the training data set directly determines the prediction performance of the machine learning potential function. Therefore, the training data set should cover structures corresponding to a wide range of potential energy surfaces and contain various different local atomic environments to ensure that the machine learning potential function can accurately simulate the material behavior under various scenarios.

[0006] Currently, the training data of machine learning potential functions rely on first-principles calculations, but this method has a huge computational cost, and constructing a training data set usually consumes high computational costs and a large amount of time. Summary of the Invention

[0007] To solve the above problems, the present invention proposes a method for preparing a training dataset of machine learning potential crystal structures in a multi-source manner, which can effectively reduce the computational cost of training data generation and improve the data generation efficiency at the same time.

[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0009] A method for preparing a training dataset of machine learning potential crystal structures in a multi-source manner, comprising the following steps:

[0010] S1: Generate the first type of training data of the specified crystal structure type of the crystal by using the Rose equation;

[0011] S2: Generate the second type of training data of the crystal structure pressure varying with volume based on experimental data;

[0012] S3: Generate representative configurations of the crystal by using traditional potential functions, and perform single-point energy calculations through first-principles to obtain the third type of training data;

[0013] S4: On the basis of the training data in steps S1, S2, and S3, use the first-principles method to expand and improve the training data, and finally generate a crystal structure dataset.

[0014] Further, in the step S1, the specified crystal structure type includes different crystal structures of pure elements, crystal structures with defects, and alloy structures.

[0015] Further, the step S1 includes:

[0016] S101: Determine the specified crystal structure type of the crystal for which the first type of training data needs to be generated;

[0017] S102: Use first-principles to optimize the structure of the specified crystal structure type in step S101 to obtain the parameters required for Rose equation calculation;

[0018] S103: Use the Rose equation to calculate the energy data at different structures and different volumes as the first type of training data.

[0019] Further, the step S2 includes:

[0020] S201: Based on the experimental data of the crystal pressure-volume relationship, keep the experimental pressure and volume ratio unchanged, replace the experimental equilibrium volume with the equilibrium volume calculated by first-principles, and then fit the polynomial equation of the pressure-volume relationship;

[0021] S202: Use the polynomial equation obtained in step S201 to generate configuration data of the pressure varying with volume as training data.

[0022] Furthermore, the step S3 includes:

[0023] S301: Using a traditional potential function to quickly generate the basic atomic configurations of the complete structure at different temperatures and pressures;

[0024] S302: For the basic atomic configurations generated in step S301, performing single-point energy calculations using the first-principles method to obtain the energy, atomic forces, and box stress data of the configurations;

[0025] S303: Using a traditional potential function to perform molecular dynamics simulations on the configurations with defects and liquid configurations under changing temperature and pressure conditions to generate the atomic configurations of the system under dynamic behavior;

[0026] S304: Screening out representative configurations from the atomic configurations generated in step S303;

[0027] S305: For the screened representative configurations, performing single-point energy calculations using the first-principles method to obtain training data.

[0028] Furthermore, the training data supplemented and improved in the step S4 includes:

[0029] Isolated atomic structures covering various elements;

[0030] Two-atom structures composed of the same or different types of atoms at different spacings;

[0031] Perturbed variants of pure element crystal structures and alloy structures.

[0032] The present invention also provides a crystal structure data set prepared by using the method as described above.

[0033] The present invention also provides the application of the crystal structure data set as described above in the training of machine learning potential functions.

[0034] The present invention has the following beneficial effects:

[0035] 1. The present invention prepares training data through a multi-source method, combining the Rose equation, experimental data, traditional potential functions, and first-principles calculations respectively. Compared with the method that completely relies on first-principles calculations, the computational cost of generating training data is significantly reduced, and the data generation efficiency is significantly improved at the same time.

[0036] 2. The present invention can quickly generate a high-quality data set covering a wide potential energy surface. The constructed machine learning potential function can accurately describe the atomic behavior under short-range interactions and complex defect conditions, and shows excellent prediction performance under extreme conditions such as high temperature and high pressure, and has strong extrapolation ability.

[0037] 3. The training data generated by the present invention covers pure element crystal structures, defect configurations, alloy structures, and extreme configurations such as isolated atoms and diatomic atoms, ensuring a wide coverage of the potential energy surface and applicability to various material systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0039] Figure 1 is the method flow for preparing a machine learning potential crystal structure training data set by a multi-source method of the present invention;

[0040] Figure 2 is the process for generating a training data set of different crystal structures of pure elements using the Rose equation of the present invention;

[0041] Figure 3 is the comparison of the training data of different crystal structures of tungsten generated by the Rose equation of the present invention and the DFT results;

[0042] Figure 4 is the process for obtaining a training data set of pure element crystal structures with different defect configurations using the Rose equation of the present invention;

[0043] Figure 5 is the comparison of the training data of BCC tungsten with different defect configurations generated by the Rose equation of the present invention and the DFT results;

[0044] Figure 6 is the process for obtaining a training data set of different alloy structures using the Rose equation of the present invention;

[0045] Figure 7 is the comparison of the training data of different tungsten-rhenium alloy structures generated by the Rose equation of the present invention and the DFT results;

[0046] Figure 8 is the process for generating a training data set of pressure varying with volume based on experimental data of the present invention;

[0047] Figure 9 is the comparison of the configuration training data of pressure varying with volume generated by fitting based on experimental data of the present invention and the experimental values;

[0048] Figure 10 is the comparison of the binding energy of vacancy clusters calculated by the tungsten potential function of the present invention and the DFT results;

[0049] Figure 11 is the comparison of the formation energy and binding energy of <111> interstitial clusters calculated by the tungsten potential function of the present invention and the DFT results;

[0050] Figure 12 Comparison of the formation energy and binding energy of <100> interstitial clusters calculated by the tungsten potential function of the present invention with the DFT results;

[0051] Figure 13 Formation energy of 1 / 2 <111> and <100> interstitial dislocation loops calculated by the tungsten potential function of the present invention;

[0052] Figure 14 Comparison of the surface energy calculated by the tungsten potential function of the present invention with the DFT results;

[0053] Figure 15 Comparison of the stacking fault energy calculated by the tungsten potential function of the present invention with the DFT results;

[0054] Figure 16 Comparison of the grain boundary energy calculated by the tungsten potential function of the present invention with the DFT results;

[0055] Figure 17 Relaxed grain boundary configuration obtained by calculating the tungsten potential function of the present invention;

[0056] Figure 18 Screw dislocation core structure calculated by the tungsten potential function of the present invention;

[0057] Figure 19 Comparison of the phonon spectrum calculated by the tungsten potential function of the present invention with the DFT and experimental results;

[0058] Figure 20 Comparison of the thermal expansion calculated by the tungsten potential function of the present invention with the experimental data;

[0059] Figure 21 Melting point predicted by the tungsten potential function of the present invention by the solid-liquid interface method;

[0060] Figure 22 Displacement threshold energy calculated by the tungsten potential function of the present invention. Detailed implementation mode

[0061] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments.

[0062] This embodiment provides a method for preparing a training data set of a machine learning potential crystal structure in a multi-source manner. The flowchart is as Figure 1 shown, and specifically includes the following steps:

[0063] S1: Generate training data of a specified structure type using the Rose equation;

[0064] Rose et al. [Phys. Rev. B 29 (1984) 2963] have demonstrated that for most materials, there is a universal relationship between the energy E per atom and the lattice constant a in their reference structures, which is specifically expressed as:

[0065] E(a * ) = -E0(1 + a * ) exp(-a * )

[0066] where,

[0067]

[0068] In the above formula, a * is a function of the lattice constant a, a0 represents the equilibrium lattice constant of the reference structure, V0 represents the volume of a single atom in the reference structure at the equilibrium lattice constant, E0 represents the negative value of the energy of a single atom in the reference structure, and B0 is the bulk modulus of the reference structure.

[0069] This study deeply analyzed the different crystal structures of pure elements, the crystal structures of pure elements containing point defects, surfaces, and grain boundaries, as well as the ordered alloy structures and disordered solid solution structures composed of different elements. The research results show that the system energy data obtained based on the above method are in good agreement with the first-principles DFT calculation results, and within a wide range of volumes, the deviation remains within a few meV / atom. This accuracy is comparable to the usual energy prediction accuracy of machine learning potentials (several meV / atom), indicating that the training data generated using this method have high accuracy and are suitable as a dataset for training machine learning potential functions.

[0070] Therefore, by applying the Rose equation and combining first-principles calculations to determine the parameters a0, V0, E0, and B0 in the Rose equation, it is possible to quickly obtain the data of energy varying with lattice parameters under different structures. These structural energy data are an important part of the machine learning potential training dataset, thus effectively reducing the workload of first-principles calculations. The following will elaborate in detail the specific steps of generating training data using this method under different structures such as the different crystal structures of pure elements, the crystal structures of pure elements containing point defects, surfaces, and grain boundaries, as well as the ordered alloy structures and disordered solid solution structures composed of different elements.

[0071] Specifically, in step S1, the specified structure types include different crystal structures of pure elements, defective crystal structures, and alloy structures.

[0072] (1) Different crystal structures of pure elements

[0073] For the pure elements in the constructed potential energy function system, a training data set of the energy-volume relationship of typical structures is generated. The typical structures include body-centered cubic (BCC), face-centered cubic (FCC), hexagonal close-packed (HCP), simple cubic (SC), diamond structure (Diamond), as well as A15 and C15 structures. In this process, only the equilibrium lattice constant, bulk modulus, equilibrium structure single-atom volume, and energy of these structures need to be calculated using first-principles, and then these parameters are substituted into the Rose equation to generate the energy data of these structures at different volumes. The calculation process is as Figure 2 shown.

[0074] First, prepare the primitive cells of body-centered cubic structure (BCC), face-centered cubic structure (FCC), hexagonal close-packed structure (HCP), simple cubic structure (SC), diamond structure (Diamond), A15 structure, and C15 structure. Use the first-principles calculation software VASP to optimize the structures of these structures to obtain the equilibrium lattice constant a0, the volume V0 of a single atom, and the energy E0. For the configurations after structure optimization, continue to use VASP to calculate the elastic constants to obtain the bulk modulus B0. Then substitute these four obtained parameters into the Rose equation, and calculate the energy data of different structures at different volumes through the Rose equation. This process is completed by executing a Fortran program. Figure 3 Shows the comparison of the calculation results of different structures of tungsten (W) using the Rose equation and the first-principles method. Among them, the red curve represents the results calculated using the Rose equation, and the dots represent the results of first-principles calculations. The abscissa represents the volume change (expressed as V / V0, where V0 is the equilibrium volume), and the ordinate is the system energy (expressed as the energy of a single atom corresponding to this structure). For different structures, the calculation results of the Rose equation are very close to the results of first-principles DFT calculations.

[0075] (2) Crystal structures with defects

[0076] Secondly, for the defect-containing systems of the ground-state structures of pure elements in the constructed potential energy function system, such as systems containing defects such as vacancies, self-interstitial atoms, surfaces, grain boundaries, etc., as well as defect configurations containing other elements in the system, such as substitutional solute atoms, interstitial solute atoms, solute atom-vacancy or self-interstitial complexes and other defect structure configurations, the Rose equation can be used to generate the data of their energy-volume relationship. The calculation process is as Figure 4As shown below. First, prepare the configurations of structures with defects, and use the first-principles calculation software VASP to optimize the structures of these configurations. According to the optimized volume and energy, obtain the equilibrium lattice constant a0, the volume V0 of a single atom, and the energy E0. On the basis of the configurations after structure optimization, change the length of the box in a certain direction while keeping the ratio of each direction of the box and the relative positions of the atoms unchanged, and use VASP to perform single-point energy calculations on these deformed configurations. Obtain the bulk modulus B0 by fitting the energy-volume curve. Then substitute these four parameters into the Rose equation, and obtain the energy data at different structures and different volumes through the Rose equation. This process is completed by executing a Fortran program. By obtaining the parameters required by the Rose equation through a small amount of first-principles calculations, a large amount of configuration data of structures with defects can be generated. Figure 5 Figure 5 shows a comparison of the calculation results of defect structures including vacancies, divacancies, self-interstitial atoms, surfaces, and grain boundaries in the pure W ground-state BCC structure using the Rose equation and the first-principles method. Among them, the red curve represents the results calculated using the Rose equation, and the dots represent the results of first-principles calculations. For different types of defect structures, the calculation results of the Rose equation are very close to the results of first-principles calculations.

[0077] (3) Alloy structure

[0078] In addition, for alloy structures, such as disordered solid solutions or ordered-phase alloy structures composed of different elements in the constructed potential function system, the relationship data between energy and volume can also be generated through the Rose equation. Its calculation process is as Figure 6 shown, which shows how to obtain the parameters required by the Rose equation through a small amount of first-principles calculations, and then generate a large amount of configuration energy data of alloy structures including multi-component systems.

[0079] The feasibility of this method has been verified in the W-Re system. As Figure 7 shown, this figure shows a comparison of the energy data of the W-Re disordered solid solution, B2 virtual ordered phase, and W-Re precipitated phases σ and χ phase structure configurations calculated using the Rose equation with the DFT calculation results. In the figure, the red curve represents the results calculated using the Rose equation, and the dots represent the DFT-based calculation results. It can be found through comparison that for the above W-Re alloy structures, the calculation results of the Rose equation are very close to the DFT calculation results, which further proves the accuracy and reliability of the Rose equation in predicting the energy of alloy structures. Generally speaking, by combining the Rose equation with first-principles calculation parameters, a large amount of energy data of different crystal structures of pure elements, defect configuration structures, and different alloy structures can be efficiently generated, and these data can be used as part of the machine learning potential training dataset.

[0080] S2: Generate training data of pressure varying with volume based on experimental data;

[0081] In addition to generating training data based on the empirical Rose equation mentioned above, for experimental data, training data is generated using the experimental pressure and volume relationship data. The specific processing method is as follows: keeping the experimental pressure (P) and volume ratio (V / V0) unchanged, replacing the experimental equilibrium volume (V0) with the equilibrium volume calculated by first-principles, then fitting the relationship between pressure and volume, and generating configuration data of pressure varying with volume using the obtained polynomial equation. The calculation process is as Figure 8 shown. This method was verified, Figure 9 and a comparison between the calculation results of this method for pure W and experimental data was given. The red curve in the figure represents the results calculated by this method, while the dots represent the experimental data. Through this method, a large amount of configuration data of pressure varying with volume can be obtained using limited P-V experimental data, while ensuring that the equilibrium volume is consistent with the equilibrium volume calculated by first-principles. For cubic structures, the stress in three axial directions is equal to the pressure, so the stress data corresponding to the configuration can be obtained as training data.

[0082] S3: Generate representative configurations using traditional potential functions, perform single-point energy calculations by first-principles, and obtain training data;

[0083] For the constructed potential function system, a series of possible basic atomic configurations are quickly generated using existing traditional potential functions based on years of experience and theoretical accumulation, such as the configurations of a complete structure at different temperatures and pressures. Subsequently, the first-principles method is used to perform single-point energy calculations on these configurations to obtain their energy, atomic force, and simulation box stress data. To broaden the coverage of the potential energy surface, molecular dynamics simulations are performed on defect-containing configurations and liquid configurations under changing temperature and pressure conditions using traditional potential functions, and representative configurations are screened out to capture the dynamic behavior of the system under different physical conditions. Then, for these selected configurations, single-point energy calculations are performed using the first-principles method to obtain configuration energy and atomic force data. Compared with directly using the first-principles molecular dynamics method, this strategy not only reduces the calculation cost but also ensures the comprehensiveness and diversity of the potential energy surface configurations, enabling the training set to include more configurations under extreme conditions, such as crystal structures at high temperature and pressure, defect diffusion transition state structures, solid-liquid interfaces, and liquid structures.

[0084] S4: Based on the training data obtained in steps S1, S2, and S3, further expand and improve the training data using first-principles, and obtain a training data set;

[0085] For the constructed potential energy function system, on the basis of combining the Rose equation, experimental data, and traditional empirical potential energy functions to generate rich training data as described above, the first-principles calculation method is further used to expand and improve the training data set. First, the isolated atomic structures covering various elements are included to accurately describe the cohesive energy of the elements. Second, the diatomic (Dimer) and triatomic structures composed of the same or different types of atoms at different spacings are introduced to reveal the basic mechanism of atomic bonding, especially focusing on the atomic interactions at extremely short distances. In addition, the perturbed variants of pure element crystal structures and alloy structures are included, where the perturbation mainly refers to the perturbation of the unit cell and atomic positions, to enhance the model's ability to accurately describe the elastic constants.

[0086] S5: Using the training data set obtained in step S4, perform machine learning potential energy function training.

[0087] According to the above method, taking tungsten as an example, we prepared a training data set for tungsten, then used the moment tensor machine learning potential energy function model for training, and conducted a detailed inspection of the obtained potential energy function.

[0088] The training data set includes:

[0089] 1) Energy data of BCC structure, FCC structure, HCP structure, SC structure, Diamond structure, A15 structure, and C15 structure at different volumes. Each crystal structure contains 200 configurations, and a total of 1400 configurations of energy data are generated based on the Rose equation.

[0090] 2) Energy data of BCC structure with point defects (including vacancy, divacancy, <111> dumbbell, <111> crowdion, <110> dumbbell, <100> dumbbell, octahedral interstitial, tetrahedral interstitial, and double interstitial), BCC structure with surface (including (100), (110), (111), (112), (210), (321)), and BCC structure with grain boundary (including ∑3(110), ∑3(112), ∑5(012), ∑5(013)). Each crystal structure contains 200 configurations, and a total of 3800 configurations of energy data are generated based on the Rose equation.

[0091] 3) Energy data and atomic force data for the configurations (500 configurations) of the complete BCC structure at different temperatures and pressures, the configurations (50 configurations) of the BCC structure with vacancies and self-interstitials at different temperatures and pressures, and the solid-liquid interface and liquid configurations (50 configurations). These configurations were generated by running molecular dynamics simulations using existing traditional potential functions [J. Nucl. Mater. 502 (2018) 141 - 153], and the generated configurations were then used to obtain energy and force data through first-principles single-point energy calculations;

[0092] 4) BCC structure P-V data, including box stress data for 500 different configurations, generated by combining first-principles calculation parameters based on experimental P-V relationships;

[0093] 5) First-principles calculation data, including the energy and force of isolated atoms, the energy and force of the dimer structure (the distance between two atoms ranges from 1.2 to 2.2, including 10 configurations), and the energy, force, and stress of the BCC perturbed structure (including 1000 configurations).

[0094] The key parameter settings in the involved first-principles static calculations are as follows: the cutoff energy parameter is 500 eV, the KSPACING parameter is 0.15, and the energy convergence accuracy parameter is 10 -5 eV. The machine learning potential function model uses the Momenttensorpotential (MTP) potential model [Mach. Learn.: Sci. Technol. 2 (2021) 025002], the model level is lev18, and the machine learning potential function of tungsten is trained for the above constructed training dataset.

[0095] The properties such as lattice constant, bulk modulus, elastic constants, formation energy and migration energy of point defects, binding energy of vacancy clusters and interstitial clusters, relative stability of interstitial dislocation loops, surface energy, stacking fault energy, grain boundary energy, screw dislocation core structure, phonon spectrum, thermal expansion, melting point, displacement threshold energy, etc. were comprehensively detected using the trained potential function of tungsten. As follows: Table 1 shows the comparison of the lattice constant, cohesive energy, bulk modulus, and elastic constants calculated by the developed machine learning tungsten potential function with experimental or DFT results, and the calculated values of the potential function are in good agreement with the experimental or DFT results. Table 2 shows the comparison of the vacancy formation energy and migration energy, binding energy of divacancies, and formation energy of self-interstitial atoms in different configurations calculated by the developed machine learning tungsten potential function with DFT results, and the calculated values of the current potential function are in good agreement with the DFT results.

[0096] Table 1 Comparison of the lattice constant, cohesive energy, bulk modulus, and elastic constants calculated by the tungsten potential function with experimental or DFT results

[0097]

[0098] [1] J. et al., Physical Review B 100 (2019) 144105.

[0099] [2] C. Kittel, Introduction to Solid State Physics, eighth ed., Wiley, New York, 2005.

[0100] [3] G. Simmons, et al., Single Crystal Elastic Constants and Calculated Aggregate Properties: a Handbook, second ed., The MIT Press, Cambridge, MA, 1971.

[0101] Table 2, Comparison of Point Defect Properties Calculated by Tungsten Potential Function with DFT Results

[0102]

[0103]

[0104] [1] P. W. Ma, et al., Physical Review Materials 3 (2019) 013605.

[0105] [2] P. W. Ma, et al., Physical Review Materials 3 (2019) 063601.

[0106] [3] D. R. Mason, et al., Journal of Physics: Condensed Matter 29 (2017) 505501.

[0107] Figure 10 Comparison of the binding energy of vacancy clusters calculated by the developed machine learning tungsten potential function with DFT results is given. The size of the vacancy clusters ranges from containing two vacancy atoms to containing ten vacancy atoms. It can be seen from the figure that the prediction results of the current potential function are in good agreement with the DFT results.

[0108] Figure 11 Comparison of the formation energy and binding energy of <111> interstitial clusters calculated by the developed machine learning tungsten potential function with DFT results is given. The size of the interstitial clusters ranges from containing two <111> configuration interstitial atoms to containing twelve <111> configuration interstitial atoms. It can be seen from the figure that the prediction results of the current potential function are in good agreement with the DFT results.

[0109] Figure 12 The comparison between the formation energies and binding energies of <100> interstitial clusters calculated by the developed machine learning tungsten potential function and the DFT results is presented. The size of the interstitial clusters ranges from containing two <100> configuration interstitial atoms to containing twelve <100> configuration interstitial atoms. It can be seen from the figure that the predicted results of the current potential function are in good agreement with the DFT results.

[0110] Figure 13 The formation energies of 1 / 2<111> interstitial dislocation loops and <100> interstitial dislocation loops calculated by the developed machine learning tungsten potential function are presented. The maximum size of the interstitial dislocation loops reaches 300 self-interstitial atoms. It can be seen from the figure that the formation energy of the 1 / 2<111> interstitial dislocation loop is lower than that of the <100> interstitial dislocation loop, indicating that 1 / 2<111> is more stable than the <100> interstitial dislocation loop at low temperatures, which is in agreement with the phenomena observed in irradiation experiments.

[0111] Figure 14 The comparison between the formation energies of different surfaces calculated by the developed machine learning tungsten potential function and the DFT results is presented. The calculated surfaces include 10 different typical surfaces such as (100), (110), and (111). It can be seen from the figure that the predicted results of the current potential function are in good agreement with the DFT results.

[0112] Figure 15 The comparison between the stacking fault energy curves calculated by the developed machine learning tungsten potential function and the DFT results is presented. It can be seen from the figure that for the BCC crystal moving along the {110}<111> and {112}<111> slip systems, the predicted results of the current potential function are in good agreement with the DFT results.

[0113] Figure 16 The comparison between the different grain boundary energies calculated by the developed machine learning tungsten potential function and the DFT results is presented. The calculated grain boundaries include four different grain boundaries of ∑3(112), Σ3(112), Σ3(112), and Σ3(112). The relaxed grain boundary configurations are as Figure 17 shown. From Figure 16 it can be seen that the grain boundary energies calculated by the current potential function are in good agreement with the DFT results.

[0114] Figure 18 The 1 / 2<111> screw dislocation core structure predicted by the developed machine learning tungsten potential function is presented. The current potential function calculates a non-degenerate (compact) structure, which is consistent with the DFT prediction results.

[0115] Figure 19 The comparison between the phonon spectrum curves calculated by the developed machine learning tungsten potential function and the experimental and DFT results is presented. It can be seen from the figure that the phonon spectrum curves predicted by the current potential function are in good agreement with the experimental and DFT results.

[0116] Figure 20 The comparison between the thermal expansion curves calculated by the developed machine learning tungsten potential function and the experimental data is given. It can be seen from the figure that the prediction results of the current potential function are in good agreement with the experimental results.

[0117] Figure 21 The calculation results of the melting process of tungsten by the developed machine learning tungsten potential function using the solid-liquid interface method are given. The melting point predicted by the current potential function is within 3525 - 3575 K, which is close to the experimental value of 3695 K.

[0118] Figure 22 The displacement threshold energies in different directions calculated by the developed machine learning tungsten potential function are given. The minimum displacement threshold energy predicted by the current potential function is 44 eV, which is in good agreement with the DFT result of 42 eV, and the average displacement threshold energy predicted is 93 eV, which is in good agreement with the recommended value of 90 eV by the American Society for Testing and Materials (ASTM).

[0119] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A method for preparing a training data set of a machine learning potential crystal structure in a multi-source manner, characterized in that, It includes the following steps: S1: Generate the first type of training data of the specified crystal structure type using the Rose equation; S2: Generate the second type of training data of the crystal structure pressure varying with volume based on experimental data; S3: Generate the representative configurations of the crystal using traditional potential functions and perform single-point energy calculations through first principles to obtain the third type of training data; S4: Based on the training data obtained in steps S1, S2, and S3, further expand and improve the training data using first principles, and finally generate the crystal structure dataset.

2. The method for preparing a machine learning potential crystal structure training data set by a multi-source method according to claim 1, wherein, In step S1, the specified crystal structure types include different crystal structures of pure elements, crystal structures with defects, and alloy structures.

3. The method for preparing a training data set of a machine learning potential crystal structure by a multi-source method according to claim 2, wherein Step S1 includes: S101: Determine the specified crystal structure types for which the first type of training data needs to be generated; S102: Use first principles to optimize the structure of the specified crystal structure types in step S101 to obtain the parameters required for Rose equation calculation; S103: Use the Rose equation to calculate the energy data at different structures and volumes as the first type of training data.

4. The method for preparing a training data set of a machine learning potential crystal structure by a multi-source method according to claim 1, wherein Step S2 includes: S201: Based on the experimental data of the crystal pressure-volume relationship, keeping the experimental pressure and volume ratio unchanged, replace the experimental equilibrium volume with the equilibrium volume calculated by first principles, and then fit the polynomial equation of the pressure-volume relationship; S202: Use the polynomial equation obtained in step S201 to generate the configuration data of the pressure varying with volume as the training data.

5. The method for preparing a machine learning potential crystal structure training data set by a multi-source method according to claim 1, wherein, Step S3 includes: S301: Use traditional potential functions to quickly generate the basic atomic configurations of the complete structure at different temperatures and pressures; S302: For the basic atomic configurations generated in step S301, perform single-point energy calculations using first principles methods to obtain the energy, atomic forces, and box stress data of the configurations; S303: Use traditional potential functions to perform molecular dynamics simulations on the defective configurations and liquid configurations under the conditions of temperature and pressure changes to generate the atomic configurations of the system under dynamic behavior; S304: Screen out the representative configurations from the atomic configurations generated in step S303; S305: For the selected representative configurations, perform single-point energy calculations using first principles methods to obtain the training data.

6. The method for preparing a training data set of a machine learning potential crystal structure in a multi-source manner according to claim 1, wherein, The training data supplemented and improved in step S4 includes: Isolated atomic structures covering various elements; Diatomic structures composed of the same or different types of atoms at different spacings; Perturbed variants of pure element crystal structures and alloy structures.

7. A crystal structure dataset prepared by using the method according to any one of claims 1-6.

8. The application of the crystal structure dataset according to claim 7 in the training of machine learning potential functions.