Nanometer material performance prediction method and system based on deep learning

By using a deep learning-based method and pre-training models and density functional theory to construct an energy landscape model, the accuracy and efficiency problems of nanomaterial performance prediction were solved, and efficient phase change path and kinetic characteristics analysis under different external field conditions was achieved.

CN120613048APending Publication Date: 2025-09-09SHENZHEN ZHUOTAN NEW MATERIALS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510694444.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing methods are inefficient and difficult to accurately model when capturing the complex relationship between the atomic-level structure and properties of nanomaterials, resulting in insufficient accuracy in performance predictions.

Method used

A deep learning-based method is used to obtain the structural data of nanomaterials, perform configuration space analysis using a pre-trained structural performance prediction model, combine the Monte Carlo method to screen candidate sampling points, use density functional theory to construct an energy landscape model, and optimize the model through transfer learning to achieve the prediction of the phase transition path and kinetic characteristics of nanomaterials.

Benefits of technology

It improves the accuracy and efficiency of nanomaterial performance prediction, can accurately capture the structure-performance relationship under different external field conditions, and realize efficient phase change path and kinetic characteristics analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613048A_ABST
    Figure CN120613048A_ABST
Patent Text Reader

Abstract

The invention discloses a nanometer material performance prediction method and system based on deep learning, and the method comprises the steps: obtaining the structure data of a nanometer material, inputting the structure data into a pre-trained structure performance prediction model, and obtaining the configuration space and performance prediction value of the nanometer material; sampling the configuration space by adopting a Monte Carlo method, and if the performance predicted value exceeds a preset threshold value, determining the configuration corresponding to the performance predicted value exceeding the preset threshold value as a candidate sampling point to obtain a candidate sampling point set; calculating an energy value of each sampling point in the candidate sampling point set by adopting a density functional theory, and constructing an energy landscape model according to the energy value of each sampling point; and obtaining energy change data from the energy landscape model, and inputting the energy change data into the optimized structural performance prediction model to obtain a phase change path and a dynamic characteristic prediction result of the nano material. The accuracy and efficiency of performance prediction of the nano material can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of nanomaterial technology, and in particular to a method and system for predicting nanomaterial properties based on deep learning. Background Art

[0002] Nanomaterials, due to their unique physical and chemical properties, have revolutionary implications for fields such as energy, catalysis, and biomedicine. Predicting their performance directly determines the efficiency and application potential of material design. However, existing methods often rely on expensive experimental testing or computational simulations to capture the complex relationship between the atomic-level structure and properties of nanomaterials. This is inefficient and difficult to cope with the vast complexity of the configuration space. Furthermore, the mapping between the atomic-level structure of nanomaterials and their properties is highly nonlinear, making it difficult for traditional methods to effectively model this complexity, resulting in inaccurate performance predictions. Summary of the Invention

[0003] In order to solve the above technical problems, the present application provides a method and system for predicting nanomaterial properties based on deep learning, which can improve the accuracy and efficiency of nanomaterial property prediction.

[0004] This application provides a method for predicting nanomaterial properties based on deep learning, including:

[0005] Acquiring structural data of the nanomaterial and inputting the structural data into a pre-trained structural performance prediction model to obtain a configuration space and performance prediction value of the nanomaterial;

[0006] The configuration space is sampled using a Monte Carlo method. If a performance prediction value exceeds a preset threshold, the configuration corresponding to the performance prediction value exceeding the preset threshold is determined as a candidate sampling point to obtain a set of candidate sampling points.

[0007] Density functional theory is used to calculate the energy value of each sampling point in the set of candidate sampling points, and an energy landscape model is constructed according to the energy value of each sampling point, wherein the energy landscape model records the energy change and atomic configuration characteristics of each sampling point under external field conditions;

[0008] Energy change data is obtained from the energy landscape model, and the energy change data is input into the optimized structural performance prediction model to obtain the phase change path and kinetic characteristics prediction results of the nanomaterial, wherein the pre-trained structural performance prediction model is optimized using a transfer learning method to obtain the optimized structural performance prediction model.

[0009] In some embodiments, the method further includes: training a structural performance prediction model, wherein the training structural performance prediction model includes:

[0010] Acquiring atomic-level structural data from a nanomaterial database, generating an initial configuration set using molecular dynamics simulation, generating a configuration space under corresponding external field conditions based on the initial configuration set, and obtaining a configuration space data set based on the configuration space;

[0011] The configuration space data set is divided into a training set and a validation set, a deep learning algorithm is used to construct a structure-performance mapping model, and a nonlinear activation function is used to process the nonlinear characteristics of the performance mapping relationship to obtain an initial neural network structure;

[0012] The training set is iteratively trained using an optimization algorithm combined with a loss function. If the prediction error of the training set is lower than the preset threshold, the optimized neural network parameters are obtained.

[0013] The configuration parameters of the verification set are back-propagated through the optimized neural network parameters to obtain a structural performance prediction model.

[0014] In some embodiments, the step of calculating the energy value of each sampling point in the set of candidate sampling points using density functional theory and constructing an energy landscape model based on the energy value of each sampling point includes:

[0015] Density functional theory is used to calculate the energy value of each sampling point in the candidate sampling point set, and an initial energy landscape model is constructed according to the energy value of each sampling point;

[0016] Obtaining energy distribution and configuration space feature data from the initial energy landscape model; if the energy value is lower than the average value and the configuration feature meets the structural parameters of the high-performance region, determining the configuration corresponding to the energy value lower than the average value and the configuration feature meeting the structural parameters of the high-performance region as a priority sampling region, thereby obtaining a priority sampling region data set;

[0017] The energy distribution and configuration characteristics are obtained from the data set of the priority sampling area. The Bayesian optimization algorithm is used to update the sampling point selection strategy based on the energy distribution and configuration characteristics to generate the first candidate sampling point and obtain the optimized sampling point set.

[0018] The minimum energy path algorithm is used to analyze the energy conversion paths between the sampling points in the optimized sampling set. If the path energy difference is lower than a preset threshold, it is determined to be a stable phase change path, and a phase change path set is obtained.

[0019] Based on the phase transition path set, the dynamic Monte Carlo method is used to simulate the atomic motion trajectory of nanomaterials under external field conditions, calculate the time evolution characteristics of the path, and obtain the dynamic characteristic data;

[0020] The time evolution characteristics are obtained from the dynamic characteristic data, combined with the phase change path set, and the sampling point selection strategy is iteratively updated using the Bayesian optimization algorithm to generate the second candidate sampling points and obtain the optimized sampling point set again.

[0021] According to the re-optimized sampling point set and the newly added external field condition data, the initial energy landscape model is updated using density functional theory to obtain the final energy landscape model.

[0022] In some embodiments, the energy distribution and configuration characteristics are obtained from the data set of the priority sampling area, and a Bayesian optimization algorithm is used to update the sampling point selection strategy based on the energy distribution and configuration characteristics to generate a first candidate sampling point to obtain an optimized sampling point set, including:

[0023] Acquiring a data set from the priority sampling area, dividing the priority sampling area into a plurality of sub-areas using a grid partitioning method, and calculating energy value distribution and geometric topological characteristics for each sub-area to obtain energy distribution and configuration characteristics;

[0024] generating a priori probability distribution based on the energy distribution, assigning a predetermined probability value if the energy value is higher than a preset threshold, and normalizing the priori probability distribution using a probability density function to obtain the initial probability distribution;

[0025] The initial probability distribution is updated by using a Bayesian optimization algorithm and configuration features, the sampling point positions are adjusted according to the geometric constraints in the configuration features, and an updated probability distribution is obtained by iterative optimization using a Gaussian process regression algorithm;

[0026] The updated probability distribution is sampled using a Monte Carlo method to generate a set of candidate sampling points. The candidate sampling point set is screened, and sampling points that satisfy the configuration characteristics and the energy distribution constraints are retained to obtain an optimized sampling point set.

[0027] In some embodiments, a minimum energy path algorithm is used to analyze the energy conversion path between the optimized sampling points. If the path energy difference is lower than a preset threshold, it is determined to be a stable phase change path, and a phase change path set is obtained, including:

[0028] Obtaining energy change data from the initial energy landscape model, dividing the energy landscape into a plurality of sub-regions using a grid partitioning method, and calculating an energy value sequence for each sub-region to obtain an energy change data set;

[0029] calculating, based on the energy change dataset, a minimum energy path algorithm for calculating conversion paths between the first candidate sampling points, extracting energy values ​​of a start point and an end point for each path, calculating a path energy difference, and obtaining a path energy difference set;

[0030] For the path energy difference set, if the path energy difference is lower than a preset threshold, it is determined to be a stable phase change path, and the stable phase change paths are grouped using a clustering method to obtain the preliminary phase change path set;

[0031] The preliminary phase change path set is processed by a topological analysis method, the geometric constraint characteristics of each path are extracted, and the paths that meet the phase change characteristics are screened to obtain the phase change path set.

[0032] In some embodiments, the method of obtaining the time evolution characteristics from the dynamic characteristic data, combining the phase change path set, and iteratively updating the sampling point selection strategy using a Bayesian optimization algorithm to generate a second candidate sampling point to obtain a further optimized sampling point set includes:

[0033] Acquiring time evolution features from the dynamic characteristic data using a feature extraction method, calculating the autocorrelation function and spectrum analysis of the time series in the time evolution features, and obtaining a first feature set;

[0034] Constructing a phase change path set based on the first feature set; if the state transition frequency in the first feature set is higher than a preset threshold, generating a phase change path set using a clustering algorithm; otherwise, generating a first path set using an interpolation method to obtain a phase change path set;

[0035] The phase change path set is processed using a Bayesian optimization algorithm, and based on a pre-established probability distribution model, a posterior probability distribution is calculated through Gaussian process regression to generate a second set of candidate sampling points;

[0036] Sampling points are screened from the second candidate sampling point set by using the path similarity metric. If the path similarity metric value of the sampling point is greater than a preset threshold, the sampling point is retained to obtain a re-optimized sampling point set.

[0037] In some embodiments, obtaining energy change data from the energy landscape model and inputting the energy change data into an optimized structural performance prediction model to obtain a prediction result of the phase transition path and kinetic characteristics of the nanomaterial includes:

[0038] receiving initial energy distribution data obtained from an energy landscape database, and refining the initial energy distribution data using a clustering algorithm to obtain a refined energy landscape dataset;

[0039] Adopting a transfer learning method to adjust the pre-trained structural performance prediction model to construct an optimized structural performance prediction model;

[0040] Processing the refined energy landscape dataset through the optimized structural performance prediction model, outputting a node sequence of a phase transition path and a corresponding final energy value, to obtain a phase transition path dataset;

[0041] For the phase change path data set, a molecular dynamics simulation method is used to calculate the motion trajectory and rate of each node to obtain the phase change path and kinetic characteristics prediction results of the nanomaterial.

[0042] In some embodiments, the present application further provides a nanomaterial performance prediction system based on deep learning, comprising:

[0043] A data acquisition module is used to obtain structural data of the nanomaterial and input the structural data into a pre-trained structural performance prediction model to obtain the configuration space and performance prediction value of the nanomaterial;

[0044] a sampling module, configured to perform sampling processing on the configuration space using a Monte Carlo method, and if a performance prediction value exceeds a preset threshold, determine the configuration corresponding to the performance prediction value exceeding the preset threshold as a candidate sampling point, thereby obtaining a set of candidate sampling points;

[0045] an energy landscape model construction module, configured to calculate the energy value of each sampling point in the candidate sampling point set using density functional theory, and construct an energy landscape model based on the energy value of each sampling point, wherein the energy landscape model records the energy changes and atomic configuration characteristics of each sampling point under external field conditions;

[0046] A prediction result module is used to obtain energy change data from the energy landscape model, input the energy change data into the optimized structural performance prediction model, and obtain the phase change path and kinetic characteristics prediction results of the nanomaterial, wherein the pre-trained structural performance prediction model is optimized using a transfer learning method to obtain the optimized structural performance prediction model.

[0047] In some embodiments, the present application also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for predicting nanomaterial properties based on deep learning as described above is implemented.

[0048] In some embodiments, the present application also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned deep learning-based nanomaterial performance prediction methods.

[0049] Compared with the existing technology, the present application has the following beneficial effects: the present application discloses a method and system for predicting the performance of nanomaterials based on deep learning, which obtains the structural data of the nanomaterial and inputs the structural data into a pre-trained structural performance prediction model to obtain the configuration space and performance prediction value of the nanomaterial; the Monte Carlo method is used to sample the configuration space, and if the performance prediction value exceeds a preset threshold, the configuration corresponding to the performance prediction value exceeding the preset threshold is determined as a candidate sampling point to obtain a set of candidate sampling points; the density functional theory is used to calculate the energy value of each sampling point in the set of candidate sampling points, and an energy landscape model is constructed based on the energy value of each sampling point; energy change data is obtained from the energy landscape model, and the energy change data is input into the optimized structural performance prediction model to obtain the phase transition path and dynamic characteristics prediction results of the atoms in the nanomaterial. The method of the above embodiment realizes the performance prediction, configuration optimization and phase transition dynamics analysis of nanomaterials under different external field conditions, thereby improving the accuracy and efficiency of nanomaterial performance prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flow chart of a method for predicting nanomaterial properties based on deep learning provided in the first embodiment of the present application;

[0051] Figure 2 This is a schematic diagram of the structure of a nanomaterial performance prediction system based on deep learning provided in the second embodiment of the present application. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0053] Under variable external field conditions (such as temperature and pressure), accurately predicting the phase transition path and kinetic properties of nanomaterials has become a key problem. Due to the inability to accurately capture the structure-performance relationship, researchers find it difficult to efficiently select the most representative sampling points in the configuration space, and thus cannot build an accurate energy landscape model. The lack of such a model directly limits the ability to predict the behavior of materials under dynamic external field conditions, such as the determination of phase transition paths and the analysis of kinetic properties. Therefore, how to capture the atomic-level structure-performance relationship of nanomaterials through efficient computational methods and accurately select key sampling points in the configuration space to construct an energy landscape model has become a key issue in achieving the prediction of phase transition paths and kinetic properties.

[0054] To solve the above problems, refer to Figure 1 The first embodiment of the present application provides a method for predicting nanomaterial properties based on deep learning, comprising the following steps:

[0055] Step 101: Acquire structural data of a nanomaterial, and input the structural data into a pre-trained structural performance prediction model to obtain the configuration space and performance prediction value of the nanomaterial.

[0056] Structural data for nanomaterials typically includes information such as crystal structure, atomic coordinates, and chemical composition. Structure-performance prediction models must be trained on historical nanomaterial structural data to ensure they capture the complex relationship between configuration and performance and can efficiently process a large number of configuration samples to generate performance predictions.

[0057] Step 102: Sampling the configuration space using the Monte Carlo method. If the performance prediction value exceeds a preset threshold, the configuration corresponding to the performance prediction value exceeding the preset threshold is determined as a candidate sampling point to obtain a set of candidate sampling points.

[0058] Random sampling is performed in the configuration space using a Monte Carlo method to generate a set of configuration samples, thereby obtaining the initial configuration set. A structural performance calculation is performed on each configuration in the initial configuration set using a prediction model to obtain a performance prediction value. If the performance prediction value exceeds a preset performance threshold, the corresponding configuration is determined as a candidate sampling point, thereby obtaining the candidate sampling point set.

[0059] For example, the Monte Carlo method generates a large number of possible configurations by randomly perturbing configuration parameters, such as the diameter or bond length of carbon nanotubes. This method relies on a random number generator to sample uniformly or according to a specific distribution in the configuration space, ensuring that the samples cover a variety of configuration features. For example, for carbon nanotube research, the diameter range can be set to 0.8-1.5nm and the bond length range can be set to 0.14-0.15nm. 1,000 configuration samples are generated using the Monte Carlo method, each sample containing a set of diameter and bond length values. This approach ensures the diversity of the configuration set and provides a broad basis for subsequent performance predictions.

[0060] In one possible implementation, the generation of the initial configuration set requires consideration of the boundary conditions of the configuration space. For example, the diameter of the carbon nanotube cannot be too small to avoid structural instability, nor too large to maintain the nanoscale. Preferably, the Monte Carlo-generated samples can be screened using these constraints to eliminate configurations that do not conform to physical laws, such as those with bond lengths outside the range of chemical bonds. This screening improves the physical plausibility of the configuration set and ensures the reliability of subsequent performance predictions.

[0061] In one embodiment, a pre-trained structural performance prediction model can be used to input configuration parameters such as diameter and bond length, and output performance prediction values, such as tensile strength. For example, for a carbon tube configuration with a diameter of 1.0nm and a bond length of 0.142nm, the model predicts its tensile strength to be 1100MPa. If the performance prediction value exceeds a preset performance threshold, the corresponding configuration is determined as a candidate sampling point. For example, the tensile strength threshold can be set to 1000MPa, and only configurations with predicted values ​​above this threshold are retained. For example, assuming that the initial configuration set contains 1000 samples, after prediction, 200 configurations have a tensile strength exceeding 1000MPa, and these configurations are selected as candidate sampling points. Specifically, a configuration with a diameter of 1.2nm and a predicted strength of 1200MPa will be retained, while a configuration with a diameter of 0.9nm and a predicted strength of 900MPa will be eliminated. It should be noted that the setting of the threshold needs to be adjusted according to application requirements. For example, high-strength materials may require a higher threshold.

[0062] In one possible implementation, the candidate sampling point set can be further used to optimize the configuration design. For example, common features among the candidate configurations, such as diameter concentration in the range of 1.1-1.3 nm, can be analyzed to guide subsequent experimental synthesis.

[0063] Using Monte Carlo methods to extract high-performance samples from a large number of random configurations significantly improves design efficiency. Candidate configurations can then be used in molecular dynamics simulations to further validate their performance, enriching the application of configuration-performance mapping. This multi-level screening and analysis approach provides systematic support for the selection of high-performance configurations for nanomaterials.

[0064] Step 103: Density functional theory is used to calculate the energy value of each sampling point in the candidate sampling point set, and an energy landscape model is constructed according to the energy value of each sampling point, wherein the energy landscape model records the energy change and atomic configuration characteristics of each sampling point under external field conditions.

[0065] In this embodiment, the external field conditions include pressure, temperature, etc. The energy landscape model records the energy changes and atomic configuration characteristics of each sampling point under the external field conditions.

[0066] Step 104: Acquire energy change data from the energy landscape model, input the energy change data into the optimized structural performance prediction model, and obtain the phase change path and kinetic characteristics prediction results of the nanomaterial, wherein the pre-trained structural performance prediction model is optimized using a transfer learning method to obtain the optimized structural performance prediction model.

[0067] In this embodiment, initial energy distribution data obtained from an energy landscape database is received and refined using a clustering algorithm to obtain a refined energy landscape dataset. A transfer learning method is used to adjust a pretrained structure-performance prediction model, construct an optimized structure-performance prediction model, and determine weighting parameters for structure and performance prediction. The refined energy landscape dataset is processed using the optimized structure-performance prediction model, outputting a node sequence of the phase transition path and the corresponding final energy value to obtain a phase transition path dataset. Molecular dynamics simulation is then used to calculate the motion trajectory and velocity of each node within the phase transition path dataset, resulting in predictions of the phase transition path and kinetic properties of the nanomaterial.

[0068] For example, when obtaining initial energy distribution data from an energy landscape database, a specialized database interface can be used to filter for energy data relevant to a specific material system. For example, for the two-dimensional graphene system, the database may store energy distribution data for various lattice defects, including energy values ​​for hundreds of configurations in electron volts. The interface then extracts the initial dataset that meets user-defined filtering criteria, such as lattice constant range or defect density.

[0069] In one possible implementation, a clustering algorithm can be used to refine the initial energy distribution data, using either K-means clustering or density clustering. Specifically, for graphene energy distribution data, K-means clustering can divide the data into several clusters, each representing the characteristics of a specific energy state. For example, setting the K value to 5 might yield clusters such as a low-energy stable state and a high-energy metastable state. The center point of each cluster represents a typical configuration, and the standard deviation of the data within the cluster can be used to assess the degree of dispersion of the energy distribution. This clustering method helps extract key features from complex data.

[0070] In some embodiments, when a transfer learning method is used to adjust a pre-trained neural network, fine-tuning can be performed based on an existing general material prediction model. For example, the pre-trained model may have been trained on a variety of two-dimensional materials and contain thousands of energy landscape samples. In the graphene system, a refined energy landscape dataset can be used to retrain some layers of the model and adjust the weights to adapt to the characteristics of the specific system. Preferably, the weight parameters can be determined by optimizing the loss function, for example, setting the weight of the structure prediction to 0.7 and the performance prediction to 0.3. This method can effectively utilize existing knowledge and reduce training time. Specifically, by processing the refined energy landscape dataset through the optimization model and outputting the node sequence of the phase transition path, an energy change path from the initial configuration to the target configuration can be generated. For example, in the phase transition process of graphene from one lattice arrangement to another, the model may output a sequence of 10 key nodes, each node corresponding to the energy value of an intermediate configuration, such as -5.2eV, -5.1eV, etc. These node sequences reflect the energy barriers during the phase transition process and help understand the phase transition mechanism.

[0071] In one embodiment, molecular dynamics simulations can be used to calculate the motion trajectories and velocities of each node along a phase transition path. Classical molecular dynamics tools can be used to simulate atomic motion. For example, for a node in a graphene phase transition path, the temperature can be set to 300K, the simulation time can be set to 100 picoseconds, and the displacement and velocity of carbon atoms can be calculated. Trajectory analysis may reveal that some atoms move significant distances during the phase transition, at rates of approximately 1 angstrom per picosecond. By comparing the velocities and energy changes at different nodes, bottleneck steps in the phase transition process, such as the nodes with the highest energy barriers, can be identified.

[0072] The deep learning-based nanomaterial performance prediction method of the above embodiment obtains the structural data of the nanomaterial and inputs the structural data into a pre-trained structure-performance prediction model to obtain the configuration space and performance prediction value of the nanomaterial; uses the Monte Carlo method to sample the configuration space, and if the performance prediction value exceeds a preset threshold, the configuration corresponding to the performance prediction value exceeding the preset threshold is determined as a candidate sampling point to obtain a set of candidate sampling points; uses density functional theory to calculate the energy value of each sampling point in the set of candidate sampling points, and constructs an energy landscape model based on the energy value of each sampling point; obtains energy change data from the energy landscape model, inputs the energy change data into the optimized structure-performance prediction model, and obtains the phase transition path and kinetic characteristics prediction results of the nanomaterial. The method of the above embodiment realizes the performance prediction, configuration optimization and phase transition kinetic analysis of nanomaterials under different external field conditions, thereby improving the accuracy and efficiency of nanomaterial performance prediction.

[0073] In some embodiments, the method further includes the step of training a structural performance prediction model, the step comprising:

[0074] Step 201: obtain atomic-level structural data from a nanomaterial database, generate an initial configuration set using molecular dynamics simulation, generate a configuration space under corresponding external field conditions based on the initial configuration set, and obtain a configuration space data set based on the configuration space.

[0075] Atomic-level structural data is obtained from a nanomaterial database. A data screening tool is used to extract atomic coordinates and chemical composition information that meets preset conditions. The integrity of the atomic-level structural data is verified using a structure analysis tool to obtain a standardized atomic-level structural data set. A molecular dynamics simulation tool is used to perform dynamic evolution calculations on the standardized atomic-level structural data set according to preset simulation parameters. The atomic motion trajectories are simulated using periodic boundary conditions and force field parameters to generate an initial configuration set. For the initial configuration set, external field conditions including temperature and pressure are set, and the simulation environment is adjusted using an environmental control algorithm. If the temperature or pressure exceeds the preset threshold, the simulation parameters are reconfigured to generate the corresponding configuration space. The configuration space is subjected to data preprocessing, and a clustering algorithm is used to separate the configuration features of the configuration space. The configuration distribution characteristics are calculated using a statistical analysis tool to obtain a configuration space data set.

[0076] In some embodiments, the atomic structure data of nanomaterials generally include information such as crystal structure, atomic coordinates, and chemical composition. For example, the atomic structure data of graphene is extracted from the Materials Project database, which includes the two-dimensional hexagonal lattice coordinates and chemical composition of carbon atoms. The lattice constant range is set to Data screening tools are used to identify qualified graphene monolayer structures, ensuring that the data conforms to experimentally verified geometric properties. Structural analysis tools such as VESTA can be used to verify data integrity, check whether atomic coordinates meet periodic boundary conditions, remove missing or redundant data, and generate standardized atomic-level structural data sets. Molecular dynamics simulation tools are used for dynamic evolution calculations, simulating the motion trajectories of atoms under specific conditions. These methods ensure high data quality and provide a reliable foundation for subsequent simulations.

[0077] Exemplarily, the graphene structure is simulated using LAMMPS software, with the temperature set to 300K and the pressure set to 1atm, and the Tersoff force field is used to describe the interaction between carbon atoms. Periodic boundary conditions are applied to the two-dimensional plane to ensure the continuity of the simulation system. The initial configuration set contains multiple randomly perturbed graphene configurations, such as introducing slight ruffles or defects to simulate the non-ideal state of real materials. This process generates a variety of initial configurations, providing rich samples for subsequent analysis. For the initial configuration set, external field conditions are set to simulate the actual environment. For example, the temperature range is set to 200-400K, the pressure is set to 0.5-1.5atm, and the environmental parameters are controlled by the NPT ensemble. If the temperature exceeds 400K, the environmental control algorithm automatically adjusts the time step to 0.5fs and reruns the simulation to ensure stability. This process generates a configuration space, which contains graphene configuration sets at different temperatures and pressures. Preferably, the configuration space removes outliers by data preprocessing, such as eliminating configurations with abnormally high energy, to ensure the physical rationality of the data. The feature separation of the configuration space relies on the clustering algorithm. In one embodiment, the K-means algorithm is used to divide the configurations into three categories based on the atomic spacing and bond angle distribution: planar configuration, slightly ruffled configuration and defective configuration. The statistical analysis tool calculates the distribution ratio of each type of configuration, for example, the planar configuration accounts for 60%, the ruffled configuration accounts for 30%, and the defective configuration accounts for 10%. These distribution characteristics reflect the differences in the stability of graphene under different environments. The clustering results can be used to predict the mechanical or thermal properties of materials under specific conditions. For example, the ruffled configuration may increase flexibility but reduce conductivity. The configuration space data set can be used to optimize the performance of graphene-based nanodevices. Under high temperature and high pressure conditions, the increase in the proportion of defective configurations may lead to device failure. The defect rate can be reduced by adjusting the preparation process parameters.

[0078] In step 202 , the configuration space dataset is divided into a training set and a validation set, a deep learning algorithm is used to construct a structure-performance mapping model, and a nonlinear activation function is used to process the nonlinear characteristics of the performance mapping relationship to obtain an initial neural network structure.

[0079] In this example, configuration parameters and corresponding performance data are obtained from a configuration space dataset. A dataset partitioning method is used to divide these configuration parameters and performance data into a training set and a validation set, resulting in a partitioned dataset. Based on this partitioned dataset, a deep learning network consisting of multiple neural network layers is constructed. Nonlinear activation functions are used to address nonlinear characteristics, resulting in an initial neural network structure.

[0080] Step 203: Iteratively train the training set using an optimization algorithm combined with a loss function. If the prediction error of the training set is lower than a preset threshold, the optimized neural network parameters are obtained.

[0081] The initial neural network structure is used to iteratively train the training set using an optimization algorithm combined with a loss function. If the prediction error of the training set is lower than a preset threshold, the optimized neural network parameters are obtained.

[0082] Step 204 : Back-propagating the configuration parameters of the validation set using the optimized neural network parameters to obtain a structural performance prediction model.

[0083] For the configuration parameters of the verification set, forward propagation calculation is performed using the optimized neural network parameters to obtain performance prediction results and a structural performance prediction model.

[0084] In this embodiment, configurational parameters typically include atomic coordinates, bond lengths, bond angles, and other quantities that describe structural features, while performance data may include mechanical strength, thermal conductivity, or electrical properties of the material. For example, the atomic coordinates and diameter of a carbon nanotube are extracted from a configurational space dataset as configurational parameters, while its tensile strength is recorded as performance data. Specifically, a data parsing tool can be used to filter carbon nanotube configurations with a diameter of 1.0 nm from the database and obtain their maximum load-bearing capacity in a tensile test, such as 1000 MPa. This data extraction method ensures a one-to-one correspondence between parameters and performance, providing a reliable foundation for subsequent modeling. A dataset partitioning method is used to divide the configurational parameters and performance data into training and validation sets, typically proportionally allocated to balance model training and evaluation. For example, 80% of the configuration-performance data pairs can be used as the training set, and the remaining 20% ​​as the validation set. For example, for 1000 sets of carbon nanotube data, 800 sets are used for training and 200 sets are used for validation. Random sampling can be used to ensure that the training set covers different diameters and configurational characteristics, avoiding data bias.

[0085] Stratified sampling can be introduced during division to ensure that the performance data distribution of the training set and the validation set is consistent, such as the tensile strength is evenly distributed within the range of 500-1500MPa. This division method provides diverse data support for model training. Constructing a neural network structure containing multiple layers of neural network layers can capture the complex relationship between configuration and performance. Preferably, a fully connected neural network containing 3 hidden layers can be designed, with 64 neurons in each layer. For example, the input layer receives the diameter and bond length data of the carbon tube, the hidden layer processes the nonlinear characteristics of the data through a nonlinear activation function such as ReLU, and the output layer predicts the tensile strength. The ReLU function can effectively avoid the gradient vanishing problem and ensure the adaptability of the model to complex configuration features. This structural design has good versatility in the prediction of nanomaterial performance. For example, in the prediction of carbon tube strength, there is a nonlinear relationship between configuration parameters such as bond angle and strength. This processing method improves the model's ability to capture nonlinear relationships.

[0086] Iterative training of the training set through an optimization algorithm combined with a loss function is the core of model optimization. For example, the Adam optimization algorithm can be used in combination with the mean square error loss function for training. In one embodiment, the learning rate is set to 0.001, and after 1000 rounds of training, the prediction error of the training set drops below 5%. For example, for the prediction of carbon tube strength, the model iteratively adjusts the weights so that the predicted value gradually approaches the true value, such as the predicted strength of 980MPa, which is close to the true value of 1000MPa. If the error does not meet the standard, the learning rate can be adjusted or the number of training rounds can be increased. This optimization method ensures the high accuracy of the model. For the configuration parameters of the validation set, the forward propagation calculation is performed through the optimized neural network parameters to obtain the performance prediction results.

[0087] In one possible implementation, the model inputs the carbon nanotube diameters and bond lengths from the validation set, and outputs the predicted tensile strength. For example, for a configuration with a 1.2 nm diameter, the model outputs a strength of 1200 MPa, close to the true value of 1180 MPa. Forward propagation utilizes pre-trained weights for rapid computation, avoiding repeated training. This approach provides an efficient approach for predicting nanomaterial properties.

[0088] In some embodiments, density functional theory is used to calculate the energy value of each sampling point in the set of candidate sampling points, and an energy landscape model is constructed according to the energy value of each sampling point, including:

[0089] Step 301 : Density functional theory is used to calculate the energy value of each sampling point in the candidate sampling point set, and an initial energy landscape model is constructed according to the energy value of each sampling point.

[0090] In some embodiments, density functional theory is used to calculate the energy of each sampling point in the candidate sampling point set, and the energy value of each sampling point is obtained and stored in a pre-established database to obtain the initial energy data set. Based on the initial energy data set, the atomic configuration characteristics of each sampling point are extracted, and combined with the energy change data under external field conditions, a feature vector set is generated through a feature data association algorithm to obtain the configuration feature data set. It is determined whether the energy value in the configuration feature data set exceeds a preset threshold. If it exceeds, the atomic position of the corresponding sampling point is optimized through a parallel processing framework, and the configuration characteristics are updated to obtain the optimized configuration data set. Using the optimized configuration data set, an initial energy landscape model is constructed through an iterative update algorithm, and the energy changes and configuration characteristics under external field conditions are recorded and stored in a database to obtain the initial energy landscape model.

[0091] For example, for a carbon nanotube configuration with a diameter of 1.2 nm, density functional theory is used to calculate its energy value, and a result of -7.5 eV / atom is obtained. After the calculation is completed, the energy value is stored in a pre-established database to form an initial energy data set. Atomic configuration features are extracted from the initial energy data set, including features such as atomic coordinates, bond angles, and bond lengths of the carbon nanotubes. For example, the bond angle distribution of a configuration is concentrated at 120 degrees and the bond length is 0.142 nm. These features are quantified and stored. Combined with external field conditions such as energy change data when the electric field strength is 0.1 V / nm, the feature data association algorithm can generate a set of feature vectors.

[0092] In one possible implementation, the bond angle and bond length data are reduced in dimension by principal component analysis to generate a feature vector containing the main configuration information, forming a configuration feature data set. It is determined whether the energy value in the configuration feature data set exceeds a preset threshold. For example, the energy threshold is set to -7.0eV / atom. If the energy value of a certain configuration is -7.5eV / atom, it exceeds the threshold and enters the optimization process. It should be noted that the threshold setting needs to be adjusted according to the material application scenario. For example, high-stability materials may require lower energy values. For configurations that exceed the threshold, the atomic position is optimized through a parallel processing framework. For example, a gradient-based optimization algorithm is used to adjust the atomic coordinates to further reduce the configuration energy, and the updated configuration features are stored as an optimized configuration data set.

[0093] In one embodiment, the bond length of the optimized configuration is adjusted from 0.142nm to 0.141nm, and the energy is reduced to -7.6eV / atom. This optimization improves the stability of the configuration. Based on the optimized configuration data set, an iterative update algorithm is used to construct an initial energy landscape model. The algorithm records the energy changes and configuration characteristics under external field conditions such as different electric field strengths through multiple iterations. For example, when the electric field increases from 0.1V / nm to 0.5V / nm, the configuration energy change trend is recorded, forming a mapping of energy changes with configuration and external field.

[0094] Preferably, the model data is stored in a database to facilitate subsequent query and analysis.

[0095] In one possible implementation, the initial energy landscape model can be further used to predict the properties of new configurations. For example, model analysis revealed that configurations with bond angles close to 120 degrees have lower energies under strong electric fields, suggesting that these configurations may have excellent electrical properties. This analysis supports the direction of configuration optimization from multiple perspectives and enriches the research perspectives of nanomaterial design.

[0096] In step 302, energy distribution and configuration space feature data are obtained from the initial energy landscape model. If the energy value is lower than the average value and the configuration feature meets the structural parameters of the high-performance area, the configuration corresponding to the energy value lower than the average value and the configuration feature meeting the structural parameters of the high-performance area is determined as a priority sampling area, thereby obtaining a priority sampling area dataset.

[0097] The energy distribution and configuration space feature data are obtained from the initial energy landscape model, and the probability density function of the energy value is calculated using a statistical analysis tool to obtain an energy distribution data set. At the same time, the atomic coordinates and bond length information of the configuration space are extracted using a geometric analysis tool to generate a configuration feature data set. For the energy distribution data set, the average energy value is determined using a mean calculation method. By comparing each energy value with the average energy value, the energy level is judged to obtain a subset below the average energy. If the energy value is lower than the average energy and the configuration feature meets the structural parameters of the preset high-performance area, the corresponding configuration is marked as a priority sampling area through a logical screening tool to obtain a candidate set of the priority sampling area. Based on the candidate set of the priority sampling area, the energy value and the configuration feature are associated and mapped using a data integration tool to generate a complete data set containing the priority sampling area.

[0098] The above method uses a probability density function to describe the probability of energy values ​​appearing in a specific interval, which helps to reveal the law of energy distribution. The geometric analysis tool extracts the atomic coordinates and bond length information of the configuration space to generate a configuration feature data set. The geometric analysis tool can calculate the bond length based on the Euclidean distance between atoms and record the three-dimensional coordinates of each atom. For example, for a molecular system containing 10 atoms, the extracted configuration features may include a C-H bond length of The C-C bond length is As well as the x, y, and z coordinates of each atom. These characteristic data can be used to characterize the spatial structure of the configuration and provide a basis for the subsequent screening of high-performance configurations. For the energy distribution data set, the mean calculation method is used to determine the average energy value, and the relationship between each energy value and the average value is compared. For example, assuming that the average energy value of 1000 sampling points is -1.5eV, after traversing the data set, it is found that the energy values ​​of approximately 600 sampling points are lower than -1.5eV. These sampling points are classified as low-energy subsets. This method is intuitive and efficient, and can quickly screen out low-energy configurations, reducing the computational complexity of subsequent analysis.

[0099] For example, if the energy value is lower than the average energy and the configuration characteristics meet the structural parameters of the preset high-performance region, the priority sampling region is marked by the logic screening tool. For example, the preset high-performance region requires the C—C bond length to be between to and the energy is lower than -1.5eV. Assume that the energy of a sampling point is -2.0eV and the CC bond length is The sampling point is marked as a priority sampling area. This screening method ensures that the configuration is not only low in energy but also has ideal geometric properties.

[0100] In some embodiments, based on the candidate set of priority sampling regions, a data integration tool is used to associate and map energy values ​​with configurational features to generate a complete data set. For example, for 100 sampling points marked as priority sampling regions, the data integration tool can generate a structured data set containing information such as energy values, atomic coordinates, and bond lengths. For example, the record of a sampling point may be: Energy = -2.1 eV, Coordinates = [0.1, 0.2, 0.3]. This correlation mapping facilitates the subsequent analysis of the relationship between configuration and energy.

[0101] Step 303: Obtain energy distribution and configuration features from the data set of the priority sampling area, and use the Bayesian optimization algorithm to update the sampling point selection strategy based on the energy distribution and configuration features to generate the first candidate sampling point and obtain an optimized sampling point set.

[0102] In this embodiment, the energy distribution reflects the stability differences of the atoms in the nanomaterial. A data set is obtained from a priority sampling area, and the priority sampling area is divided into multiple sub-areas using a grid partitioning method. The energy value distribution and geometric topological characteristics are calculated for each sub-area to obtain the energy distribution and configuration characteristics. A priori probability distribution is generated based on the energy distribution. If the energy value is higher than a preset threshold, a predetermined probability value is assigned. The priori probability distribution is normalized using a probability density function to obtain the initial probability distribution. The initial probability distribution is updated using a Bayesian optimization algorithm and configuration characteristics. The sampling point positions are adjusted according to the geometric constraints in the configuration characteristics. The Gaussian process regression algorithm is used for iterative optimization to obtain an updated probability distribution. The Monte Carlo method is used to sample the updated probability distribution to generate a set of candidate sampling points. The candidate sampling point set is screened and the sampling points that meet the configuration characteristics and the energy distribution constraints are retained to obtain an optimized set of sampling points.

[0103] In one possible implementation, assuming that the priority sampling area is a two-dimensional plane containing multiple atomic configurations, it can be divided into 100 sub-grids with a spacing of 0.5 nanometers on each side. The energy value distribution is calculated in each sub-grid to obtain statistical features such as average energy and variance, and geometric topological features such as interatomic distance or bond angle distribution are extracted. Specifically, the calculation of the energy value distribution can be achieved through statistical tools. For example, for a certain sub-grid, assuming that the calculated energy value range is -5 to -3 electron volts, a probability density function can be generated by the histogram method, and then it is concluded that the proportion of areas with energy higher than -4 electron volts is 30%. If the preset threshold is -4 electron volts, the area above this threshold is given a higher prior probability, such as 0.7, while the probability of the area below the threshold is 0.3.

[0104] In one possible implementation, the configurational features are assumed to include constraints on atomic spacing and bond angles, such as requiring spacing to be between 0.1 and 0.2 nanometers. Using a Bayesian approach, the probability distribution is iteratively updated based on observed subgrid data, increasing the probability of regions satisfying the geometric constraints. For example, a subgrid with an initial probability of 0.5 might have a probability of 0.8 after the update if the spacing constraint is satisfied. This dynamic adjustment improves sampling efficiency.

[0105] In one possible implementation, 1,000 candidate sampling points are generated based on the updated probability distribution, and points that meet the requirements of energy below -4 electron volts and spacing between 0.1 and 0.2 nanometers are screened out to obtain an optimized sampling point set, such as 200 points. This screening process retains high-quality sampling points.

[0106] In some embodiments, generating an optimized set of sampling points can also incorporate additional constraints, such as symmetry requirements for atomic arrangement. For example, if a high-performance region requires hexagonal symmetry in its local configuration, sampling points that do not meet this symmetry can be eliminated during screening. This multi-constraint screening further improves the accuracy of the dataset.

[0107] Step 304 : using a minimum energy path algorithm to analyze the energy conversion paths between the sampling points in the optimized sampling set; if the path energy difference is lower than a preset threshold, it is determined to be a stable phase change path, and a phase change path set is obtained.

[0108] Energy change data is obtained from the initial energy landscape model. The energy landscape is divided into multiple subregions using a gridding method. For each subregion, a sequence of energy values ​​is calculated to obtain an energy change dataset. Based on this energy change dataset, a minimum energy path algorithm is used to calculate the transition paths between the first candidate sampling points. For each path, the energy values ​​at the start and end points are extracted, and the path energy difference is calculated to obtain the path energy difference set. For this path energy difference set, if the path energy difference is below a preset threshold, it is determined to be a stable phase change path. These stable phase change paths are then grouped using a clustering method to obtain the preliminary phase change path set. This preliminary phase change path set is processed using a topological analysis method to extract the geometric constraint characteristics of each path. Paths that meet these phase change characteristics are screened to obtain the phase change path set.

[0109] For example, in the process of obtaining energy change data from the initial energy landscape model, the energy landscape can be understood as a two-dimensional or three-dimensional distribution diagram that describes the change of energy with configuration. Assuming that the behavior of molecules adsorbed on the metal surface is being studied, the energy landscape represents the change of potential energy of the molecules at different positions on the surface. The energy landscape is divided into multiple sub-regions using a grid partitioning method, and the size of each sub-region can be 0.5 nm × 0.5 nm. For each sub-region, a sequence of energy values ​​is calculated through molecular dynamics simulation. For example, the energy values ​​of molecules in a certain sub-region are recorded 100 times to obtain an energy change data set. This data set can reflect the energy fluctuations of atoms in nanomaterials at different positions.

[0110] In one possible implementation, a minimum energy path algorithm is used to calculate the transition paths between sampling points. For example, assume there are two sampling points, A and B, corresponding to two low-energy adsorption sites of molecules on a metal surface. The algorithm calculates the path from A to B, extracting the energy value of -5.0 eV at the starting point of the path and -4.8 eV at the end point, and calculating the path energy difference to be 0.2 eV. The energy difference set of all paths reflects the possible configurational transition trends. Specifically, for the processing of the path energy difference set, if the path energy difference is below a preset threshold, such as 0.3 eV, the path is considered to correspond to a stable phase transition path. This is because a low energy difference indicates that the energy barrier on the path is small, and atoms are more likely to move along this path. For example, the energy differences of paths AB and CD are 0.2 eV and 0.25 eV, respectively, both below the threshold, and are judged to be stable phase transition paths. Clustering methods, such as K-means clustering, are used to group these paths, and similar paths are classified into one category to obtain a preliminary phase transition path set.

[0111] In this embodiment, a topological analysis method is used to extract the geometric constraint characteristics of the path and further screen the paths. For example, in a metal surface adsorption scenario, some paths may pass through highly symmetric bridge positions or top positions. The geometric characteristics of these paths are extracted through topological analysis. Only paths that meet the phase transition characteristics, such as path curvature less than a certain threshold or path length within a reasonable range, are retained, resulting in an optimized phase transition path set. This screening ensures that the path set is more consistent with actual physical constraints.

[0112] The above method uses the dual constraints of energy change and geometric characteristics to generate a set of phase change paths that can more accurately describe the dynamic behavior of the system.

[0113] Step 305 : Based on the phase transition path set, a kinetic Monte Carlo method is used to simulate the atomic motion trajectory of the nanomaterial under external field conditions, and the time evolution characteristics on the path are calculated to obtain kinetic characteristic data.

[0114] In this example, external field conditions such as electric field strength and temperature parameters are set. Using the kinetic Monte Carlo method, random sampling is used to obtain atomic motion trajectories from a set of phase transition paths. The position and velocity of the atoms in the nanomaterial are calculated over time to generate time-evolution characteristics. Based on these time-evolution characteristics, statistical analysis tools are used to calculate the transition frequency and diffusion coefficient of the atomic motion. If the transition frequency exceeds a preset threshold, the path is subdivided to generate kinetic characteristic data.

[0115] Using path statistics, the probability distribution of the phase transition path set is derived from the kinetic property data. Cluster analysis is then performed on this probability distribution to determine the stability and dominance of each path, generating path statistics. Simulation accuracy assessment tools are then used to analyze the deviation between these path statistics and the kinetic property data. If the deviation exceeds a preset threshold, the random sampling step size and external field condition parameters are adjusted to obtain optimized time evolution characteristics.

[0116] In one embodiment, the polarization reversal process of the ferroelectric material PbTiO3 can be simulated, and the atomic motion trajectory can be obtained by random sampling. Specifically, the electric field strength is set to 10kV / cm and the temperature field is set to 300K. The displacement of the Ti atom in the oxygen octahedron is calculated over time to obtain the time evolution characteristics from the upper polarization state to the lower polarization state. For the polarization reversal of the above ferroelectric material, the transition frequency of the Ti atom can be calculated to be about 10 12 Hz, diffusion coefficient is 10 -6 cm 2 / s. When the transition frequency is higher than the preset threshold 10 11When the kinetic characteristics are higher than Hz, the path needs to be subdivided, reducing the original 10ps time step to 1ps to obtain more accurate dynamic characteristic data. When the path statistics deviate from the dynamic characteristic data by more than 5%, it is necessary to adjust the random sampling step size from 2fs to 0.5fs, and reset the electric field strength to 8kV / cm and the temperature field to 320K to obtain more accurate time evolution characteristics.

[0117] Preferably, during the phase transition of ferroelectric materials, the influence of different factors on the phase transition path can be studied by adjusting the external field condition parameters. When the temperature increases from 300K to 400K, the proportion of direct flipping pathways decreases to 45%, while the proportion of rotational pathways increases to 40%, indicating that increasing temperature promotes competition among multiple phase transition pathways, thereby affecting the macroscopic polarization response characteristics of the material.

[0118] Step 306 , obtaining time evolution characteristics from the dynamic characteristic data, combining with the phase change path set, and using the Bayesian optimization algorithm to iteratively update the sampling point selection strategy to generate second candidate sampling points, thereby obtaining a re-optimized sampling point set.

[0119] Time evolution features are obtained from the dynamic characteristic data, and the data is processed using a feature extraction method. A first feature set is obtained by calculating the autocorrelation function and spectral analysis of the time series. A phase change path set is constructed based on the first feature set. If the state transition frequency in the first feature set is higher than a preset threshold, a phase change path set is generated using a clustering algorithm. Otherwise, a phase change path set is generated using an interpolation method to obtain a phase change path set. A Bayesian optimization algorithm is used to process the phase change path set. Based on a pre-established probability distribution model, a posterior probability distribution is calculated using Gaussian process regression to generate a second set of candidate sampling points. Sampling points are screened from the second set of candidate sampling points using a path similarity metric. If the path similarity metric value of the sampling point is higher than a preset threshold, the sampling point is retained; otherwise, it is eliminated, resulting in a further optimized sampling point set.

[0120] In one possible implementation, when extracting time evolution features from dynamic characteristic data, the core of the feature extraction method is to capture the dynamic behavior of atomic motion. The autocorrelation function of the time series is used to measure the correlation of atomic position changes over time, reflecting the periodicity or regularity of the motion; spectrum analysis converts the time series into the frequency domain through Fourier transform to reveal the frequency distribution of atomic motion. For example, when simulating the phase transition of nanomaterials, assuming that a time series records the position changes of atoms under an electric field strength of 1V / nm, the autocorrelation function shows a periodic fluctuation of 0.1ps, indicating that the atoms have short-term oscillation behavior. Spectral analysis further confirms that the main frequency is concentrated at 10THz, generating a first feature set containing information such as period, frequency, and amplitude.

[0121] For example, assuming a preset threshold of 1 GHz, if the atomic transition frequency in the feature set reaches 2 GHz, it indicates that the path changes dramatically, and it is suitable to use a clustering algorithm such as K-means for path segmentation. Clustering can be used to classify similar trajectories into the same path category. For example, 100 trajectories can be divided into 3 categories, each representing a phase transition mode. If the frequency is lower than the threshold, such as 0.5 GHz, an interpolation method is used to supplement the path points through linear interpolation to ensure path smoothness. For example, 50 points are generated by interpolation from 10 sampling points to form a complete first path set, which ensures the continuity of the path.

[0122] Preferably, when processing a set of phase transition paths, the Bayesian optimization algorithm calculates a posterior probability distribution based on Gaussian process regression to determine the optimal sampling points. For example, in a nanomaterial simulation, the initial path set contains 1000 points. After 50 iterations of Bayesian optimization, 200 candidate sampling points are generated, with the probability distribution concentrated in the high transition frequency region. These candidate points are sorted by posterior probability to ensure that the sampling points cover the key phase transition regions, thereby improving the representativeness of the path.

[0123] In one embodiment, a path similarity metric is used to screen sampling points, typically using Euclidean distance or a dynamic time warping algorithm to calculate differences between paths. For example, a similarity threshold is set to 0.8, and the similarity metric between two paths is calculated. If the value is 0.9, the sampling point is retained; if it is 0.7, it is removed. Assuming the initial candidate set has 200 points, 150 points are retained after screening. The optimized sampling point set more accurately reflects the dynamic characteristics of the phase transition path. This screening method reduces redundant points and improves computational efficiency.

[0124] Each step of the above method revolves around the dynamic characteristics of the phase transition of nanomaterials. Bayesian optimization improves sampling efficiency, and similarity measurement ensures the accuracy of the path. Together, they provide reliable data support for the dynamic behavior of the phase transition path and are suitable for the motion trajectory analysis of nanomaterials under external field conditions.

[0125] Step 307 : Based on the re-optimized sampling point set and the newly added external field condition data, density functional theory is used to update the initial energy landscape model to obtain a final energy landscape model.

[0126] The sampling point set is screened by a preset threshold, and a subset that conforms to the density distribution is obtained from the re-optimized sampling point set. A parallel computing method is used to perform density functional calculations to obtain preliminary energy data. Based on the preliminary energy data, in combination with quantum chemical calculation tools, external field condition data is obtained from an external database. The external field condition data is integrated into the data fusion algorithm to generate an updated energy data set. If the updated energy data set does not meet the preset model accuracy threshold, an iterative optimization algorithm is used to adjust the parameters, and a refined energy landscape data set is generated through multiple rounds of calculations. A consistency check tool is used to determine the compatibility of the refined energy landscape data set with the environmental variables, and key features are extracted from the refined energy landscape data set to obtain the final refined energy landscape model.

[0127] For example, a probability density-based method can be used to screen the sampling point set. Assuming that the research object is the atomic position data in the molecular dynamics simulation, the density threshold is set to 0.8, and the sampling points with a density higher than this threshold are retained. In one possible implementation, the kernel density estimation method is used to calculate the local density of each sampling point, and a subset that meets the conditions is screened out. This method can effectively focus on high-density areas and reduce the interference of redundant data on subsequent calculations. Specifically, after obtaining a subset that meets the density distribution from the sampling point set, density functional calculations can be performed through parallel computing. For example, in a set of 1,000 sampling points, 200 high-density points are screened out, and density functional theory calculations are run in parallel using a multi-core processor to obtain preliminary energy data for each sampling point.

[0128] In one embodiment, the external field condition data is obtained from an external database in combination with quantum chemical calculation tools based on the preliminary energy data. For example, the electric field strength data is extracted from the material database using Gaussian software, assuming that the external field strength is Preferably, the field data is integrated with the preliminary energy data through a data fusion algorithm, such as a weighted average method, to generate an updated energy data set. This fusion method can more realistically reflect the impact of the external environment on the energy distribution. If the updated energy data set does not reach the preset model accuracy threshold, for example, the deviation from the target accuracy exceeds 5%, an iterative optimization algorithm is used to adjust the parameters. For example, the molecular spacing parameters are gradually adjusted by the gradient descent method, and after 3 rounds of iterations, the data set deviation is reduced to within 2%, thereby generating a refined energy landscape data set. This iterative process can gradually approach the optimal solution. For example, a consistency verification tool is used to determine the compatibility of the refined energy landscape data set with environmental variables. Assuming that the environmental variables include temperature and pressure, the verification tool determines its adaptability by comparing the stability of the data set under 300K and 1atm conditions. If the data set remains consistent under multiple conditions, it is considered to have good adaptability. This verification method ensures the robustness of the model.

[0129] In one possible implementation, key features, such as energy minima and transition state locations, are extracted from the refined energy landscape dataset. Suppose five key features are extracted, corresponding to the stable and transition states of the molecular configuration. These features can be used to construct the final refined energy landscape model. For example, visualization tools can be used to create an energy landscape diagram, visualizing the energy variations of the nanomaterial's atoms under different configurations.

[0130] Reference Figure 2 The second embodiment of the present application provides a nanomaterial performance prediction system based on deep learning, including:

[0131] The data acquisition module 401 is used to acquire the structural data of the nanomaterial and input the structural data into a pre-trained structural performance prediction model to obtain the configuration space and performance prediction value of the nanomaterial;

[0132] Sampling module 402 is used to sample the configuration space using the Monte Carlo method. If the performance prediction value exceeds a preset threshold, the configuration corresponding to the performance prediction value exceeding the preset threshold is determined as a candidate sampling point to obtain a set of candidate sampling points.

[0133] An energy landscape model construction module 403 is configured to calculate the energy value of each sampling point in the candidate sampling point set using density functional theory, and to construct an energy landscape model based on the energy value of each sampling point, wherein the energy landscape model records the energy changes and atomic configuration characteristics of each sampling point under external field conditions;

[0134] The prediction result module 404 is used to obtain energy change data from the energy landscape model, input the energy change data into the optimized structural performance prediction model, and obtain the phase change path and kinetic characteristics prediction results of the nanomaterial, wherein the structural performance prediction model is optimized using a transfer learning method to obtain the optimized structural performance prediction model.

[0135] It should be noted that the nanomaterial performance prediction system based on deep learning provided in the embodiment of the present application is used to execute all the process steps of the nanomaterial performance prediction method based on deep learning in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.

[0136] The present application also provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data acquisition program. When the processor executes the computer program, the steps in the above-mentioned method for predicting nanomaterial properties based on deep learning are implemented, such as Figure 1Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned system embodiments are realized, such as the data acquisition module.

[0137] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the electronic device.

[0138] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.

[0139] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire electronic device using various interfaces and lines.

[0140] The memory can be used to store the computer programs and / or modules, and the processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0141] Wherein, if the module / unit integrated in the electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0142] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive work.

[0143] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application by those skilled in the art should be included within the scope of protection of this application.

Claims

1. A method for predicting nanomaterial properties based on deep learning, characterized in that: include: Acquiring structural data of the nanomaterial and inputting the structural data into a pre-trained structural performance prediction model to obtain a configuration space and performance prediction value of the nanomaterial; The configuration space is sampled using a Monte Carlo method. If a performance prediction value exceeds a preset threshold, the configuration corresponding to the performance prediction value exceeding the preset threshold is determined as a candidate sampling point to obtain a set of candidate sampling points. Density functional theory is used to calculate the energy value of each sampling point in the set of candidate sampling points, and an energy landscape model is constructed according to the energy value of each sampling point, wherein the energy landscape model records the energy change and atomic configuration characteristics of each sampling point under external field conditions; Energy change data is obtained from the energy landscape model, and the energy change data is input into the optimized structural performance prediction model to obtain the phase change path and kinetic characteristics prediction results of the nanomaterial, wherein the pre-trained structural performance prediction model is optimized using a transfer learning method to obtain the optimized structural performance prediction model.

2. The method according to claim 1, characterized in that Also includes: The step of training the structural performance prediction model comprises: Acquiring atomic-level structural data from a nanomaterial database, generating an initial configuration set using molecular dynamics simulation, generating a configuration space under corresponding external field conditions based on the initial configuration set, and obtaining a configuration space data set based on the configuration space; The configuration space data set is divided into a training set and a validation set, a deep learning algorithm is used to construct a structure-performance mapping model, and a nonlinear activation function is used to process the nonlinear characteristics of the performance mapping relationship to obtain an initial neural network structure; The training set is iteratively trained using an optimization algorithm combined with a loss function. If the prediction error of the training set is lower than the preset threshold, the optimized neural network parameters are obtained. The configuration parameters of the verification set are back-propagated through the optimized neural network parameters to obtain a structural performance prediction model.

3. The method according to claim 1, characterized in that The step of calculating the energy value of each sampling point in the candidate sampling point set by using density functional theory and constructing an energy landscape model according to the energy value of each sampling point includes: Density functional theory is used to calculate the energy value of each sampling point in the candidate sampling point set, and an initial energy landscape model is constructed according to the energy value of each sampling point; Obtaining energy distribution and configuration space characteristic data from the initial energy landscape model; if the energy value is lower than the average value and the configuration characteristics meet the structural parameters of the high-performance area, determining the configuration corresponding to the energy value lower than the average value and the configuration characteristics meeting the structural parameters of the high-performance area as a priority sampling area, and obtaining a priority sampling area data set; The energy distribution and configuration characteristics are obtained from the data set of the priority sampling area. The Bayesian optimization algorithm is used to update the sampling point selection strategy based on the energy distribution and configuration characteristics to generate the first candidate sampling point and obtain the optimized sampling point set. The minimum energy path algorithm is used to analyze the energy conversion paths between the sampling points in the optimized sampling set. If the path energy difference is lower than a preset threshold, it is determined to be a stable phase change path, and a phase change path set is obtained. Based on the phase transition path set, the dynamic Monte Carlo method is used to simulate the atomic motion trajectory of nanomaterials under external field conditions, calculate the time evolution characteristics of the path, and obtain the dynamic characteristic data; The time evolution characteristics are obtained from the dynamic characteristic data, combined with the phase change path set, and the sampling point selection strategy is iteratively updated using the Bayesian optimization algorithm to generate the second candidate sampling points and obtain the optimized sampling point set again. According to the re-optimized sampling point set and the newly added external field condition data, the initial energy landscape model is updated using density functional theory to obtain the final energy landscape model.

4. The method according to claim 3, characterized in that The energy distribution and configuration characteristics are obtained from the data set of the priority sampling area, and the sampling point selection strategy is updated using the Bayesian optimization algorithm according to the energy distribution and configuration characteristics to generate the first candidate sampling point, thereby obtaining the optimized sampling point set, including: Acquiring a data set from the priority sampling area, dividing the priority sampling area into a plurality of sub-areas using a grid partitioning method, and calculating the energy value distribution and geometric topological characteristics of each sub-area to obtain energy distribution and configuration characteristics; generating a priori probability distribution based on the energy distribution, assigning a predetermined probability value if the energy value is higher than a preset threshold, and normalizing the priori probability distribution using a probability density function to obtain the initial probability distribution; The initial probability distribution is updated by using a Bayesian optimization algorithm and configuration features, the sampling point positions are adjusted according to the geometric constraints in the configuration features, and an updated probability distribution is obtained by iterative optimization using a Gaussian process regression algorithm; The updated probability distribution is sampled using a Monte Carlo method to generate a set of candidate sampling points. The candidate sampling point set is screened, and sampling points that meet the configuration characteristics and the energy distribution constraints are retained to obtain an optimized sampling point set.

5. The method according to claim 3, characterized in that The minimum energy path algorithm is used to analyze the energy conversion path between the optimized sampling points. If the path energy difference is lower than the preset threshold, it is determined to be a stable phase change path, and the phase change path set is obtained, including: Obtaining energy change data from the initial energy landscape model, dividing the energy landscape into a plurality of sub-regions using a grid partitioning method, and calculating an energy value sequence for each sub-region to obtain an energy change data set; calculating, based on the energy change dataset, a minimum energy path algorithm for calculating conversion paths between the first candidate sampling points, extracting energy values ​​of a start point and an end point for each path, calculating a path energy difference, and obtaining a path energy difference set; For the path energy difference set, if the path energy difference is lower than a preset threshold, it is determined to be a stable phase change path, and the stable phase change paths are grouped using a clustering method to obtain the preliminary phase change path set; The preliminary phase change path set is processed by a topological analysis method, the geometric constraint characteristics of each path are extracted, and the paths that meet the phase change characteristics are screened to obtain the phase change path set.

6. The method according to claim 3, characterized in that The time evolution characteristics are obtained from the dynamic characteristic data, combined with the phase change path set, and the sampling point selection strategy is iteratively updated using the Bayesian optimization algorithm to generate the second candidate sampling points, thereby obtaining a further optimized sampling point set, including: Acquiring time evolution features from the dynamic characteristic data using a feature extraction method, calculating the autocorrelation function and spectrum analysis of the time series in the time evolution features, and obtaining a first feature set; Constructing a phase change path set based on the first feature set; if the state transition frequency in the first feature set is higher than a preset threshold, generating a phase change path set using a clustering algorithm; otherwise, generating a first path set using an interpolation method to obtain a phase change path set; The phase change path set is processed using a Bayesian optimization algorithm, and based on a pre-established probability distribution model, a posterior probability distribution is calculated through Gaussian process regression to generate a second set of candidate sampling points; Sampling points are screened from the second candidate sampling point set by using the path similarity metric. If the path similarity metric value of the sampling point is greater than a preset threshold, the sampling point is retained to obtain a re-optimized sampling point set.

7. The method according to claim 1, characterized in that The step of obtaining energy change data from the energy landscape model and inputting the energy change data into the optimized structural performance prediction model to obtain the phase transition path and kinetic characteristics prediction results of the nanomaterial includes: receiving initial energy distribution data obtained from an energy landscape database, and refining the initial energy distribution data using a clustering algorithm to obtain a refined energy landscape dataset; Adopting a transfer learning method to adjust the pre-trained structural performance prediction model to construct an optimized structural performance prediction model; Processing the refined energy landscape dataset through the optimized structural performance prediction model, outputting a node sequence of a phase transition path and a corresponding final energy value, to obtain a phase transition path dataset; For the phase change path data set, a molecular dynamics simulation method is used to calculate the motion trajectory and rate of each node to obtain the phase change path and kinetic characteristics prediction results of the nanomaterial.

8. A nanomaterial performance prediction system based on deep learning, characterized in that: include: A data acquisition module is used to obtain structural data of the nanomaterial and input the structural data into a pre-trained structural performance prediction model to obtain the configuration space and performance prediction value of the nanomaterial; a sampling module, configured to perform sampling processing on the configuration space using a Monte Carlo method, and if a performance prediction value exceeds a preset threshold, determine the configuration corresponding to the performance prediction value exceeding the preset threshold as a candidate sampling point, thereby obtaining a set of candidate sampling points; an energy landscape model construction module, configured to calculate the energy value of each sampling point in the candidate sampling point set using density functional theory, and construct an energy landscape model based on the energy value of each sampling point, wherein the energy landscape model records the energy changes and atomic configuration characteristics of each sampling point under external field conditions; A prediction result module is used to obtain energy change data from the energy landscape model, input the energy change data into the optimized structural performance prediction model, and obtain the phase change path and kinetic characteristics prediction results of the nanomaterial, wherein the pre-trained structural performance prediction model is optimized using a transfer learning method to obtain the optimized structural performance prediction model.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for predicting nanomaterial properties based on deep learning as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the nanomaterial performance prediction method based on deep learning according to any one of claims 1 to 7.