A method and system for directed generation of metastable sulfide crystal structures
Patent Information
- Application Number
- CN202611289577.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]有鉴于此,本发明的目的在于克服现有技术的上述缺陷,提供一种亚稳态硫化物晶体结构定向生成方法及系统,用于解决现有技术中传统生成模型无法定向生成、条件约束与生成质量之间存在矛盾以及生成式搜索缺乏收敛保证的技术问题,具体包括:
[0018]1、通过构建条件扩散模型并将目标物理量作为引导条件输入,结合物理约束损失函数训练去噪网络,实现了根据性能目标反向定制亚稳态硫化物新晶体结构,克服了传统生成模型无法定向生成的缺陷。通过物理约束损失函数中的原子间距惩罚项和成键规则惩罚项,有效抑制了非物理结构的生成;
Smart Images

Figure CN122822103A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of computational materials science and artificial intelligence, specifically to a method and system for the directional generation of metastable sulfide crystal structures. Background Technology
[0002] In the development of solid-state electrolytes and electrode materials, sulfide crystals have attracted widespread attention due to their excellent ionic conductivity. Traditional methods for discovering new crystal structures rely on trial-and-error experiments and high-throughput computational screening, which is inefficient. In recent years, deep learning-based crystal structure generation methods have made some progress, with representative techniques including crystal generation methods based on generative adversarial networks (GANs) and crystal diffusion variational autoencoders (VADEs). These models learn crystal structure distributions through forward noise addition and backward denoising processes, achieving some improvement in addressing the problems of structure generation and mode collapse.
[0003] However, existing deep learning-based crystal structure generation schemes generally employ an unconditional generation mode. The model only outputs the generated structure and cannot generate it in a targeted manner based on the target physical properties. The vast majority of the generated structures are thermodynamically unstable and cannot be synthesized experimentally. To improve the stability of the generated structures, some schemes introduce conditional control or physical constraints, such as adding conditional guidance or auxiliary loss functions during the generation process.
[0004] Existing technologies for generating sulfide crystal structures suffer from numerous technical shortcomings. They fail to balance the physical rationality of the structure with the guidance of target performance, resulting in low search efficiency. Specifically, traditional generative adversarial network methods are prone to generating non-physical structures such as atomic overlap and unreasonable chemical bonds due to training instability and pattern collapse. Furthermore, unconditional diffusion models lack performance target guidance, resulting in an extremely low proportion of generated structures that satisfy metastable conditions. In addition, when conditional control is introduced, strong conditional constraints degrade the diversity and physical rationality of generated structures, creating an inherent contradiction between condition satisfaction and generation quality. This leads to a lack of reliable structural guarantees in the scenario of directed generation of new materials. Summary of the Invention
[0005] In view of this, the purpose of this invention is to overcome the above-mentioned defects of the prior art and provide a method and system for the directional generation of metastable sulfide crystal structures. This method addresses the technical problems in the prior art, such as the inability of traditional generation models to generate structures in a directional manner, the contradiction between conditional constraints and generation quality, and the lack of convergence guarantees in generative searches. Specifically, it includes:
[0006] The crystal structure representation vector, electronic descriptor, and target physical performance parameters of the sulfide crystal structure are obtained. The preset target physical performance parameters are mapped to a conditional vector via a conditional encoder. The crystal structure representation vector and the electronic descriptor are concatenated to form an extended crystal representation vector. The conditional vector and the extended crystal representation vector are paired to generate conditional-structure training pairs. Noise is added to the extended crystal representation vector in the conditional-structure training pairs to obtain a noise scheduling sequence. A denoising network is constructed based on the noise scheduling sequence and the conditional vector. The denoising network is trained using a loss function containing physical constraints to obtain a physical constraint conditional diffusion model. The physical constraint diffusion model and the condition vector are used to perform inverse denoising sampling, which includes both conditional and unconditional denoising guidance, to output preliminary denoising results. Adaptive guidance weights are determined based on the degree of matching between the preliminary denoising results and the physical constraints. Gradient correction is applied to the preliminary denoising results according to these adaptive guidance weights to generate candidate crystal structure variables. These candidate crystal structure variables are then filtered using a surrogate model to output near-optimal structures. Importance weights are determined based on the probability ratio of the near-optimal structures under guided and unconditional sampling distributions. The near-optimal structures are then resampled according to these importance weights to correct distribution shifts, verifying and outputting the metastable sulfide crystal structure.
[0007] Further, the acquisition of the crystal structure representation vector, electronic descriptor, and target physical property parameters of the sulfide crystal structure includes: extracting the atom type sequence, fractional coordinate matrix, and lattice parameters of the crystal from a database of sulfide crystal structures; performing one-hot encoding on the atom type sequence, performing continuous value encoding on the fractional coordinate matrix and the lattice parameters, and concatenating the encoding results of the one-hot encoding and the continuous value encoding to obtain the crystal structure representation vector; extracting local electronic descriptors from each crystal structure in the crystal structure representation vector, wherein the local electronic descriptors include Bader charge and atomic density of states; obtaining the formation energy of each crystal structure in the crystal structure representation vector through first-principles calculations, and obtaining the room temperature ionic conductivity through molecular dynamics simulations; and combining the formation energy and the room temperature ionic conductivity of each crystal structure to form a physical property tag set.
[0008] Further, the step of mapping the preset target physical performance parameters into a conditional vector via a conditional encoder, concatenating the crystal structure representation vector with the electronic descriptor to form an extended crystal representation vector, and pairing the conditional vector with the extended crystal representation vector to generate a conditional-structure training pair includes: mapping the preset target physical performance parameters into a conditional vector via a multilayer perceptron conditional encoder, wherein the multilayer perceptron conditional encoder is composed of alternately stacked fully connected layers, batch normalization layers, and activation layers; concatenating the crystal structure representation vector with the electronic descriptor according to dimension to obtain an extended crystal representation vector; and pairing the conditional vector with the extended crystal representation vector according to batch dimension to form a conditional-structure training pair.
[0009] Further, the step of adding noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noise scheduling sequence includes: gradually adding Gaussian noise to the extended crystal representation vector in the conditional-structure training pair according to a preset noise scheduling table, wherein the addition of Gaussian noise only affects the fractional coordinate matrix and lattice parameters, and the atom type remains unchanged; after a preset time step, transforming the extended crystal representation vector into a standard Gaussian distribution to obtain the noise scheduling sequence.
[0010] Further, the step of constructing a denoising network based on the noise scheduling sequence and the conditional vector, and training the denoising network using a loss function with physical constraints to obtain a physically constrained diffusion model, includes: constructing an isovariant graph neural network as the denoising network, wherein the denoising network takes the noisy extended crystal representation and time step of the current time step in the noise scheduling sequence as input, and the conditional vector as a guiding signal, to jointly predict structural noise and electronic descriptor noise; constructing a joint loss function with physical constraints, wherein the joint loss function with physical constraints consists of a main denoising loss, an atomic spacing penalty term, a bonding rule penalty term, and an electronic descriptor denoising loss, wherein the main denoising loss is the mean square error between the predicted value of the structural noise and the actual noise, the atomic spacing penalty term is used to penalize atomic spacings smaller than the minimum allowable spacing, the bonding rule penalty term is used to penalize structures that deviate from the chemical bonding rules, and the electronic descriptor denoising loss is the mean square error between the predicted value of the electronic descriptor noise and the actual noise; and training the denoising network by minimizing the joint loss function with physical constraints to obtain the physically constrained diffusion model.
[0011] Further, the step of performing inverse denoising sampling based on the physical constraint diffusion model and the condition vector, including both conditional and unconditional denoising guidance, and outputting preliminary denoising results, includes: when training the physical constraint diffusion model, replacing the condition vector with a zero vector with a preset probability for training, so that the physical constraint diffusion model supports unconditional denoising; simultaneously initializing structural variables and electronic descriptor variables with standard Gaussian noise, and performing inverse denoising sampling on the structural variables and electronic descriptor variables based on the physical constraint diffusion model and the condition vector, wherein the structural variables include fractional coordinate variables and lattice parameter variables, and the fractional coordinate variables... The lattice parameter variables are intermediate random variables used to generate the fractional coordinate matrix and the lattice parameters during the reverse denoising sampling process, respectively. At each denoising time step of the reverse denoising sampling, the currently noisy extended crystal representation and the conditional vector are received through the physical constraint diffusion model, and the constrained guided conditional denoising result and the unconstrained unconditional denoising result are calculated, respectively. Based on the conditional denoising result and the unconditional denoising result, the conditional denoising result and the unconditional denoising result are linearly interpolated according to a preset basic guiding weight, and the preliminary denoising result is output. The preset basic guiding weight is a pre-set fixed weight value used for the linear interpolation.
[0012] Further, determining the adaptive guidance weights based on the degree of matching between the preliminary denoising results and the physical constraints includes: calculating the degree of violation of the physical constraints using a constraint evaluation function based on the preliminary denoising results to obtain a constraint violation vector, wherein the constraint violation vector includes the atomic spacing constraint violation degree, coordination number constraint violation degree, and bond length distribution constraint violation degree; calculating the adaptive guidance weight vector using an adaptive guidance scheduler based on the constraint violation vector and the current denoising time step; wherein the adaptive guidance weight vector is positively correlated with the constraint violation vector and negatively correlated with the current denoising time step.
[0013] Further, the step of performing gradient correction on the preliminary denoising result according to the adaptive guiding weights to generate candidate crystal structure variables includes: calculating the gradient of the preliminary denoising result using a constraint function to obtain a gradient correction term for physical constraints; weighting and superimposing the gradient correction term onto the preliminary denoising result based on the adaptive guiding weights to obtain a corrected denoising result; inputting the corrected denoising result and the conditional vector into the denoising network to output a denoised extended crystal representation; and separating the denoised extended crystal representation to obtain candidate crystal structure variables and candidate electronic descriptor variables.
[0014] Further, the step of filtering the candidate crystal structure variables through a surrogate model and outputting near-optimal structures includes: inputting the generated candidate crystal structure variables into a preset surrogate model, predicting the target physical performance parameters of the candidate crystal structure variables through the surrogate model; and filtering the candidate crystal structure variables based on the predicted target physical performance parameters, retaining the candidate crystal structure variables that meet the preset performance threshold as near-optimal structures.
[0015] Further, the step of determining the importance weight based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions, resampling the near-optimal structure according to the importance weight to correct the distribution shift, verifying and outputting the metastable sulfide crystal structure includes: calculating the ratio of the probability density of the near-optimal structure under guided sampling distribution to the probability density under unconditional sampling distribution to obtain the importance weight; resampling the near-optimal structure according to the importance weight to correct the distribution shift and obtain the corrected near-optimal structure; verifying the physical feasibility of the corrected near-optimal structure and outputting the verified metastable sulfide crystal structure.
[0016] This invention also provides a system for the directional generation of metastable sulfide crystal structures, comprising: a data acquisition module for acquiring a crystal structure representation vector, an electronic descriptor, and target physical performance parameters of a sulfide crystal structure; mapping the preset target physical performance parameters to a conditional vector via a conditional encoder; concatenating the crystal structure representation vector with the electronic descriptor to form an extended crystal representation vector; and pairing the conditional vector with the extended crystal representation vector to generate a conditional-structure training pair; and a model training module for adding noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noise scheduling sequence; constructing a denoising network based on the noise scheduling sequence and the conditional vector; and training the denoising network using a loss function containing physical constraints to obtain physical constraints. The system comprises: a conditional diffusion model; a sampling module for performing reverse denoising sampling with and without conditional denoising guidance based on the physical constraint conditional diffusion model and the conditional vector, and outputting preliminary denoising results; a gradient correction module for determining adaptive guidance weights based on the degree of matching between the preliminary denoising results and the physical constraints, and performing gradient correction on the preliminary denoising results according to the adaptive guidance weights to generate candidate crystal structure variables; a screening and output module for screening the candidate crystal structure variables through a surrogate model and outputting near-optimal structures; determining importance weights based on the probability ratio of the near-optimal structures under guided sampling and unconditional sampling distributions, resampling the near-optimal structures according to the importance weights to correct distribution shifts, verifying and outputting metastable sulfide crystal structures.
[0017] The present invention has the following beneficial effects:
[0018] 1. By constructing a conditional diffusion model and using the target physical quantity as the guiding condition input, and training a denoising network with a physical constraint loss function, a novel metastable sulfide crystal structure was achieved through reverse customization based on performance targets, overcoming the limitation of traditional generative models that cannot generate structures in a directional manner. The generation of non-physical structures was effectively suppressed through the interatomic spacing penalty term and bonding rule penalty term in the physical constraint loss function.
[0019] 2. By using an adaptive constraint-guided sampling method, the gradient correction strength of physical constraints is dynamically adjusted according to the constraint violation degree and time step during the reverse denoising process. Weak constraints are applied in the early stage of denoising to maintain structural diversity, and strong constraints are applied in the later stage of denoising to ensure physical rationality. This solves the contradiction between diversity and rationality caused by fixed weight guidance.
[0020] 3. By using the structure-electronic descriptor joint denoising framework, the electronic descriptor is used as a physical information channel to guide the structure denoising trajectory. This improves the condition satisfaction rate and the electronic rationality of the generated structure without sacrificing diversity, and alleviates the problem of condition constraints degrading the generation quality.
[0021] 4. By using an iterative search and corrective resampling method based on the probability quality improvement index, the generative search objective is transformed from single-point optimization to improving the overall probability quality distribution of the generator in the near-optimal solution region. By correcting the distribution offset through corrective resampling, a theoretical guarantee for search efficiency is provided, and invalid sampling is avoided. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the overall process of a method for the directional generation of metastable sulfide crystal structures according to the present invention.
[0024] Figure 2 This is a flowchart illustrating the adaptive constraint-guided dynamic gradient correction sampling method of the present invention;
[0025] Figure 3 This is a schematic diagram of the structure of a metastable sulfide crystal structure directional generation system according to the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0027] Example 1
[0028] like Figure 1 As shown, the present invention provides a method for the directional generation of metastable sulfide crystal structures, comprising:
[0029] S1: Obtain the crystal structure representation vector, electronic descriptor, and target physical performance parameters of the sulfide crystal structure. Map the preset target physical performance parameters into a conditional vector using a conditional encoder. Concatenate the crystal structure representation vector and the electronic descriptor to form an extended crystal representation vector. Pair the conditional vector with the extended crystal representation vector to generate a conditional-structure training pair.
[0030] The sulfide crystal data preprocessing and conditional representation construction method uses the original sulfide crystal database as input. First, it extracts the atomic type sequence, fractional coordinate matrix, and lattice parameter six-tuple for each crystal to obtain a crystal structure representation vector. Atom types are encoded using one-hot encoding, while fractional coordinates and lattice parameters are encoded using continuous values. These three are concatenated to form the crystal structure representation vector. Based on the crystal structure representation vector, the formation energy of each crystal structure is calculated using first-principles calculations, and the room-temperature ionic conductivity is obtained through molecular dynamics simulations. These two are combined to form a physical property tag set. Local electronic descriptors, including Bader charge and atomic density of states, are extracted from each crystal structure. The crystal structure representation vector is concatenated with the electronic descriptors to form an extended crystal representation vector. User-defined target physical property parameters (such as desired ionic conductivity greater than 10 mSiemens / cm and formation energy threshold less than -0.5 eV / atom) are mapped into conditional vectors using a multilayer perceptron conditional encoder. The obtained conditional vectors are then paired with the extended crystal representation vectors to form conditional-structure training pairs.
[0031] S2: Add noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noise scheduling sequence. Construct a denoising network based on the noise scheduling sequence and the conditional vector. Train the denoising network using a loss function with physical constraints to obtain a physical constraint conditional diffusion model.
[0032] The conditional diffusion model construction and physical constraint training process combine conditional-structure training pairs to perform forward diffusion and denoising network training on the extended crystal representation vector, constructing a physically constrained conditional diffusion model. During the forward diffusion process, Gaussian noise is progressively added to the extended crystal representation vector according to a fixed noise schedule, transforming it into a standard Gaussian distribution after a preset time step. At each time step where Gaussian noise is added, the noise addition operation only affects the fractional coordinates and lattice parameters, while the atom types remain unchanged. Based on the noise schedule sequence and the conditional vector, an equivariant graph neural network is constructed as the denoising network. The denoising network takes the noisy extended crystal representation and the time step as input, and the conditional vector as an additional guiding signal, jointly predicting structural noise and electronic descriptor noise. The message passing layer of the denoising network uses equivariant graph convolution to ensure rotational and translational equivariance. The training loss function consists of four parts: the main denoising loss (mean square error between predicted and actual structural noise), an atomic spacing penalty term (penalizing structures with interatomic distances smaller than the minimum allowable spacing), a bonding rule penalty term (penalizing structures that deviate from chemical bonding rules), and an electronic descriptor denoising loss (mean square error between predicted and actual electronic descriptor noise). Training is completed by minimizing the joint loss function containing physical constraints, resulting in a physically constrained diffusion model.
[0033] The conditional control mechanism maps the expected value of the target physical performance from a scalar form to a high-dimensional conditional representation compatible with the activation space of the denoising network. This allows the denoising network to sense and respond to the performance target when estimating noise, thus biasing the inverse denoising trajectory towards the crystal structure region that satisfies the target performance. From a probabilistic perspective, the unconditional diffusion model learns the indiscriminate distribution p(x) of all sulfide crystal structures, while the conditional diffusion model learns the conditional distribution p(x|c) of the crystal structure given the performance target c. The core difference lies in the fact that the latter assigns different generation probability weights to structures in different performance ranges, giving structures in the high-performance range a higher sampling probability.
[0034] Specifically, in each message-passing layer of the E(3)-GNN, the conditional vector is mapped to a modulated signal of the same dimension as the node's hidden features through a learnable linear projection layer. This signal is then injected additively into the node feature update step. In other words, while aggregating the interaction information between atoms, the denoising network also superimposes the target performance information as a global bias onto the node features of each atom. This design allows the conditional information to continuously influence the evolution direction of the atomic features in each message-passing layer, rather than simply imposing conditional constraints at the network's input or output, thus achieving deep fusion of the conditional information. After multiple rounds of message passing, the node features of each atom simultaneously encode local atomic environment information (from neighbor aggregation) and global performance target information (from conditional vector injection), enabling the denoising head's noise prediction to consider both the local rationality of the structure and the global performance target.
[0035] The physical constraint loss function further enhances the physical plausibility of the generated structures. The interatomic spacing penalty term, by imposing a secondary loss on atomically overlapping structures during training, forces the denoising network to avoid directions that cause atoms to be too close together when predicting noise. The bonding rule penalty term, by penalizing structures whose coordination number and bond length deviate from standard values, enables the denoising network to implicitly learn chemical bonding constraints. The joint optimization of these two types of physical constraint losses with the main denoising loss allows the model to automatically follow physical rules when learning conditional distributions, thus generating physically plausible candidate structures without additional post-processing steps during the inference phase.
[0036] The model trained by the joint training of conditional control mechanism and physical constraint loss function shows significant advantages in performance screening and structural rationality. After introducing a conditional control mechanism, the proportion of generated structures that meet the target performance threshold (ionic conductivity greater than 10 mSiemens / cm and formation energy less than -0.5 eV / atom) is increased by about 8 times compared to traditional unconditional diffusion models (such as CDVAE) (from about 1% to about 8%), which can significantly reduce the amount of computation required for subsequent first-principles calculations and verification. After further superimposing the physical constraint loss function, the proportion of atomic overlap in the generated structures is reduced by about 90% compared to traditional generative adversarial networks (GANs) (from about 15% to about 1.5%), and the proportion of reasonable bond lengths (bond length deviation within ±15% of the standard value) is increased to about 92%. At the same time, the introduction of physical constraint loss does not significantly impair the diversity of generated structures. Using the Shannon entropy of spatial group distribution as a diversity indicator, the Shannon entropy of the generated structures of this invention is increased by about 20% compared to traditional variational autoencoders (VAEs) (from about 1.2 to about 1.44), indicating that the model maintains rich structural diversity while ensuring physical rationality. The Shannon entropy is used to measure the diversity of the generated crystal structure in terms of space group distribution. It is calculated based on the frequency distribution of each space group in the generated structure set. The larger the Shannon entropy value, the richer and more uniform the space group types covered by the generated structure are, that is, the higher the structural diversity.
[0037] When designing solid-state electrolyte materials, the target performance parameters were set as follows: ionic conductivity greater than 20 mSiemens / cm and formation energy gap convexity less than 0.08 eV / atom. These parameters were encoded into a 128-dimensional conditional vector using a conditional encoder and input into a physical constraint diffusion model for inverse denoising sampling to generate candidate sulfide crystal structures. After rapid screening using a surrogate model based on a graph neural network, the proportion of near-optimal structures satisfying the above conditions was significantly higher than that in the unconditional sampling case, indicating that the conditional control mechanism can effectively improve the sampling efficiency of the target region.
[0038] S3: Based on the physical constraint diffusion model and the condition vector, perform reverse denoising sampling that includes conditional denoising guidance and unconditional denoising guidance, and output preliminary denoising results.
[0039] Based on a pre-trained physical constraint diffusion model, the target performance condition vector is introduced into the inverse denoising sampling process to generate preliminary denoising results. The condition-guided inverse denoising sampling method takes the physical constraint diffusion model and the target performance condition vector as input. It starts with standard Gaussian noise, simultaneously initializing structural and electronic descriptor variables, and then performs inverse denoising sampling. At each denoising time step, the denoising network receives the current noisy extended crystal representation and the condition vector, calculating both conditional and unconditional denoising results. A classifier-guided mechanism enhances conditional control: during training, the condition vector is replaced with a zero vector with a preset probability; during sampling, the conditional and unconditional denoising results are linearly interpolated according to the basic guiding weights, outputting the preliminary denoising result.
[0040] S4: Determine the adaptive guiding weights based on the degree of matching between the preliminary denoising results and the physical constraints, and perform gradient correction on the preliminary denoising results according to the adaptive guiding weights to generate candidate crystal structure variables.
[0041] Based on the degree of violation of physical constraints, the preliminary denoising result is adaptively gradient-corrected to generate candidate crystal structure variables. The adaptive constraint-guided dynamic gradient correction sampling method takes the matching degree between the preliminary denoising result and the physical constraints as input. First, it calculates the degree of violation of each physical constraint through the constraint evaluation function, obtaining a constraint violation vector. Based on the constraint violation vector and the current denoising time step, an adaptive guiding scheduler calculates adaptive guiding weights. When the denoising time step is large (high noise stage), the guiding weights decay cosinely to a smaller value; when the denoising time step is small (low noise stage), the guiding weights increase cosinely to a larger value. Constraints with high violation degrees are automatically weighted, while constraints that are satisfied are weighted less. Based on the adaptive guiding weights, the gradient correction terms of each constraint are weighted and superimposed onto the preliminary denoising result to obtain the corrected denoising result. The corrected denoising result and the condition vector are input into the denoising network, which outputs the denoised extended crystal representation, separating the candidate crystal structure variables and candidate electronic descriptor variables.
[0042] The degree of matching between the preliminary denoising result and the physical constraints is quantified and calculated using a matching degree function, the formula of which is:
[0043]
[0044] Where M(x) is the matching degree value of the initial denoising result x, and the value range is (0,1]. The closer the matching degree value is to 1, the higher the matching degree with the physical constraints. To constrain the degree vector of violation; Represents the L2 norm operator; Represents an exponential function; The degree of violation of the interatomic spacing constraint; The degree of violation of the coordination number constraint; Let represent the constraint violation degree of the bond length distribution. This formula quantifies the degree of matching by mapping the negative exponential of the L2 norm of the constraint violation degree. When the constraint violation degree is zero, the degree of matching is 1, indicating a perfect match. As the constraint violation degree increases, the degree of matching decays exponentially and approaches 0.
[0045] S5: The candidate crystal structure variables are screened by a surrogate model to output the near-optimal structure; the importance weight is determined based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions, and the near-optimal structure is resampled according to the importance weight to correct the distribution shift, and the metastable sulfide crystal structure is verified and output.
[0046] Candidate crystal structure variables are screened and corrected using a surrogate model, and the final metastable sulfide crystal structure is verified and output. A probability-quality-enhanced guided efficient search and convergence guarantee method uses candidate crystal structure variables as input. First, the candidate crystal structure variables are input into a pre-defined surrogate model for rapid evaluation, predicting the formation energy and ionic conductivity. Structures that simultaneously satisfy the metastable condition and performance target are selected as near-optimal structures. The importance weight is determined based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions. Near-optimal structures are then resampled according to their importance weights to correct the distribution offset caused by guided sampling. The physical feasibility of the resampled structures is verified, including accurate evaluation of the formation energy and ionic conductivity through first-principles calculations and verification of thermodynamic stability through molecular dynamics simulations. Verified structures must simultaneously satisfy the metastable condition (formation energy distance from the convex hull 0 to 0.1 eV / atom) and the performance target (ionic conductivity greater than 10 mSiemens / cm) to be output as the final metastable sulfide crystal structure.
[0047] Example 2
[0048] Based on Example 1, the acquisition of the crystal structure representation vector, electronic descriptor, and target physical property parameters of the sulfide crystal structure includes:
[0049] S1.1: Extract the atomic type sequence, fractional coordinate matrix, and lattice parameters of the crystal from the database of sulfide crystal structures.
[0050] After obtaining the raw data from the sulfide crystal structure database, it is necessary to extract the atom type sequence, fractional coordinate matrix, and lattice parameters for each crystal to obtain a unified crystal structure representation vector. The atom type sequence describes the element type of the atom at each position in the crystal, and one-hot encoding is used to map the element type to a binary vector, ensuring that the representations of different elements are orthogonal. The fractional coordinate matrix describes the relative position of each atom in the lattice, and the six-tuple lattice parameters describe the geometric dimensions and angles of the lattice; both are encoded using continuous values to maintain their numerical precision.
[0051] S1.2: Perform one-hot encoding on the atomic type sequence, perform continuous value encoding on the fractional coordinate matrix and the lattice parameters, and concatenate the encoding results of the one-hot encoding and the continuous value encoding to obtain the crystal structure representation vector.
[0052] The one-hot encoding result is concatenated with the continuous value encoding result to obtain a fixed-dimensional crystal structure representation vector, which facilitates subsequent neural network processing. The dimension of the crystal structure representation vector is jointly determined by the atomic type encoding dimension, fractional coordinate dimension, and lattice parameter dimension, ensuring that the structural information of the crystal can be fully expressed.
[0053] S1.3: Extract local electronic descriptors from each crystal structure in the crystal structure representation vector, the local electronic descriptors including Bader charge and atomic density of states.
[0054] The extraction of local electronic descriptors uses each crystal structure in the crystal structure representation vector as input, and calculates the Bader charge and atomic density of states for each atom using first-principles calculations. The Bader charge reflects the amount of valence electron transfer in the crystal, and the atomic density of states describes the characteristic vector of the electronic density of states of each atom near the Fermi level. The crystal structure representation vector is concatenated with the electronic descriptor to form an extended crystal representation vector, providing electronic structure information for subsequent joint denoising.
[0055] Specifically, Bader charge calculation is based on the Bader charge analysis method. It uses the electron density distribution calculated from first-principles calculations as input, determines the Bader volume for each atom by dividing the electron density into zero-flux surfaces, and integrates the electron density within the Bader volume to obtain the Bader charge value for that atom. The numerical meaning of Bader charge is as follows: a positive value indicates that the atom has lost electrons (characteristic of a cation), and a negative value indicates that the atom has gained electrons (characteristic of an anion). Its absolute value reflects the degree of ionicity of the atom in the crystal's chemical bonding. For sulfide crystals, sulfur atoms typically exhibit a negative Bader charge, while metal cations exhibit a positive Bader charge, and the distribution pattern of Bader charge is closely related to the ionic conductivity mechanism of the crystal.
[0056] The calculation of the atomic density of states is also based on first-principles calculations. Using the Kohn-Sham orbitals after convergence of the self-consistent field iteration as a foundation, the total density of states is decomposed into the contribution of each atom through a projection operation, resulting in an atomically resolved density of states spectrum. In this embodiment of the invention, the atomic density of states uses the Fermi level as a reference zero point, extracting the density of states eigenvector within a ±2 eV range near the Fermi level. The eigenvector has a fixed dimension of 64, obtained through uniform sampling within this energy window. The atomic density of states eigenvector encodes the local contribution of each atom to the electronic structure of the crystal, which is of great significance for evaluating the electronic rationality of the crystal.
[0057] By concatenating the Bader charge (scalar) and atomic density of states eigenvector (64-dimensional vector) of each atom, a local electronic descriptor vector (65-dimensional) for each atom is formed. Taking a crystal containing N atoms as an example, its local electronic descriptor matrix has a shape of N×65, corresponding one-to-one with the fractional coordinate matrix (N×3) and the atom type matrix (N×number of elements) in the atomic dimension, facilitating atomic-level information fusion in the extended crystal representation vector. By introducing local electronic descriptors, the extended crystal representation vector not only contains the geometric structure information of the crystal but also the electronic structure information directly related to ionic conductivity, providing a richer physical information channel for the denoising network.
[0058] S1.4: The formation energy of each crystal structure in the crystal structure representation vector is obtained by first-principles calculation, and the room temperature ionic conductivity is obtained by molecular dynamics simulation. The formation energy of each crystal structure and the room temperature ionic conductivity are combined to form a physical property tag set.
[0059] When constructing the physical performance tag set, first-principles calculations and molecular dynamics simulations were employed. First-principles calculations were used to obtain the formation energy of each crystal structure, defined as the energy difference between the crystal structure and the stable phase of its constituent elements; a lower formation energy indicates a more stable crystal structure. Molecular dynamics simulations were used to obtain the room-temperature ionic conductivity, a key indicator for evaluating the performance of sulfide crystals as solid-state electrolytes. The formation energy and room-temperature ionic conductivity of each crystal structure were combined to form the physical performance tag set, which served as the target for conditional vector mapping.
[0060] Example 3
[0061] Based on Example 1, after obtaining the physical performance tag set and crystal structure representation vector of Example 2, the following steps are further included: mapping the preset target physical performance parameters into a conditional vector via a conditional encoder; concatenating the crystal structure representation vector with the electronic descriptor to form an extended crystal representation vector; and pairing the conditional vector with the extended crystal representation vector to generate a conditional-structure training pair.
[0062] S1.5: The preset target physical performance parameters are mapped into condition vectors through a multilayer perceptron conditional encoder, wherein the multilayer perceptron conditional encoder is composed of alternating stacked fully connected layers, batch normalization layers and activation layers.
[0063] The target physical performance parameters are mapped into conditional vectors using a multilayer perceptron conditional encoder. The multilayer perceptron conditional encoder consists of alternating stacked fully connected layers, batch normalization layers, and activation layers. Its function is to map target performance parameters of arbitrary dimensions into conditional vectors of fixed dimensions to align with the input space of the denoising network. Specifically, the multilayer perceptron conditional encoder contains three fully connected layers with hidden dimensions of 256 and 128 respectively, and an output dimension of 128. Each fully connected layer is followed by a batch normalization layer and a ReLU activation function. The output dimension of the conditional encoder matches the dimensions of the intermediate layers of the denoising network, ensuring that the conditional information effectively guides the denoising process.
[0064] S1.6: Concatenate the crystal structure representation vector with the electronic descriptor according to their dimensions to obtain the extended crystal representation vector.
[0065] By concatenating along the feature axis, the crystal structure representation vector and the electronic descriptor are fused into an extended crystal representation vector. This extended crystal representation vector adds electronic structure information to the original crystal structure representation, enabling the denoising network to consider electronic plausibility while generating the crystal structure. The concatenation operation is performed along the feature axis, forming a higher-dimensional joint representation.
[0066] S1.7: Pair the conditional vector with the extended crystal representation vector according to the batch dimension to form a conditional-structure training pair.
[0067] Before model training, the conditional vector and the extended crystal representation vector need to be paired along the batch dimension to generate conditional-structure training pairs. In each training sample, the conditional vector describes the desired target performance, and the extended crystal representation vector describes the actual crystal structure. These paired pairs are used to train the conditional diffusion model. During training, the model learns the probability distribution of the extended crystal representation vector given the conditional vector, thereby achieving the directed generation of crystal structures based on the performance target.
[0068] Example 4
[0069] Based on Example 1, the step of adding noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noisy scheduling sequence includes:
[0070] S2.1: Gaussian noise is gradually added to the extended crystal representation vector in the conditional-structure training pair according to a preset noise schedule. The addition of Gaussian noise only affects the fractional coordinate matrix and lattice parameters, while the atom type remains unchanged.
[0071] According to a pre-defined noise scheduling table, Gaussian noise is progressively added to the extended crystal representation vector in the conditional-structure training pair. The noise scheduling table defines the variance scheduling strategy for the Gaussian noise added at each time step, employing cosine scheduling, and setting the total number of denoising time steps to 1000. During the forward pass of the diffusion model, as the time steps increase, the structural information in the extended crystal representation vector is gradually masked by noise, eventually approaching a standard Gaussian distribution. The addition of Gaussian noise only affects the fractional coordinate matrix and lattice parameters; the atom types remain unchanged because atom types are highly discrete categorical variables, and directly adding continuous noise to them would destroy their semantic meaning.
[0072] It should be noted that the noise variance scheduling table for cosine scheduling is defined as: at time step t, the cumulative noise variance... Where s is the offset coefficient with a value of 0.008, T is the total number of denoising time steps, and π is the mathematical constant pi. Compared with linear scheduling, cosine scheduling increases noise more slowly in the early stage of diffusion, avoiding premature loss of crystal structure information due to excessively rapid noise growth in early time steps; in the later stage of diffusion, the noise variance approaches 1, making the noisy state of the final time step sufficiently close to the standard Gaussian distribution, ensuring that the starting condition of the reverse denoising sampling satisfies the theoretical assumptions of the diffusion model.
[0073] The addition of Gaussian noise only affects the fractional coordinate matrix and lattice parameters, while the atom type remains unchanged. This design is based on the representational characteristics of crystal structures: the fractional coordinates and lattice parameters are continuous-valued variables, suitable for diffusion processes in continuous space; however, the atom type, after one-hot encoding, is a discrete categorical variable. Directly adding continuous Gaussian noise to it would destroy the class distinguishability, causing the denoising network to fail to correctly recover the original element type. By fixing the atom type and performing diffusion only in the coordinate space, the denoising network only needs to learn to generate reasonable atomic arrangements and lattice geometry given the elemental composition, which is highly consistent with the goal of conditional generation. In addition, fixing the atom type implicitly constrains the chemical space of the generated structure, ensuring that the elemental composition of all generated structures comes from the sulfide chemical systems already appearing in the training database, avoiding the generation of unreasonable element combinations.
[0074] S2.2: After a preset time step, the extended crystal representation vector is transformed into a standard Gaussian distribution to obtain a noise scheduling sequence.
[0075] As the noise addition process progresses, after a preset time step, the extended crystal representation vector is completely noise-enhanced and transformed into a standard Gaussian distribution, thus forming a noise scheduling sequence. This noise scheduling sequence records the complete evolution trajectory from the original crystal structure to a completely noisy state, forming the basis for training the denoising network. In the reverse denoising process, the denoising network learns the ability to recover the original crystal structure from any intermediate state in the noise scheduling sequence, with conditional vectors guiding the recovery process towards meeting the target performance.
[0076] Example 5
[0077] Based on Example 1, the step of constructing a denoising network based on the noise scheduling sequence and the conditional vector, training the denoising network using a loss function with physical constraints, and obtaining a physical constraint-based diffusion model further includes:
[0078] S2.3: Construct an equivariant graph neural network as the denoising network, wherein the denoising network takes the noisy extended crystal representation and time step of the current time step in the noise scheduling sequence as input, and the conditional vector as the guiding signal, and jointly predicts structural noise and electronic descriptor noise.
[0079] Equivariant graph neural network (E(3)-GNN) is used as a denoising network to ensure that the neural network is equivariant to rotation and translation transformations of the input crystal structure, that is, for any rotation matrix R and translation vector v, it satisfies This ensures that the generated crystal structure is not affected by the choice of coordinate system and meets the requirements of physical symmetry. The E(3)-GNN contains four message passing layers with a hidden dimension of 128. Each message passing layer uses equivariant graph convolution operation, with node features encoding atomic element attributes and edge features encoding inter-atomic distance information. The denoising network takes the denoised extended crystal representation and time step of the noise scheduling sequence as input and the conditional vector as the guiding signal to jointly predict structural noise and electronic descriptor noise. The network uses a shared E(3)-GNN backbone and is divided into two prediction heads after the message passing layers: a structural denoising head predicts the noise of fractional coordinates and lattice parameters, and an electronic denoising head predicts the noise of Bader charge and atomic density of states.
[0080] Specifically, E(3)-GNN can treat each atom in the crystal as a node in a graph, and treat the chemical bonds between atoms or pairs of atoms with a distance smaller than the cutoff radius (6 Å in this embodiment) as edges in the graph, thus constructing a crystal graph. The node features are initialized as the concatenation of the element embedding vector of the atom (mapped to a 32-dimensional continuous vector by one-hot encoding through a learnable embedding layer) and the time step embedding vector (mapped to a 32-dimensional vector by sinusoidal position encoding), for a total of 64 dimensions. The edge features are initialized as the concatenation of the inter-atomic distance obtained by expanding the radial basis function (16 Gaussian radial basis functions are used in this embodiment) into a 16-dimensional continuous vector, and the inter-atomic direction vector (3-dimensional unit vector), for a total of 19 dimensions.
[0081] In each message passing layer, each node aggregates information from its neighboring nodes. The aggregation operation satisfies the equivariance property E(3), meaning that scalar features (such as energy and charge) remain unchanged under coordinate system rotation (invariants), and vector features (such as force and gradient) are transformed by the same rotation matrix under coordinate system rotation (equivariants). Specifically, the message function takes the features of the sending node, the features of the receiving node, and the edge features as inputs, and outputs scalar messages and vector messages; the aggregation function sums the scalar messages of all neighbors and sums the vector messages of all neighbors by direction weight; the update function fuses the aggregated messages with the node's own features and updates the scalar features and vector features respectively through a gating mechanism.
[0082] In crystal structure generation tasks, the introduction of E(3) equivariance is crucial. The physical properties of a crystal structure are independent of the chosen coordinate system; the formation energy and ionic conductivity of the same crystal should remain completely consistent when placed in different orientations. If the denoising network does not satisfy equivariance, the model may identify the same crystal structure in different coordinate systems as different samples, leading to redundancy in training data and a lack of geometric consistency in the generated structures. By employing E(3)-GNN, this invention ensures the equivariant response of the denoising network to arbitrary rotation and translation transformations, thereby fully utilizing the symmetry information in the training data without data augmentation, improving the model's generalization ability and the physical rationality of the generated structures.
[0083] S2.4: Construct a joint loss function with physical constraints, wherein the joint loss function with physical constraints consists of a main denoising loss, an atomic spacing penalty term, a bonding rule penalty term, and an electronic descriptor denoising loss. The main denoising loss is the mean square error between the predicted value of the structural noise and the actual noise. The atomic spacing penalty term is used to penalize atomic spacings smaller than the minimum allowable spacing. The bonding rule penalty term is used to penalize structures that deviate from the chemical bonding rules. The electronic descriptor denoising loss is the mean square error between the predicted value of the electronic descriptor noise and the actual noise.
[0084] The joint loss function with physical constraints is constructed by taking the main denoising loss, interatomic spacing penalty term, bonding rule penalty term, and electronic descriptor denoising loss as inputs, and then summing the four parts of the loss according to their weights to obtain the joint loss function. The main denoising loss is the mean square error between the predicted value of the structural noise and the actual noise, and its calculation formula is as follows:
[0085]
[0086] in, Mainly to reduce noise loss, This represents real structural noise. For the structural noise predicted by the denoising network, x t Let c be the noisy extended crystal representation at time step t, where t is the current denoising time step and c is the condition vector. For the parameters of the denoising network, For mathematical expectation operators.
[0087] Based on the main denoising loss, a joint loss function containing physical constraints is further constructed, and the calculation formula is as follows:
[0088]
[0089] in, For the joint loss function, Mainly to reduce noise loss, This is a penalty term for interatomic spacing. For key rule penalty items, λ1 represents the electronic descriptor denoising loss, λ2 represents the weighting coefficient corresponding to the interatomic spacing penalty term, λ3 represents the weighting coefficient corresponding to the bonding rule penalty term, and λ4 represents the weighting coefficient corresponding to the electronic descriptor denoising loss term.
[0090] When setting the weights of the joint loss function, the main denoising loss, as the core driver of the model's learning of the structural distribution, has a fixed weight of 1. Considering that atomic overlap will directly lead to divergence in structural force calculations, and its harmfulness is greater than deviation from chemical bonding rules, the weight of the interatomic spacing penalty term λ1 (0.5) is greater than the weight of the bonding rule penalty term λ2 (0.3). At the same time, to avoid excessively high weights interfering with the convergence of the main structural denoising, the weight λ3 of the electronic descriptor denoising loss, which serves as an auxiliary generation target, is set to 0.2.
[0091] The weight coefficients mentioned above are preferred values in this embodiment of the invention. Those skilled in the art can adjust the weight coefficients according to the scale and quality of the actual training data. In the actual parameter tuning process, it is recommended to pre-train the denoising network for several rounds under the unconstrained condition of λ1=λ2=λ3=0, and then gradually introduce physical constraint penalty terms after the main denoising loss converges. The weights of each penalty term should be gradually increased from small values (such as 0.1) to avoid excessive constraint penalties that could cause the denoising loss to diverge. The effectiveness of the constraint penalty terms can be evaluated by the proportion of non-physical structures (atomic overlap rate, bond length anomaly rate) on the validation set, and the weight coefficients can be adjusted accordingly.
[0092] The interatomic spacing penalty term penalizes interatomic spacing smaller than the minimum permissible spacing, achieved by applying a secondary penalty to structures with interatomic distances smaller than the minimum permissible spacing. The bonding rule penalty term penalizes structures that deviate from chemical bonding rules, calculated by comparing the degree of deviation of the coordination number and bond length distributions of the generated structure from known chemical rules. The electronic descriptor denoising loss is the mean square error between the predicted value of the electronic descriptor noise and the actual noise, and has the same form as the main denoising loss.
[0093] S2.5: Train the denoising network by minimizing the joint loss function containing physical constraints to obtain the physical constraint diffusion model.
[0094] The training data for E(3)-GNN comes from sulfide crystal entries in the Materials Project Database and the International Crystal Structure Database, totaling approximately 5000 sulfide crystal structures. These are divided into a training set of 4000 structures, a validation set of 500 structures, and a test set of 500 structures in an 8:1:1 ratio. The parameters are set as follows: learning rate 0.001, batch size 64, training epochs 500, AdamW optimizer, weight decay 0.01, and dropout rate 0.1. The weight coefficients λ1, λ2, and λ3 are set to 0.5, 0.3, and 0.2, respectively. The denoising network is trained by minimizing the joint loss function containing physical constraints to obtain the physical constraint diffusion model.
[0095] Example 6
[0096] Based on Example 1, the reverse denoising sampling, which includes both conditional and unconditional denoising guidance, is performed based on the physical constraint diffusion model and the condition vector to output preliminary denoising results, including:
[0097] S3.1: When training the physical constraint diffusion model, the condition vector is replaced with a zero vector with a preset probability for training, so that the physical constraint diffusion model supports unconditional denoising.
[0098] The classifier's free-guided training method takes the training process of a physically constrained diffusion model as input, and replaces the condition vector with a zero vector with a preset probability during training. Those skilled in the art will understand that setting the preset probability to 0.1 to 0.2, by randomly discarding conditional information during training, allows the model to learn both conditional and unconditional distributions simultaneously, enabling flexible control of the strength of conditional guidance during the sampling phase. If conditional vectors are always used for training, the model will be unable to perform effective unconditional generation during sampling, limiting the control space of conditional guidance.
[0099] The condition vector is randomly discarded within each training batch. For each sample in the batch, a preset probability (0.15 in this embodiment) is used to determine whether its condition vector should be replaced with a zero vector. The replacement operation is performed after the conditional encoder output and before the denoising network receives the data; that is, the condition vector of the sample is covered with an all-zero vector with the exact same dimension as the condition vector. This ensures that each training batch contains both conditional samples (conditional vectors encoding the target performance) and unconditional samples (conditional vectors with zero vectors), allowing the model to receive both types of supervision signals in a single forward propagation without requiring additional training loops.
[0100] Specifically, the zero vector has a clear semantic meaning in the conditional coding space, namely, the "unconditional" prior state. If replaced with a random vector, the model will have difficulty distinguishing between "unconditional guidance" and "guidance based on a specific but unknown condition," leading to instability in unconditional distribution learning. During the sampling phase, the zero vector is also used as the conditional input for unconditional denoising, consistent with the training phase, ensuring the training-inference consistency of the classifier's free guidance mechanism. Furthermore, the preset probability value also affects the final generation effect. A probability that is too low (e.g., less than 0.05) will result in insufficient unconditional generation capability, limiting the adjustable range of guidance weights during sampling; a probability that is too high (e.g., greater than 0.3) will weaken the learning quality of the conditional distribution and reduce the accuracy of conditional guidance. This embodiment of the invention selects 0.15 as the balance point, achieving the best comprehensive index of condition satisfaction rate and structural diversity on the validation set.
[0101] S3.2: Initialize the structural variables and electronic descriptor variables simultaneously with standard Gaussian noise. Based on the physical constraint diffusion model and the condition vector, perform inverse denoising sampling on the structural variables and the electronic descriptor variables. The structural variables include fractional coordinate variables and lattice parameter variables, which are intermediate random variables used to generate the fractional coordinate matrix and the lattice parameter during the inverse denoising sampling process.
[0102] The structural and electronic descriptor variables are initialized using standard Gaussian noise. The structural variables include fractional coordinates and lattice parameters, which are intermediate random variables used to generate the fractional coordinate matrix and lattice parameters during the inverse denoising sampling process. The electronic descriptor variables include Bader charge and atomic density of states. All variables are initialized to standard Gaussian noise of the same dimension as the extended crystal representation vector. The sampling process involves progressively denoising from this noise state to ultimately restore a valid crystal structure representation.
[0103] S3.3: At each denoising time step of the reverse denoising sampling, the currently noisy extended crystal representation and the condition vector are received through the physical constraint diffusion model, and the constrained guided conditional denoising result and the unconstrained unconditional denoising result are calculated respectively.
[0104] Linear interpolation of conditional and unconditional denoising results takes the conditional and unconditional denoising results as inputs, performs linear interpolation on them according to preset basic guiding weights, and outputs a preliminary denoising result. The linear interpolation formula is:
[0105]
[0106] in, This is the initial denoising result obtained from single-time-step interpolation; This refers to the unconstrained and unconditional denoising results output by the model. This represents the conditional denoising result obtained by combining physical constraints and conditional vector guidance. `g` is the preset basic guidance weight, with a value range of [0,5], preferably [1,3]. A larger weight value results in a more significant guidance effect from the physical conditions; however, an excessively large weight can lead to a decrease in the diversity and physicochemical rationality of the generated crystal structures. `t` is the current denoising time step in the reverse denoising sampling process. A larger basic guidance weight `g` results in a stronger conditional guidance effect, but an excessively large `g` can lead to a decrease in the diversity and rationality of the generated structures.
[0107] S3.4: Based on the conditional denoising result and the unconditional denoising result, perform linear interpolation on the conditional denoising result and the unconditional denoising result according to the preset basic guiding weight, and output the preliminary denoising result. The preset basic guiding weight is a pre-set fixed weight value used for the linear interpolation.
[0108] At each denoising time step of the reverse denoising sampling, the currently noisy extended crystal representation and condition vector are received through a physically constrained diffusion model. Constrained guided conditional denoising results and unconstrained unconditional denoising results are then calculated, respectively. The conditional denoising result is the model's predicted denoising result given the condition vector, while the unconditional denoising result is the model's predicted denoising result when the condition vector is zero. Combining the two using a linear interpolation formula achieves a balance between the strength of the conditional guidance and the diversity of generation.
[0109] Example 7
[0110] like Figure 2 As shown, based on Example 1, the step of determining the adaptive guiding weights based on the degree of matching between the preliminary denoising results and physical constraints includes:
[0111] S4.1: Based on the preliminary denoising results, the degree of violation of physical constraints is calculated through the constraint evaluation function to obtain the constraint violation vector, which includes the interatomic spacing constraint violation degree, coordination number constraint violation degree, and bond length distribution constraint violation degree.
[0112] The constraint evaluation function calculates the degree of violation of physical constraints based on the preliminary denoising results, determining the constraint violation vector. The constraint violation vector includes the interatomic spacing constraint violation, coordination number constraint violation, and bond length distribution constraint violation. The interatomic spacing constraint violation is calculated by statistically analyzing the proportion of atomic pairs with spacings smaller than the minimum permissible spacing to the total number of atomic pairs. The bond length distribution constraint violation is calculated by comparing the bond length distribution of the desired structure with the standard bond length distribution of similar known compounds using the Kohlbek-Leibler divergence. The specific formula for the constraint evaluation function is as follows:
[0113]
[0114] in, To constrain the degree vector of violation; This is the preliminary denoising result; The degree of violation of the interatomic spacing constraint; The degree of violation of the coordination number constraint; The degree of violation of the bond length distribution constraint.
[0115] The degree of violation of the coordination number constraint The degree of deviation is calculated by comparing the actual coordination number of each atom with the expected range of coordination numbers in elemental chemistry. The specific calculation formula is as follows:
[0116]
[0117] Where N is the total number of atoms in the crystal; C i The actual coordination number of the i-th atom; is the expected standard coordination number of the element to which the i-th atom belongs in elemental chemistry; The allowable range of coordination number deviation; To perform the operation to find the maximum value.
[0118] S4.2: Based on the constraint violation vector and the current denoising time step, calculate the adaptive guidance weight vector through the adaptive guidance scheduler; wherein the adaptive guidance weight vector is positively correlated with the constraint violation vector and negatively correlated with the current denoising time step.
[0119] The adaptive guidance scheduler takes the constraint violation vector and the current denoising time step as input to calculate the adaptive guidance weight vector. The adaptive guidance weight vector is positively correlated with the constraint violation vector; that is, constraints with higher violation degrees receive larger guidance weights to focus on correcting constraints with severe violations. The adaptive guidance weight vector is negatively correlated with the current denoising time step; that is, in the early stages of denoising (larger denoising time steps, high noise stage), the guidance weights are smaller, allowing the structure to explore diverse solution spaces under more relaxed constraints; in the later stages of denoising (smaller denoising time steps, low noise stage), the guidance weights are larger to ensure that the final generated structure strictly satisfies the physical constraints. The specific formula for calculating the adaptive guidance weight is as follows:
[0120]
[0121] in, For the k-th constraint, the adaptive guiding weights are given at time step t. The maximum value of the guiding weight (ranging from 0.1 to 1.0) is given, where t is the current denoising time step and T is the total number of denoising time steps. This is the scaling factor for the sigmoid function (logic stearic function). Let be the violation degree of the k-th constraint. The formula consists of two parts: a cosine scheduling term and a sigmoid adaptive term. The cosine scheduling term achieves a smooth decay with the denoising time step, while the sigmoid adaptive term achieves a nonlinear response to the violation degree.
[0122] The adaptive guidance weight vector is positively correlated with the constraint violation vector, and negatively correlated with the current denoising time step. Specifically, the aforementioned dual monotonicity relationship jointly determines the temporal distribution pattern of constraint guidance intensity throughout the denoising process. In the high-noise stage with a large denoising time step, the cosine scheduling term outputs a weight value close to zero. At this time, even if the constraint violation is high, the overall guidance weight remains at a low level, allowing the structure to explore a wide area under relaxed constraints and preserving generation diversity. In the low-noise stage with a small denoising time step, the cosine scheduling term outputs a weight value close to the maximum guidance intensity. At this time, the level of constraint violation directly determines the correction intensity of each constraint, ensuring that the final output structure strictly meets the physical rationality requirements. The coupled design of the two monotonicities avoids the collapse of structural diversity due to excessive constraints in the early stage of denoising, while avoiding the leakage of non-physical structures due to excessively weak constraints in the later stage of denoising.
[0123] The adaptive constraint-guided sampling mechanism is based on the characteristic of signal-to-noise ratio (SNR) changing with time step during diffusion model denoising. In the early stage of denoising (time step t is close to T), the signal component in the noisy crystal structure is extremely weak, and the noise component dominates. At this stage, the denoising network mainly determines the overall topology of the crystal (space group type, atomic composition mode), and the generated decisions have the greatest impact on the diversity of the final structure. In the later stage of denoising (time step t is close to 0), the SNR in the noisy crystal structure is high. At this stage, the denoising network mainly refines the local details of atomic positions and lattice parameters, which has the greatest impact on the physical rationality of the final structure. Applying strong constraints in the early stage of denoising will lead to premature locking of the structural topology type, causing all generated structures to tend to have similar space groups and atomic arrangements, i.e., diversity collapse. Applying weak constraints in the later stage of denoising allows unreasonable atomic spacing and bond lengths to be retained in the final structure, reducing the quality of generation.
[0124] The combined design of the cosine decay term and the sigmoid adaptive term has unique advantages. Compared with linear scheduling, cosine scheduling changes more smoothly at both ends of the time step (t close to T and t close to 0), and transitions more rapidly in the middle stage, conforming to the gradual constraint strategy of "relaxed in the early stage, transitioning in the middle stage, and strict in the later stage". The nonlinear response of the sigmoid function to the constraint violation degree allows constraints that have been satisfied (violation degree close to 0) to automatically withdraw from gradient correction, avoiding unnecessary perturbation to reasonable regions, while constraints with high violation degree (violation degree much greater than 0) receive near-linear guidance, realizing the automatic allocation of constraint correction resources. The product form of the two terms ensures the decoupling of time-series scheduling and violation degree response, allowing the two hyperparameters α_max (controlling the overall constraint strength) and β (controlling the violation degree sensitivity) to be adjusted independently during parameter tuning without redesigning the scheduling function.
[0125] Unlike the original fixed-weight classifier-free guidance (CFG) mechanism, the guidance strength of fixed-weight CFG remains constant throughout the denoising process. However, the selection of the guidance weight *w* in the fixed-weight mechanism presents a challenge: too small a guidance weight results in insufficient condition satisfaction, while too large a weight reduces diversity and structural rationality. Furthermore, the optimal guidance weight value is significantly influenced by specific performance targets and dataset distribution, requiring retuning for each new generation task. The fixed-weight CFG mechanism's constant guidance strength throughout the denoising process cannot adapt to the structural generation requirements under varying noise levels. In the high-noise stage of initial denoising, excessively strong fixed constraints can cause the model to converge prematurely to local optima, compromising the diversity of generated structures. In the low-noise stage of later denoising, a small fixed weight may fail to effectively correct subtle structural defects that deviate from physical rules.
[0126] Adaptive constraint-guided sampling automatically adapts to the constraint requirements at different stages through a mechanism that dynamically schedules sampling based on the denoising time steps and constraint violation rate, achieving a smooth transition from relaxed exploration to strict constraints. Taking a specific denoising sampling process as an example, let the total time steps T = 1000 and the maximum guiding strength α... max =0.8, sensitivity coefficient β=5. At time step 800 (t / T=0.8), the cosine scheduling term outputs approximately 0.19. At this point, even if the violation of a certain interatomic spacing constraint is 0.6, the adaptive guiding weight of this constraint is only approximately 0.8×0.19×sigmoid(5×0.6)≈0.8×0.19×0.82≈0.12, indicating a very small guiding correction magnitude. The structure can freely explore different topology types at this stage. At time step 200 (t / T=0.2), the cosine scheduling term outputs approximately 0.81. Under the same conditions, the adaptive guiding weight increases to approximately 0.8×0.81×0.82≈0.53, significantly enhancing the guiding correction magnitude and effectively correcting the structure to a reasonable region that satisfies the interatomic spacing constraint. At time step 50 (t / T=0.05), the cosine scheduling term outputs approximately 0.97, and the adaptive guiding weight further increases to approximately 0.8×0.97×0.82≈0.64. The strongest constraint is applied in the final stage of denoising, ensuring the physical rationality of the final output structure.
[0127] Example 8
[0128] Based on Example 1, the step of performing gradient correction on the preliminary denoising result according to the adaptive guiding weights to generate candidate crystal structure variables includes:
[0129] S4.3: Calculate the gradient of the preliminary denoising result using the constraint function to obtain the gradient correction term of the physical constraint. Based on the adaptive guiding weight, weight the gradient correction term and add it to the preliminary denoising result to obtain the corrected denoising result.
[0130] The constraint function uses mathematical expressions to describe each physical constraint. Gradient correction terms for the physical constraints are calculated based on the preliminary denoising results; these gradient correction terms point in the direction that reduces the degree of constraint violation. The gradient correction terms are obtained by taking the partial derivative of the constraint function with respect to the current crystal structure parameters (fractional coordinates, lattice parameters). Based on adaptive guiding weights, the gradient correction terms are weighted and superimposed onto the preliminary denoising results to obtain the corrected denoising result. This weighted superposition operation applies a larger correction force to constraints with high violation rates and a smaller correction force to constraints with low violation rates or those already satisfied, achieving adaptive constraint guidance.
[0131] S4.4: Input the corrected denoising result and the conditional vector into the denoising network, and output the denoised extended crystal representation.
[0132] The denoising network receives the corrected denoising result and conditional vector, processes them, and outputs the denoised extended crystal representation. Upon receiving the corrected intermediate result, the denoising network uses it as the output of the current time step and passes it to the next denoising time step. The entire inverse denoising process consists of a preset number of time steps. Each time step performs linear interpolation of the conditional and unconditional denoising results, adaptive constraint-guided weight calculation, and gradient correction operations, forming a complete inverse denoising sampling process.
[0133] S4.5: Separate the denoised extended crystal representation to obtain candidate crystal structure variables and candidate electronic descriptor variables.
[0134] The denoised extended crystal representation vector is split according to the original dimensionality partitioning method to obtain candidate crystal structure variables and candidate electronic descriptor variables. Candidate crystal structure variables include atomic type sequences, fractional coordinate matrices, and lattice parameters, while candidate electronic descriptor variables include Bader charge and atomic density of states. The candidate crystal structure variables obtained after separation represent the new crystal structures generated by the diffusion model. Subsequent screening using surrogate models and physical verification will determine whether they are valid metastable sulfide crystal structures.
[0135] A deep physical coupling exists between crystal structure and electronic structure. The structure-electronic descriptor joint denoising framework introduces this deep physical coupling into the denoising process. The spatial arrangement of atoms in a crystal (structural variables) directly determines the bonding mode and chemical bond orbital overlap between atoms, thus determining the local electronic structure (Bader charge and atomic density of states) of each atom. Conversely, specific electronic structure modes (such as charge transfer between specific elements and the characteristics of electronic density of states at the Fermi level) correspond to structural parameters (bond length, bond angle, and coordination polyhedron shape) within a specific range. The denoising network generates physically consistent electronic descriptors simultaneously when generating structural variables. The generation process performs a constrained search within the joint space formed by structural variables and electronic descriptors, avoiding unconstrained generation only within the structural space. The joint denoising trajectory is jointly restricted by structural variable constraints and electronic descriptor constraints, improving the stability of the generated structure within a physically reasonable region.
[0136] Specifically, the joint denoising framework enhances the condition by requiring that, within the latent space of the diffusion model, sulfide crystal structures satisfying the target ion conductivity (greater than 10 mSiemens / cm) often possess specific local electronic structure characteristics. For example, the Bader charge of conductive ions (such as Li and Na) is close to +1, and the atomic density of states along the carrier migration path exhibits low impedance characteristics at the Fermi level. When the denoising network simultaneously learns the predictions of structural noise and electronic descriptor noise, the denoising trajectory of the electronic descriptor applies implicit constraints to the structural denoising head through the network's shared backbone layer, guiding the structural denoising trajectory towards a structural region consistent with the target electronic characteristics. This is equivalent to introducing an implicit information channel connecting conditions and structure through the physical laws of electronic structure, in addition to the original condition vector (scalar encoding of target performance), naturally improving the condition satisfaction rate without adding external constraint terms.
[0137] The candidate electronic descriptor variable serves as a self-verification index for the electronic rationality of the generated structure. For each generated crystal structure, the candidate electronic descriptor (Bader charge and atomic density of states feature vector) is compared with the electronic descriptors of known sulfide crystals with similar chemical compositions in the training database, and the Euclidean distance is calculated. If the nearest neighbor distance between the candidate electronic descriptor and the database reference value exceeds a preset threshold (set to the 95th percentile of the distance between all sample pairs in the training set in this embodiment), the generated structure is marked as "low confidence," with a lower priority than structures whose electronic descriptors are close to the reference value. However, it is not directly discarded but is still retained in the subsequent proxy model screening process to avoid missing truly novel crystal structures due to the precision limitations of the electronic descriptor. In the verification experiment of this embodiment, the success rate of verification by first-principles calculation for generated structures marked as "low confidence" is significantly lower than that of unmarked structures, indicating that the candidate electronic descriptor has an effective pre-screening effect on the quality of the generated structure.
[0138] It is worth noting that the joint denoising framework, compared to a single denoising framework that only generates structural variables, has a limited increase in the number of model parameters. The number of parameters sharing the E(3)-GNN backbone is approximately 2.1M, and the structural denoising head and electronic denoising head each increase by approximately 0.1M parameters, for a total of approximately 2.3M parameters, which is only about 9.5% higher than the single denoising framework. However, in terms of training time, since the joint loss function requires backpropagation of the gradients of both structural noise and electronic descriptor noise, the computational cost of each training step increases by approximately 15%. The additional computational overhead has a significant cost-effectiveness advantage relative to the improvement in condition satisfaction rate it brings, indicating that the joint denoising framework is an efficient condition generation enhancement scheme.
[0139] Example 9
[0140] Based on Example 1, the step of filtering the candidate crystal structure variables using a surrogate model to output the near-optimal structure includes:
[0141] S5.1: Input the generated candidate crystal structure variables into a preset surrogate model, and predict the target physical performance parameters of the candidate crystal structure variables through the surrogate model.
[0142] The surrogate model predicts the target physical property parameters based on candidate crystal structure variables using a graph neural network model. The surrogate model is a pre-trained graph neural network model containing three message-passing layers with a hidden dimension of 64. Each message-passing layer is followed by a global average pooling operation, and finally, the predicted formation energy and ionic conductivity are output through two fully connected output heads. During training, the surrogate model uses first-principles calculations as labels. The training set contains first-principles calculation results for approximately 10,000 sulfide crystal structures. The learning rate is set to 0.0005, the batch size to 128, and the number of training epochs to 300. The Adam optimizer is used. The surrogate model achieves near-first-principles accuracy but with inference speeds several orders of magnitude faster, making it suitable for rapidly screening a large number of candidate structures.
[0143] The accuracy of the surrogate model can be demonstrated by first-principles calculations on an independent test set (containing approximately 1000 sulfide crystal structures, with no overlap with the training set), where the mean absolute error (MAE) of the predicted formation energy is no higher than 0.05 eV / atom, and the mean absolute error (MAE) of the predicted logarithmic value (log10σ) of the room-temperature ionic conductivity is no higher than 0.3 (corresponding to approximately twice the order of magnitude of the prediction error). Regarding screening accuracy, the surrogate model achieves at least 80% accuracy for near-optimal structures (meeting the conditions of formation energy less than -0.5 eV / atom and ionic conductivity greater than 10 mSiemens / cm), meaning that at least 80% of the candidate structures screened by the surrogate model can be verified through subsequent first-principles calculations.
[0144] Specifically, the surrogate model is computationally faster than first-principles calculations by approximately four to five orders of magnitude. Taking a single crystal structure as an example, first-principles calculations based on density functional theory typically take hours to tens of hours, while the surrogate model's single inference time is in the millisecond range. This allows the surrogate model to rapidly screen thousands of candidate crystal structures within seconds, significantly reducing the scale of subsequent precise first-principles calculations, thereby significantly improving overall search efficiency while ensuring the reliability of the results. The prediction accuracy of the surrogate model depends on the distribution of the training data. When the chemical composition of the candidate crystal structure differs significantly from the training set, the prediction error may exceed the preset range. During the physical verification stage, first-principles calculations are still required to accurately verify the candidate crystal structures.
[0145] S5.2: Based on the predicted target physical performance parameters, the candidate crystal structure variables are screened, and the candidate crystal structure variables that meet the preset performance threshold are retained as near-optimal structures.
[0146] Based on the target physical performance parameters predicted by the surrogate model, candidate crystal structure variables are screened, and those that meet preset performance thresholds are retained as near-preferred structures. The preset performance thresholds include the formation energy threshold (less than -0.5 eV / atom) and the ionic conductivity threshold (greater than 10 mSiemens / cm). Candidate structures that meet these threshold conditions are marked as near-preferred structures, meaning they possess both thermodynamic stability and excellent ion transport performance as predicted by the surrogate model.
[0147] A probabilistic quality improvement metric is introduced to measure search efficiency. This metric is defined as the ratio of the proportion of near-optimal structures under guided sampling to the proportion of near-optimal structures under unconditional sampling. The calculation formula is as follows:
[0148]
[0149] in, As a probability quality improvement indicator, To guide the sampling of candidate structure sets The probability, The candidate structure set under unconditional sampling The probability, This represents the set of near-optimal structures. The probability quality improvement index reflects the degree to which the conditional guidance improves the probability quality of the near-optimal solution region. When the probability quality improvement index reaches a preset threshold (e.g., 10), iterative sampling stops. The number of near-optimal structures is usually much smaller than the number of candidate structures initially sampled, thus significantly reducing the computational cost of subsequent physical verification.
[0150] The process of filtering candidate crystal structure variables using a surrogate model to output near-optimal structures can be implemented in two modes: single-round filtering or iterative filtering. In single-round filtering, a batch of candidate crystal structure variables is predicted by the surrogate model and then directly filtered according to a preset performance threshold. Candidate crystal structure variables that meet the preset performance threshold are retained as near-optimal structures. This is suitable for scenarios with a small number of candidate structures or where search efficiency requirements are not high. In iterative filtering, after each round of filtering, a probability quality improvement index is calculated. If the probability quality improvement index does not reach the preset threshold, the guiding parameters of the condition vector are updated based on the distribution characteristics of the current near-optimal structure in the performance space, triggering the next round of sampling and filtering. If the probability quality improvement index reaches the preset threshold, the iteration stops, and the candidate crystal structure variables that meet the preset performance threshold in the current round are output as near-optimal structures. It should be noted that, regardless of the mode used, the near-optimal structure must satisfy the surrogate model's predicted formation energy of less than -0.5 eV / atom and ionic conductivity of greater than 10 mSiemens / cm.
[0151] The masslift metric measures the efficiency improvement of conditional guided sampling compared to unconditional sampling within the near-optimal region. From an information theory perspective, a masslift value of 1 means that conditional guidance is completely ineffective, and guided sampling and unconditional sampling have the same probability of hitting the near-optimal region. A masslift value of 10 means that guided sampling has a 10 times higher probability of hitting the near-optimal region than unconditional sampling. In other words, with the same sampling budget, guided sampling can generate 10 times more effective candidate structures than unconditional sampling, which is equivalent to reducing the equivalent computational cost to 1 / 10 of the original. The advantage of the masslift metric is that it is independent of the absolute size of the near-optimal region. Even if the near-optimal region accounts for a very small proportion of the entire chemical space (e.g., the near-optimal rate of unconditional sampling is only 0.1%), as long as guided sampling can improve the near-optimal rate to 1%, the masslift value is 10, and the improvement in search efficiency is quantifiable.
[0152] Furthermore, the search index γ is obtained through an empirical diagnostic formula using a finite candidate pool. The estimate is given by K, where K is the total number of candidate structures sampled in each round. The search exponent γ represents the absolute number of near-optimal structures that increases by approximately e times (the natural logarithm base) when the sampling size is increased under the current search strategy. γThe search efficiency is linearly related to the sampling size when γ is close to 1, indicating that the current guidance strategy can effectively cover the near-optimal solution region. When γ is much less than 1, the marginal benefit of search efficiency with increasing sampling size decreases, indicating that the current search strategy needs to be adjusted. This embodiment of the invention monitors the changing trend of the search index γ to determine whether the guidance parameters of the condition vector need to be adjusted, thus realizing the adaptive update of the search strategy. Specifically, when the γ value of two consecutive rounds of sampling does not change significantly (the change is less than 0.05) and the masslift has not yet reached the preset threshold, the system automatically increases the basic guidance weight g by 0.5 to strengthen the condition guidance. When the masslift reaches the preset threshold (set to 10 in this embodiment), the iteration stops and enters the subsequent correction and resampling stage.
[0153] Example 10
[0154] Based on Example 9, the step of determining the importance weight based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions, resampling the near-optimal structure according to the importance weight to correct the distribution shift, and verifying and outputting the metastable sulfide crystal structure includes:
[0155] S5.3: Calculate the ratio of the probability density of the near-optimal structure under the guided sampling distribution to the probability density under the unconditional sampling distribution to obtain the importance weight.
[0156] The importance weight is calculated by taking the probability density of the near-optimal structure under guided sampling and unconditional sampling distributions as input, and then calculating the ratio of the probability density of the near-optimal structure under the guided sampling distribution to that under the unconditional sampling distribution to obtain the importance weight. The formula for calculating the importance weight is:
[0157]
[0158] in, Let i be the importance weight of the i-th near-optimal structure. To guide the sampling of the i-th near-optimal structure x i The probability density, For the i-th near-optimal structure under unconditional sampling The probability density, Let be the i-th near-optimal structure. The importance weight reflects the relative probability of each near-optimal structure in the two distributions and is used to correct for the distribution shift introduced by conditional guidance. If the probability density of a near-optimal structure in the guided sampling distribution is much higher than that in the unconditional sampling distribution, its importance weight is large, indicating that the structure strongly depends on conditional guidance for generation; conversely, if the probability density of a near-optimal structure in the two distributions is similar, its importance weight is close to 1, indicating that the structure also has a high generation probability in the unconditional case and is more robust.
[0159] S5.4: Resample the near-optimal structure according to the importance weights, correct the distribution offset, and obtain the corrected near-optimal structure.
[0160] After obtaining the importance weights and near-optimal structures through resampling, weighted random sampling is performed on the near-optimal structures according to the normalized importance weights to correct the distribution bias, resulting in corrected near-optimal structures. The resampling operation makes the distribution of the corrected near-optimal structure set closer to the target conditional distribution, while avoiding the distribution distortion caused by guided sampling. The resampled near-optimal structure set statistically approximates the target conditional distribution, providing a more reliable set of structure candidates for subsequent physical verification.
[0161] S5.5: Perform physical feasibility verification on the modified near-optimal structure and output the verified metastable sulfide crystal structure.
[0162] The physical feasibility verification stage obtains the modified near-optimal structure and performs physical feasibility verification on it, outputting a verified metastable sulfide crystal structure. Physical feasibility verification includes accurately evaluating the formation energy and ionic conductivity of each candidate structure through first-principles calculations, verifying the thermodynamic stability of the structure through molecular dynamics simulations, and verifying the rationality of chemical bonds through bond valence and calculations. The verified structure must simultaneously satisfy the metastable condition (formation energy distance from the convex hull 0 to 0.1 eV / atom) and the performance target (ionic conductivity greater than 10 mSiemens / cm) to be considered the final metastable sulfide crystal structure.
[0163] The modified resampling applies importance sampling theory, when the target distribution... (Guided sampling distribution) and reference distribution When there is bias in the (unconditional sampling distribution), an importance weight can be assigned to each sampling point. This is to correct for this bias, so that the weighted sample set statistically approximates the target distribution. .
[0164] in, The importance weight corresponding to the i-th sampling point; To guide the sampling distribution, it represents the distribution of the conditionally denoised sampling output of the model; It is an unconditional sampling distribution; This is the crystal structure sample obtained from the i-th sampling.
[0165] Guided sampling distribution (i.e., the distribution of conditionally denoised sampling) is the target conditional distribution. The biased approximation in the sampling method leads to an over-generation of structures with extremely high performance (significantly exceeding the target threshold), while relatively ignoring the boundary regions that just meet the threshold. This results in a performance distribution in the final structure set shifting towards the high end, failing to accurately reflect the target condition distribution. After corrective resampling, the structure set is more evenly distributed in the target performance space, covering the complete performance range from just meeting the threshold to significantly exceeding it. This provides a more diverse range of candidate structures for subsequent material synthesis experiments, avoiding the systematic bias of focusing only on extremely optimized structures while ignoring structures in the intermediate performance range.
[0166] In practice, importance weights are calculated using the log probability estimation of the trajectory from the diffusion model. The log probability density under the guided sampling distribution is then used. The log probability density under the unconditional sampling distribution is approximated by summing the denoised log probabilities at each time step on the reverse denoising trajectory. Similarly, the trajectory is approximated using unconditional denoising, where, This represents the logarithmic probability density of the i-th sampling point under the guided sampling distribution; This represents the logarithmic probability density of the i-th sampling point under the unconditional sampling distribution; This represents the i-th sampling point.
[0167] Since directly calculating the exact probability density in high-dimensional space is extremely computationally intensive, this embodiment of the invention adopts the Monte Carlo approximation method, which extracts several denoised trajectories with different randomness for each near-optimal structure, and takes the average of the log probabilities of each trajectory as an approximate estimate of the log probability density, thus achieving a balance between computational accuracy and efficiency.
[0168] In the specific implementation of an iterative search and correction resampling process, the number of candidate crystal structures generated in each round of sampling is set to K=500, and the proportion of near-optimal structures under unconditional sampling is approximately 1%. After the first round of condition-guided sampling, the surrogate model identified approximately 35 near-optimal structures, with a near-optimal rate of about 7% and a masslift of approximately 7, which did not reach the preset threshold of 10. The system automatically adjusted the basic guiding weight g from 2 to 2.5 and performed a second round of sampling, identifying approximately 52 near-optimal structures with a near-optimal rate of about 10.4% and a masslift of approximately 10.4, reaching the preset threshold and stopping the iteration. Subsequently, importance weights were calculated for the 52 near-optimal structures. Resampling was performed using normalized weights, and the top 20 were selected for precise first-principles calculations. This resulted in approximately 11 new metastable sulfide crystal structures. The first-principles computational resources consumed in the entire search process were equivalent to the cost of directly and precisely calculating 20 structures. In contrast, obtaining the same number of near-optimal structures through a completely random search requires an average of approximately 200 precise calculations, resulting in a search efficiency improvement of about 10 times. This aligns with the theoretical predictions of the masslift index, validating the effectiveness of the theoretical guarantee of search efficiency. Here, K represents the number of candidate crystal structures generated in each round of sampling. denoted as the proportion of near-optimal structures under unconditional sampling when generating K candidate structures; g represents the basic guiding weight of the diffusion model; masslift represents the probability quality improvement index, which is calculated by the ratio of the proportion of near-optimal structures under guided sampling to the proportion of near-optimal structures under unconditional sampling.
[0169] Example 11
[0170] like Figure 3 As shown, the present invention provides a system for the directional generation of metastable sulfide crystal structures, comprising:
[0171] The data acquisition module 301 is used to acquire the crystal structure representation vector, electronic descriptor and target physical performance parameters of the sulfide crystal structure, map the preset target physical performance parameters into a condition vector through a condition encoder, concatenate the crystal structure representation vector and the electronic descriptor into an extended crystal representation vector, and pair the condition vector with the extended crystal representation vector to generate a condition-structure training pair.
[0172] The model training module 302 is used to add noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noise scheduling sequence, construct a denoising network based on the noise scheduling sequence and the conditional vector, train the denoising network using a loss function with physical constraints, and obtain a physical constraint conditional diffusion model.
[0173] The sampling module 303 is used to perform reverse denoising sampling, which includes conditional denoising guidance and unconditional denoising guidance, based on the physical constraint diffusion model and the condition vector, and output preliminary denoising results.
[0174] The gradient correction module 304 is used to determine the adaptive guiding weights based on the degree of matching between the preliminary denoising result and the physical constraints, and to perform gradient correction on the preliminary denoising result according to the adaptive guiding weights to generate candidate crystal structure variables.
[0175] The filtering output module 305 is used to filter the candidate crystal structure variables through a surrogate model and output the near-optimal structure; determine the importance weight based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions; resample the near-optimal structure according to the importance weight to correct the distribution shift; verify and output the metastable sulfide crystal structure.
[0176] The data acquisition module 301 is configured to extract crystal structure data from a database of sulfide crystal structures and construct conditional-structure training pairs. This module includes a data preprocessing unit and a conditional encoding unit. The data preprocessing unit is responsible for cleaning and formatting the raw crystal structure data, including extracting atomic type sequences, fractional coordinate matrices, and lattice parameters, performing one-hot encoding and continuous value encoding, and extracting local electronic descriptors (Bader charge and atomic density of states). The conditional encoding unit is responsible for mapping the user-defined target physical performance parameters into fixed-dimensional conditional vectors using a multilayer perceptron conditional encoder.
[0177] Model training module 302 is configured to train a physically constrained diffusion model. This module includes a forward diffusion unit, a denoising network construction unit, and a loss calculation unit. The forward diffusion unit progressively adds Gaussian noise to the extended crystal representation vector according to a noise schedule, acting only on the fractional coordinate matrix and lattice parameters. The denoising network construction unit constructs an E(3)-GNN as a denoising network, ensuring rotational and translational equivariance, and jointly predicts structural noise and electronic descriptor noise. The loss calculation unit calculates the joint loss function with physical constraints, including the main denoising loss, interatomic spacing penalty term, bonding rule penalty term, and electronic descriptor denoising loss.
[0178] Sampling module 303 is configured to perform inverse denoising sampling. This module includes a noise initialization unit, a condition-guided sampling unit, and a basic-guided interpolation unit. The noise initialization unit initializes both structural variables and electronic descriptor variables with standard Gaussian noise simultaneously. The condition-guided sampling unit calculates both conditional and unconditional denoising results at each denoising time step. The basic-guided interpolation unit performs linear interpolation on the conditional and unconditional denoising results according to preset basic guiding weights, outputting a preliminary denoising result.
[0179] The gradient correction module 304 is configured to perform adaptive constraint guidance. This module includes a constraint evaluation unit, an adaptive scheduling unit, and a gradient correction unit. The constraint evaluation unit calculates the degree of violation of physical constraints using the constraint evaluation function, obtaining a constraint violation vector. The adaptive scheduling unit calculates the adaptive guidance weight vector based on the constraint violation vector and the current denoising time step. The gradient correction unit calculates the gradient of the constraint function and weights the gradient correction terms onto the preliminary denoising result, outputting the denoised extended crystal representation, separating the candidate crystal structure variables and candidate electronic descriptor variables.
[0180] The screening output module 305 is configured to screen near-optimal structures and output the final results. This module includes a surrogate screening unit, an importance weight calculation unit, a resampling unit, and a physical verification unit. The surrogate screening unit uses a surrogate model to quickly evaluate candidate crystal structure variables and screen out near-optimal structures that meet preset performance thresholds. The importance weight calculation unit calculates the ratio of the probability density of the near-optimal structure under guided sampling and unconditional sampling distributions to obtain the importance weight. The resampling unit resamples the near-optimal structures according to their importance weights to correct for distribution shifts. The physical verification unit performs first-principles calculations to verify the resampled structures and outputs the verified metastable sulfide crystal structures.
[0181] The overall workflow of the screening output module 305 includes: a proxy screening unit first receiving candidate crystal structure variables from the gradient correction module, quickly predicting their formation energy and ionic conductivity using a proxy model, and outputting a near-optimal structure set after screening according to a preset performance threshold. The near-optimal structure set is simultaneously passed to the importance weight calculation unit and the resampling unit. Based on the near-optimal structure set, the importance weight calculation unit estimates the log probability density of each near-optimal structure under guided sampling distribution and unconditional sampling distribution, calculates the ratio of the two to obtain the importance weight vector, and passes the normalized importance weight vector to the resampling unit. The resampling unit performs weighted random sampling on the near-optimal structure set according to the normalized importance weight, corrects the distribution offset introduced by guided sampling, obtains the corrected near-optimal structure set, and passes it to the physical verification unit. The physical verification unit sequentially performs first-principles calculations, molecular dynamics simulations, and bond valence calculations on each structure in the corrected near-optimal structure set, and outputs the final metastable sulfide crystal structure for the structure that simultaneously satisfies the metastable condition and performance target. The above four sub-units form a sequential pipeline structure. Intermediate results are passed between adjacent units in the form of a structure candidate set. Data is passed in one direction to ensure that the processing results of each sub-unit have complete physical meaning and traceability.
[0182] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for the directional generation of metastable sulfide crystal structures, characterized in that, include: The crystal structure representation vector, electronic descriptor, and target physical performance parameters of the sulfide crystal structure are obtained. The preset target physical performance parameters are mapped into a conditional vector through a conditional encoder. The crystal structure representation vector and the electronic descriptor are concatenated to form an extended crystal representation vector. The conditional vector and the extended crystal representation vector are paired to generate a conditional-structure training pair. The extended crystal representation vector in the conditional-structure training pair is denoised to obtain a noise scheduling sequence. A denoising network is constructed based on the noise scheduling sequence and the conditional vector. The denoising network is trained using a loss function with physical constraints to obtain a physical constraint conditional diffusion model. Based on the physical constraint diffusion model and the condition vector, reverse denoising sampling, including conditional denoising guidance and unconditional denoising guidance, is performed to output preliminary denoising results. Based on the degree of matching between the preliminary denoising results and physical constraints, adaptive guiding weights are determined, and the preliminary denoising results are gradient corrected according to the adaptive guiding weights to generate candidate crystal structure variables. The candidate crystal structure variables are filtered using a surrogate model to output the near-optimal structure; The importance weight is determined based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions. The near-optimal structure is then resampled according to the importance weight to correct the distribution offset, and the metastable sulfide crystal structure is verified and output.
2. The method according to claim 1, characterized in that, The process of obtaining the crystal structure representation vector, electronic descriptor, and target physical property parameters of the sulfide crystal structure includes: Extract the atomic type sequence, fractional coordinate matrix, and lattice parameters of crystals from a database of sulfide crystal structures; One-hot encoding is performed on the atomic type sequence, continuous value encoding is performed on the fractional coordinate matrix and the lattice parameters, and the encoding results of the one-hot encoding and the continuous value encoding are concatenated to obtain the crystal structure representation vector; Local electronic descriptors, including Bader charge and atomic density of states, are extracted from each crystal structure in the crystal structure representation vector. The formation energy of each crystal structure in the crystal structure representation vector is obtained by first-principles calculation, and the room temperature ionic conductivity is obtained by molecular dynamics simulation. The formation energy and the room temperature ionic conductivity of each crystal structure are combined to form a set of physical property tags.
3. The method according to claim 1, characterized in that, The step of mapping the preset target physical performance parameters into a conditional vector via a conditional encoder, concatenating the crystal structure representation vector with the electronic descriptor to form an extended crystal representation vector, and pairing the conditional vector with the extended crystal representation vector to generate conditional-structure training pairs includes: The preset target physical performance parameters are mapped into condition vectors through a multilayer perceptron condition encoder, wherein the multilayer perceptron condition encoder is composed of alternating stacks of fully connected layers, batch normalization layers and activation layers. The crystal structure representation vector and the electronic descriptor are concatenated dimensionally to obtain the extended crystal representation vector; The conditional vector and the extended crystal representation vector are paired according to the batch dimension to form conditional-structure training pairs.
4. The method according to claim 1, characterized in that, The step of adding noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noisy scheduling sequence includes: Gaussian noise is gradually added to the extended crystal representation vector in the conditional-structure training pair according to a preset noise schedule. The addition of Gaussian noise only affects the fractional coordinate matrix and lattice parameters, while the atom type remains unchanged. After a preset time step, the extended crystal representation vector is transformed into a standard Gaussian distribution to obtain a noise scheduling sequence.
5. The method according to claim 1, characterized in that, The process of constructing a denoising network based on the noise scheduling sequence and the conditional vector, training the denoising network using a loss function with physical constraints, and obtaining a physically constrained diffusion model includes: An isomorphic graph neural network is constructed as the denoising network, wherein the denoising network takes the noisy extended crystal representation and time step of the current time step in the noise scheduling sequence as input, and the conditional vector as the guiding signal, and jointly predicts structural noise and electronic descriptor noise. A joint loss function with physical constraints is constructed, wherein the joint loss function with physical constraints consists of a main denoising loss, an atomic spacing penalty term, a bonding rule penalty term, and an electronic descriptor denoising loss. The main denoising loss is the mean square error between the predicted value of the structural noise and the actual noise. The atomic spacing penalty term is used to penalize atomic spacings smaller than the minimum allowable spacing. The bonding rule penalty term is used to penalize structures that deviate from the chemical bonding rules. The electronic descriptor denoising loss is the mean square error between the predicted value of the electronic descriptor noise and the actual noise. The denoising network is trained by minimizing the joint loss function containing physical constraints to obtain the physical constraint diffusion model.
6. The method according to claim 1, characterized in that, Based on the physical constraint diffusion model and the condition vector, the process performs inverse denoising sampling, including both conditional and unconditional denoising guidance, and outputs preliminary denoising results, including: When training the physical constraint diffusion model, the condition vector is replaced with a zero vector with a preset probability to enable the physical constraint diffusion model to support unconditional denoising. The structural variables and electronic descriptor variables are initialized simultaneously with standard Gaussian noise. Based on the physical constraint diffusion model and the condition vector, inverse denoising sampling is performed on the structural variables and the electronic descriptor variables. The structural variables include fractional coordinate variables and lattice parameter variables, which are intermediate random variables used to generate the fractional coordinate matrix and the lattice parameter during the inverse denoising sampling process, respectively. At each denoising time step of the reverse denoising sampling, the currently noisy extended crystal representation and the condition vector are received through the physical constraint diffusion model, and the constraint-guided conditional denoising result and the unconstrained unconditional denoising result are calculated respectively. Based on the conditional denoising result and the unconditional denoising result, the conditional denoising result and the unconditional denoising result are linearly interpolated according to a preset basic guiding weight, and the preliminary denoising result is output. The preset basic guiding weight is a pre-set fixed weight value used for the linear interpolation.
7. The method according to claim 1, characterized in that, The step of determining the adaptive guiding weights based on the degree of matching between the preliminary denoising results and physical constraints includes: Based on the preliminary denoising results, the degree of violation of physical constraints is calculated through the constraint evaluation function to obtain the constraint violation vector, which includes the interatomic spacing constraint violation, coordination number constraint violation, and bond length distribution constraint violation. Based on the constraint violation vector and the current denoising time step, an adaptive guidance weight vector is calculated by an adaptive guidance scheduler; wherein the adaptive guidance weight vector is positively correlated with the constraint violation vector and negatively correlated with the current denoising time step.
8. The method according to claim 1, characterized in that, The step of performing gradient correction on the preliminary denoising result according to the adaptive guiding weights to generate candidate crystal structure variables includes: The gradient of the initial denoising result is calculated using the constraint function to obtain the gradient correction term of the physical constraint. Based on the adaptive guiding weight, the gradient correction term is weighted and superimposed on the initial denoising result to obtain the corrected denoising result. The corrected denoising result and the conditional vector are input into the denoising network, and the denoised extended crystal representation is output. The denoised extended crystal representation is separated to obtain candidate crystal structure variables and candidate electronic descriptor variables.
9. The method according to claim 8, characterized in that, The step of filtering the candidate crystal structure variables using a surrogate model and outputting the near-optimal structure includes: The generated candidate crystal structure variables are input into a preset surrogate model, and the target physical performance parameters of the candidate crystal structure variables are predicted by the surrogate model. Based on the predicted target physical performance parameters, the candidate crystal structure variables are screened, and the candidate crystal structure variables that meet the preset performance threshold are retained as near-optimal structures.
10. The method according to claim 9, characterized in that, The process of determining the importance weight based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions, resampling the near-optimal structure according to the importance weight to correct the distribution shift, and verifying and outputting the metastable sulfide crystal structure includes: The importance weight is obtained by calculating the ratio of the probability density of the near-optimal structure under the guided sampling distribution to the probability density under the unconditional sampling distribution. The near-optimal structure is resampled according to the aforementioned importance weights to correct the distribution offset and obtain the corrected near-optimal structure. The physical feasibility of the modified near-optimal structure is verified, and the verified metastable sulfide crystal structure is output.
11. A system for the directional generation of metastable sulfide crystal structures, characterized in that, include: The data acquisition module is used to acquire the crystal structure representation vector, electronic descriptor and target physical performance parameters of sulfide crystal structure, map the preset target physical performance parameters into a condition vector through a condition encoder, concatenate the crystal structure representation vector and the electronic descriptor into an extended crystal representation vector, and pair the condition vector and the extended crystal representation vector to generate a condition-structure training pair. The model training module is used to add noise to the extended crystal representation vector in the conditional-structure training pair to obtain a noise scheduling sequence, construct a denoising network based on the noise scheduling sequence and the conditional vector, train the denoising network using a loss function with physical constraints, and obtain a physical constraint conditional diffusion model. The sampling module is used to perform reverse denoising sampling, which includes conditional denoising guidance and unconditional denoising guidance, based on the physical constraint diffusion model and the condition vector, and output the preliminary denoising result. The gradient correction module is used to determine the adaptive guiding weights based on the degree of matching between the preliminary denoising result and the physical constraints, and to perform gradient correction on the preliminary denoising result according to the adaptive guiding weights to generate candidate crystal structure variables. The filtering output module is used to filter the candidate crystal structure variables through a surrogate model and output the near-optimal structure; The importance weight is determined based on the probability ratio of the near-optimal structure under guided sampling and unconditional sampling distributions. The near-optimal structure is then resampled according to the importance weight to correct the distribution offset, and the metastable sulfide crystal structure is verified and output.