Identification of nucleating agents using machine learning

By generating material embeddings through machine learning models and selecting nucleating agents based on similarity assessment, the problem of difficult nucleating agent screening in the synthesis of new materials is solved, and rapid and efficient synthesis of new materials is achieved.

CN120937085APending Publication Date: 2025-11-11GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480025219.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-12
Filing Date
2024-04-12
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently identify and screen suitable nucleating agents for the synthesis of new materials, resulting in time-consuming and expensive synthesis processes.

Method used

By using an embedded machine learning model to generate embeddings of the target material and candidate nucleating agents, a suitable nucleating agent is selected from the candidate nucleating agent set based on similarity evaluation, and the target material is rapidly synthesized by combining physical synthesis techniques.

Benefits of technology

It significantly accelerates the synthesis process of new materials, reduces computational resource consumption, improves synthesis efficiency and accuracy, and lowers costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937085A_ABST
    Figure CN120937085A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting nucleating agents for synthesis of a target material. According to one aspect, a method includes: receiving data identifying a target material; generating embedding of the target material in the potential space using an embedded machine learning model; and selecting one or more nucleating agents for the target material from the set of candidate nucleating agents based at least in part on the embedding of the target material in the potential space.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This manual relates to using machine learning models to process data.

[0002] Machine learning models receive input and generate outputs, such as predicted outputs, based on the received input. Some machine learning models are parametric models, and they generate outputs based on the received input and the values ​​of the model parameters.

[0003] Some machine learning models are deep models, which employ multiple layers to generate outputs from received inputs. For example, a deep neural network is a type of deep machine learning model that includes an output layer and one or more hidden layers, each of which applies a non-linear transformation to the received input to generate an output. Summary of the Invention

[0004] This specification describes a system implemented as a computer program on one or more computers at one or more locations, which can select nucleating agents for synthesizing target materials.

[0005] "Material" is any substance or mixture of substances that makes up a physical entity. For example, a material can be a metal (e.g., an alloy), a ceramic, a polymer, a composite material, or a molecular crystal.

[0006] Nucleation is the initial step in the formation of a new phase in a material, such as a specific arrangement of atoms (or molecules) with various chemical properties in space with a periodic translational structure, also known as a "crystal structure." During nucleation, particles of the material begin to form and disperse through a random process involving the emergence of small regions or "nuclei" of new phases that are thermodynamically more stable under current conditions than existing phases.

[0007] A nucleating agent is a substance added to a material to promote nucleation. Nucleating agents work by providing sites (nucleation sites) on which a phase (e.g., a crystalline phase) of the material can form more easily and rapidly than in a bulk material (i.e., an existing phase), thereby inducing and / or accelerating the formation of a phase in the material.

[0008] An "embedding" of an entity (e.g., a material) is representing the entity as an ordered collection of numerical values, such as a vector, matrix, or other tensor of numerical values.

[0009] If the first neural network is included in the second neural network, then the first neural network can be referred to as a "subnetwork" of the second neural network.

[0010] According to one aspect, a method executed by one or more computers is provided, the method comprising: receiving data identifying a target material; using an embedding machine learning model to generate an embedding of the target material in a latent space; and selecting one or more nucleating agents for the target material from a set of candidate nucleating agents, based at least in part on the embedding of the target material in the latent space.

[0011] In some implementations, the data for identifying the target material includes: (i) data for identifying the chemical composition of the target material, and (ii) data for identifying the crystal structure of the target material.

[0012] In some implementations, using an embedding machine learning model to generate an embedding of the target material in the latent space includes: using the embedding machine learning model to process model inputs, including data identifying the target material, based on the values ​​of the embedding machine learning model parameter set, to generate an embedding of the target material.

[0013] In some implementations, the embedded machine learning model is an embedded neural network that has been jointly trained with a predictive neural network, wherein the predictive neural network is configured to: receive the embeddings of input material generated by the embedded neural network; and process the embeddings of the input material according to the values ​​of the set of parameters of the predictive neural network to generate predictions characterizing the input material.

[0014] In some implementations, the prediction characterizing the input material includes the prediction of the energy of the input material.

[0015] In some implementations, the predictions characterizing the input material include corresponding predictions of the forces acting on each atom in the unit cell of the input material.

[0016] In some implementations, the predictions characterizing the input material include the predicted reconstruction of the input material's chemical composition and crystal structure.

[0017] In some implementations, joint training of the embedded neural network and the predictive neural network includes: obtaining a set of training examples, wherein each training example corresponds to a corresponding training material and includes: (i) a training input representing the training material, and (ii) a target output of the predictive neural network; and jointly training the embedded neural network and the predictive neural network on the set of training examples.

[0018] In some implementations, joint training of the embedded neural network and the predictive neural network on a set of training examples includes, for each training example: using the embedded neural network to process the training input of the training example to generate an embedding of the training material represented by the training input; using the predictive neural network to process the embedding of the training material to generate a predicted output; determining the gradient of an objective function with respect to the parameter set of the embedded neural network and the parameter set of the predictive neural network, the objective function measuring the difference between (i) the predicted output generated by the predictive neural network and (ii) the target output specified by the training example; and using the gradient to adjust the values ​​of the parameter sets of the embedded neural network and the predictive neural network.

[0019] In some implementations, selecting one or more nucleating agents for the target material from a set of candidate nucleating agents, based at least in part on the embedding of the target material in the latent space, includes: for each candidate nucleating agent in the set of candidate nucleating agents, obtaining the corresponding embedding of the candidate nucleating agent generated using an embedding machine learning model; for each candidate nucleating agent in the set of candidate nucleating agents, determining a corresponding distance in the latent space between (i) the embedding of the candidate nucleating agent and (ii) the embedding of the target material; and selecting one or more nucleating agents for the target material based at least in part on the distance.

[0020] In some implementations, selecting one or more nucleating agents for the target material based at least in part on the distance includes: ranking each candidate nucleating agent in the candidate nucleating agent set based on the corresponding distance of each candidate nucleating agent in the potential space from the target material; filtering the candidate nucleating agent set based on the ranking to remove one or more candidate nucleating agents from the candidate nucleating agent set; and selecting one or more candidate nucleating agents from the remaining candidate nucleating agents in the candidate nucleating agent set as nucleating agents for the target material.

[0021] In some implementations, selecting one or more remaining candidate nucleating agents from the candidate nucleating agent set as nucleating agents for the target material includes: determining, for each remaining candidate nucleating agent in the candidate nucleating agent set, whether the candidate nucleating agent can stably coexist with the target material; and filtering the candidate nucleating agent set to remove any candidate nucleating agents that cannot stably coexist with the target material; after filtering the candidate nucleating agent set, selecting one or more remaining candidate nucleating agents from the candidate nucleating agent set as nucleating agents for the target material based on: (i) ranking each candidate nucleating agent based on its corresponding distance from the target material in the potential space, and (ii) whether each candidate nucleating agent can stably coexist with the target material.

[0022] In some implementations, embedded machine learning models include neural networks.

[0023] In some implementations, the embedded machine learning model includes a graph neural network, which comprises multiple message-passing neural network layers.

[0024] In some implementations, the method further includes generating graph data representing a graph that is included in the network input to a graph neural network, wherein: the graph includes multiple nodes and multiple edges; each node in the graph represents a corresponding atom in a unit cell of the target material; and each edge in the graph connects a corresponding pair of nodes that represent a pair of atoms separated by a three-dimensional spatial distance less than a threshold in the unit cell of the target material.

[0025] In some implementations, each node in the graph is associated with a corresponding set of node features that characterize the atom represented by the node; and for each node in the graph, the set of node features includes features that identify the element type of the atom.

[0026] In some implementations, the method further includes providing each of the selected nucleating agents for synthesizing the target material.

[0027] In some implementations, the method further includes, for each of one or more of the selected nucleating agents, physically synthesizing the target material by means of a physical synthesis technique involving nucleating the target material using the selected nucleating agent.

[0028] According to another aspect, a system is provided comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the methods described herein.

[0029] According to another aspect, one or more non-transitory computer storage media are provided, which store instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the methods described herein.

[0030] Specific embodiments of the subject matter described in this specification may be implemented in order to achieve one or more of the following advantages.

[0031] The discovery and synthesis of new materials have driven technological advancements in fields such as energy storage and conversion (e.g., where batteries, supercapacitors, and fuel cells rely on new materials to achieve higher energy densities and faster charging times), electronics and photonics (e.g., where new materials enable thinner, more flexible, and more energy-efficient devices), and semiconductors (where next-generation transistors, memory devices, and integrated circuits can rely on new materials).

[0032] Various computational methods can be used to generate libraries of candidate novel materials, each with a specific crystal structure and chemical composition. However, the physical synthesis of compounds with specific structures and compositions is difficult to induce intentionally, and the physical experiments for synthesizing new materials are typically carried out through expensive and time-consuming trial-and-error processes.

[0033] The system described in this specification addresses these problems. Specifically, the system can identify nucleating agents for inducing, controlling, and / or accelerating the formation of target materials with specific chemical compositions and crystal structures. Therefore, the system can significantly accelerate the process of successfully synthesizing new materials with specific chemical compositions and crystal structures.

[0034] Specifically, given a target material with a specific chemical composition and crystal structure, the system can automatically screen a library of candidate nucleating agents to identify one or more suitable nucleating agents for the synthesis of the target material. If a candidate nucleating agent has a similar crystal structure and composition to the target material, that candidate nucleating agent is more likely to efficiently synthesize the target material. This is because compositional similarity reduces the likelihood of undesirable reactions or phase separation that could inhibit crystal growth or lead to defects, and because similar crystal structures can lower the surface energy barrier that must be overcome to initiate nucleation. Therefore, the system can screen a library of candidate nucleating agents to identify those with a composition and structure similar to the target material.

[0035] However, numerically assessing the similarity between the composition and structure of materials is challenging, for example, because different materials have different amounts of chemical composition, different cell sizes, different numbers of atoms in the cell, and so on. Therefore, for example, data representing the chemical structure and composition of a first material may have a different numerical dimension than data representing the chemical structure and composition of a second material, and few metrics exist to measure the similarity between different numerical data defined in different dimensions.

[0036] To address this issue, in some implementations, the system can evaluate the similarity between the target material and the candidate nucleating agent by generating corresponding embeddings of the target material and the candidate nucleating agent in a shared latent space. The corresponding embeddings of the target material and the candidate nucleating agent in the shared latent space have the same dimension, and the system can use conventional numerical similarity measures—such as those based on the L1 or L2 norm—to measure the similarity between the embeddings.

[0037] The system can generate embeddings of materials (e.g., target materials or candidate nucleating agents) by processing data defining the chemical composition and structure of materials using an embedding machine learning model. The system can train the embedding machine learning model using machine learning training techniques to generate useful embeddings that encode rich informational content characterizing the material. For example, for an embedding machine learning model implemented as an embedding neural network, the system can jointly train the embedding neural network with a predictive neural network. Specifically, the system can train the embedding neural network to process data on the chemical composition and structure of the material to generate embeddings that enable the predictive neural network to effectively perform machine learning tasks when processed by the predictive neural network. Machine learning tasks could be, for example, predicting the properties of the material, predicting the energy of the material, predicting forces on atoms (or molecules) in the material, or reconstructing the chemical structure and composition of the material.

[0038] The system can efficiently and rapidly screen large libraries of candidate nucleating agents while consuming far fewer computational resources, such as memory and computing power, than alternative methods. For example, alternative methods for evaluating the feasibility of synthesizing a target material using candidate nucleating agents could involve performing molecular dynamics simulations on a chemical system including both the candidate nucleating agent and the target material. However, such molecular dynamics simulations would be highly computationally intensive, for example, because they would require simulating complex interactions between a large number of atoms over long timescales with short time steps. In contrast, the system described in this specification can evaluate the feasibility of candidate nucleating agents by using an embedding machine learning model to generate embeddings of the candidate nucleating agent and the target material, and then evaluating the similarity between the individual embeddings. As mentioned above, the number of operations required to perform two forward passes through the embedding machine learning model and then evaluate the similarity metric is orders of magnitude fewer than that required to perform molecular dynamics simulations.

[0039] In one aspect of this disclosure, a method executed by one or more computers is provided, the method comprising: receiving data identifying a target material; using an embedding machine learning model to generate an embedding of the target material in a latent space; and selecting one or more nucleating agents for the target material from a set of candidate nucleating agents, at least in part based on the embedding of the target material in the latent space. The method may further comprise physically synthesizing the target material by means of a physical synthesis technique involving nucleating the target material using the selected nucleating agent. In some cases, the amount of nucleating material used may be minute compared to the amount of target material synthesized.

[0040] Each of the target material and / or nucleating agent can be, for example, a metal (e.g., steel, solder, or alloy), or a ceramic (e.g., glass-ceramic), or a polymer, or a composite, or a molecular crystal (e.g., a molecular solid comprising discrete molecules held together by intermolecular forces). In some cases, both the target material and / or the nucleating agent are crystalline solids.

[0041] In some cases, when nucleating agents are used in the physical synthesis of a target material, the nucleating agent can be selected such that one of several different polymorphs of the target material is formed. For example, a polymorph formed using a selected nucleating agent may have worse thermodynamic stability compared to one or more polymorphs formed without the selected nucleating agent. In some cases, the target material may be a molecular crystal of an active pharmaceutical ingredient, and a polymorph formed using a selected nucleating agent may have physicochemical properties (e.g., solubility) that make it more suitable as a pharmaceutical product than one or more polymorphs of the target material formed without the selected nucleating agent.

[0042] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the description, drawings, and claims. Attached Figure Description

[0043] Figure 1 An example nucleating agent selection system is shown.

[0044] Figure 2 This is a flowchart of an example process for selecting nucleating agents for synthesizing target materials.

[0045] Figure 3 This is a flowchart of an example process for generating embeddings of materials, such as target materials or candidate nucleating agents, using an embedding machine learning model implemented as a graph neural network.

[0046] Figure 4 This is a flowchart of an example process for training an embedded machine learning model implemented as an embedded neural network.

[0047] Figure 5 This is a flowchart of an example process for determining whether a candidate nucleating agent can stably coexist with the target material (and / or with one or more reactants to be used during the synthesis of the target material).

[0048] Figure 6 An example is shown for determining the corresponding distance between (i) the embedding of the target material and (ii) the corresponding embedding of each candidate nucleating agent in the nucleating agent set in the potential space.

[0049] Figure 7 The nucleating agent Li3PO4 selected by the nucleating agent selection system for the target material Li2Si2O5 is shown.

[0050] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation

[0051] Figure 1 An example nucleating agent selection system 100 is shown. The nucleating agent selection system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations, wherein the systems, components and techniques described below are implemented.

[0052] System 100 is configured to process data identifying the chemical composition and crystal structure of target material 102 to generate data identifying one or more nucleating agents 104 predicted to be effective for the synthesis of target material 102.

[0053] Data identifying the chemical composition of target material 102 can identify the elements included in target material 102 and their corresponding proportions in the target material, for example, through a chemical formula such as A2B2C7, where A, B, and C are different elements, and the numerical subscripts indicate the relative proportions of these elements in the target material.

[0054] Data for identifying the crystal structure of a target material may include data identifying the corresponding three-dimensional (3D) spatial location of each atom in the unit cell of the target material and the element type of that atom. The data for identifying the crystal structure of the target material may further include, for example, one or more of the following: data identifying the lattice parameters of the unit cell (e.g., the length of the cell edges and / or the angles between the cell edges), the lattice type (e.g., cubic, tetragonal, orthogonal, hexagonal, trigonal, monoclinic, or triclinic), or the space group (e.g., information defining how the unit cells are arranged and repeated in three-dimensional space to form the target material).

[0055] System 100 may receive data identifying target material 102, for example, through a user interface (e.g., a graphical user interface) or an application programming interface (API) provided by the system. The data identifying target material 102 may be provided to system 100, for example, by a user or an upstream system.

[0056] System 100 can generate data to identify each nucleating agent 104, for example, by specifying the chemical composition and crystal structure of the nucleating agent. System 100 can provide the data identifying the nucleating agent 104, for example, by presenting the data identifying the nucleating agent 104 on a display of a user device, or by storing the data identifying the nucleating agent 104 in a memory, or by transmitting the data identifying the nucleating agent 104 via a data communication network.

[0057] The system includes: (i) data for identifying a set 104 of candidate nucleating agents, (ii) an embedded machine learning model 106, and (iii) a filtering engine 112, each of which will be described in detail below (and throughout this specification).

[0058] The candidate nucleating agent set 104 may include any appropriate number of candidate nucleating agents, such as at least 100, at least 1,000, at least 100,000, or at least 1,000,000 candidate nucleating agents. In some cases, such as in addition to the data identifying the target material 102, the system 100 also receives the candidate nucleating agent set 104 as additional input.

[0059] The embedded machine learning model 106 is configured to process model inputs, including data identifying the material (e.g., the chemical composition and crystal structure of the material), based on the values ​​of the set of machine learning model parameters of the machine learning model, to generate an embedding of the material in a latent space. The embedding of the material can be, for example, a vector, matrix, or other tensor with any suitable dimension.

[0060] The embedded machine learning model 106 can have any suitable machine learning model architecture. For example, the embedded machine learning model 106 can be implemented as a neural network, such as a graph neural network or a convolutional neural network. (Reference) Figure 3 An example process for generating material embeddings using an embedding machine learning model implemented as a graph neural network is described.

[0061] System 100 can train the set of machine learning model parameters of embedded machine learning model 106 using machine learning training techniques, so that embedded machine learning model 106 generates embeddings that encode rich information content of the representation material. For example, for embedded machine learning model 106 implemented as an embedded neural network, the system can jointly train the embedded neural network and the predictive neural network. Specifically, system 100 can train the embedded neural network to process the data representing the material to generate embeddings that enable the predictive neural network to effectively perform machine learning tasks when processed by the predictive neural network. (Reference) Figure 4 An example process for jointly training an embedded neural network and a predictive neural network is described.

[0062] System 100 uses an embedding machine learning model 106 to generate: (i) an embedding 108 for the target material 102, and (ii) a corresponding embedding 110 for each nucleating agent in the candidate nucleating agent set 104. Specifically, system 100 uses the embedding machine learning model 106 to process data characterizing the target material 102 to generate the target material embedding 108, and for each candidate nucleating agent 104, system 100 uses the embedding machine learning model to process data characterizing that candidate nucleating agent 104 to generate a corresponding nucleating agent embedding 110.

[0063] In some cases, system 100 may pre-calculate and store the corresponding nucleating agent embedding 110 for each nucleating agent in the nucleating agent set 104, i.e., instead of recalculating the same nucleating agent embedding 110 each time system 100 identifies a nucleating agent for a new target material 102. Pre-calculating the nucleating agent embedding 110 can reduce the consumption of computational resources, for example, by preventing redundancy and repeated generation of the nucleating agent embedding 110.

[0064] Filtering engine 112 processes target material embedding 108 and nucleating agent embedding 110 to select one or more nucleating agents 104 for target material 102. Specifically, filtering engine 112 applies one or more filtering operations to a candidate nucleating agent set—at least one of these filtering operations being based on target material embedding 108 and nucleating agent embedding 110—to filter the candidate nucleating agent set 104 to remove candidate nucleating agents from the candidate nucleating agent set 104 that meet the filtering criteria defined by the filtering operations. Filtering engine 112 can then provide any remaining candidate nucleating agents 104 in the candidate nucleating agent set 104 as nucleating agents 104 for target material 102 after the filtering operations.

[0065] The filtering engine 112 applies at least one filtering operation based on the target material embedding 108 and the nucleating agent embedding 110 to the candidate nucleating agent set 104. Specifically, the filtering engine 112 can filter the candidate nucleating agent set 104 to maintain only those candidate nucleating agents that are relatively close to the target material in the latent space and therefore have a similar chemical composition and structure to the target material. Candidate nucleating agents with a similar chemical composition and structure to the target material are more likely to be effective nucleating agents for the target material, for example, because the similarity of composition reduces the likelihood of undesirable reactions or phase separation that could inhibit crystal growth or lead to defects, and because the similarity of crystal structure can reduce the surface energy barrier that must be overcome to initiate nucleation.

[0066] The following is for reference. Figure 2 A detailed description is provided of an example of a filtering operation that can be performed by the filtering engine 112 to filter the candidate nucleating agent set 104.

[0067] The nucleating agent 104 identified by system 100 can be used in any of a number of downstream applications. Several examples of downstream applications involving the nucleating agent 104 identified by system 100 are described below.

[0068] In some implementations, system 100 may perform additional computational verification for each nucleating agent 104. For example, for each nucleating agent 104, system 100 may perform molecular dynamics simulations to simulate the chemical system including target material 102 (or reactants used to generate target material 102) and nucleating agent 104 to assess whether nucleating agent 104 effectively promotes nucleation of the target material. Typically, the additional computational verification of nucleating agent 104 consumes significantly more computational resources than system 100 requires to initially identify nucleating agents 104 from the candidate nucleating agent set 104 using embedded machine learning model 106. System 100 can thus significantly reduce computational resource consumption by initially screening the candidate nucleating agent set 104 using embedded machine learning model 106 and filtering engine 112 and then performing additional computational verification only for the remaining nucleating agents.

[0069] In some implementations, for each of one or more nucleating agents 104, the target material 102 is physically synthesized by a physical synthesis technique involving the nucleation of the target material 102 using the nucleating agent 104. The physical synthesis technique may be, for example, solid-state reaction synthesis, ceramic synthesis, carbothermal synthesis, combustion synthesis, hydrothermal synthesis, sol-gel synthesis, coprecipitation synthesis, precursor synthesis, vapor deposition synthesis, high-pressure synthesis, or electrochemical synthesis.

[0070] Figure 2This is a flowchart of an example process 200 for selecting a nucleating agent for synthesizing a target material. For convenience, process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, such as those appropriately programmed according to this specification. Figure 1 The nucleating agent selection system 100 can execute process 200.

[0071] The system receives data to identify the target material (202). The system may receive data, for example, through a user interface or API, from a user or from an upstream system. The data to identify the target material may include data defining the chemical composition and crystal structure of the target material.

[0072] The system uses an embedded machine learning model to process the model input, including data for identifying the target material, based on the values ​​of the embedded machine learning model parameter set, to generate an embedding of the target material in the latent space (204). Reference Figure 3 An example process is described using an embedding machine learning model implemented as a graph neural network to generate embeddings of materials (e.g., target materials). References Figure 4 An example process for training an embedded machine learning model that is implemented as a neural network is described.

[0073] The system obtains the corresponding embedding of each candidate nucleating agent in the latent space for each candidate nucleating agent in the candidate nucleating agent set, which is generated using an embedding machine learning model (206). Specifically, for each candidate nucleating agent, the embedding of the candidate nucleating agent is generated by processing the model input, which includes data identifying the candidate nucleating agent, according to the values ​​of the embedding machine learning model parameter set using the embedding machine learning model.

[0074] In some cases, the system has pre-computed and stored the embeddings of candidate nucleating agents. In these cases, the system accesses the pre-computed embeddings from memory instead of using an embedding machine learning model to regenerate the embeddings of candidate nucleating agents.

[0075] The system determines, for each candidate nucleating agent in the candidate nucleating agent set, the corresponding distance in the latent space between (i) the embedding of the candidate nucleating agent and (ii) the embedding of the target material (208). The system can use any appropriate numerical distance measurement, such as Euclidean distance measurement, or distance measurement based on the L1 norm, or distance measurement based on the L2 norm, etc., to measure the distance between two embeddings in the latent space.

[0076] The system filters the set of candidate nucleating agents based on the distance between the candidate nucleating agents and the target material in the potential space (i.e., as calculated at step 208) (210).

[0077] For example, the system can determine that any candidate nucleating agent that is at least a threshold distance from the target material in the potential space should be filtered (removed) from the candidate nucleating agent set.

[0078] As another example, the system can rank candidate nucleating agents based on their respective distances from the target material in the latent space, for example, from lowest to highest distance. The system can then filter the candidate nucleating agent set based on this ranking to remove one or more candidate nucleating agents. For example, the system can determine that N candidate nucleating agents should be maintained at the minimum distance to the target material, and all other candidate nucleating agents should be removed from the candidate nucleating agent set (where N is a positive integer value). As another example, the system can determine that X% of the candidate nucleating agents at the minimum distance to the target material should be maintained, and all other candidate nucleating agents should be removed from the candidate nucleating agent set (where X is a percentage between 0% and 100%).

[0079] Therefore, the system can filter the set of candidate nucleating agents to maintain only those that are relatively close to the target material in the latent space and thus have a similar chemical composition and structure. Candidate nucleating agents with a similar chemical composition and structure to the target material are more likely to be effective nucleating agents for the target material.

[0080] Optionally, the system filters the candidate nucleating agent set based on one or more additional filtering criteria (212).

[0081] For example, the system can determine for each remaining candidate nucleating agent in the candidate nucleating agent set whether that candidate nucleating agent can stably coexist with the target material. The system can then filter the candidate nucleating agent set to remove any candidate nucleating agents that cannot stably coexist with the target material. Candidate nucleating agents that cannot stably coexist with the target material are unlikely to be used effectively in the synthesis of the target material, for example, because adding the candidate nucleating agent to the target material may trigger undesirable chemical reactions. The system can use any appropriate computational technique to determine whether a candidate nucleating agent can stably coexist with the target material (or reactants to be used during the synthesis of the target material). Reference Figure 5 An example procedure is described for determining whether a candidate nucleating agent can stably coexist with the target material (and / or with the reactants to be used during the synthesis of the target material).

[0082] As another example, the system can determine for each of the remaining candidate nucleating agents in the candidate nucleating agent set whether the target material can stably coexist with each of the reactants in the set of reactants to be used during the synthesis of the target material. Candidate nucleating agents that cannot stably coexist with the planned reactants of the target material are unlikely to be effective for the synthesis of the target material, for example, because the candidate nucleating agent might disrupt the synthesis reaction by reacting with the reactants. The system can use any suitable computational technique to determine whether a candidate nucleating agent can stably coexist with the target material (or the reactants to be used during the synthesis of the target material). Reference Figure 5 An example procedure is described for determining whether a candidate nucleating agent can stably coexist with the target material (and / or with the reactants to be used during the synthesis of the target material).

[0083] After filtering the set of candidate nucleating agents, the system can select some or all of the remaining candidate nucleating agents in the set as nucleating agents for the target material (214). The system can then provide the selected nucleating agents (i.e., selected by the novel operation and configuration of the system, as described in this specification) for, for example, additional computational verification using molecular dynamics simulations, or for the physical synthesis of the target material.

[0084] Figure 3 This is a flowchart of an example process 300 for generating embeddings of materials, such as target materials or candidate nucleating agents, using an embedding machine learning model implemented as a graph neural network. For convenience, process 300 will be described as being executed by a system of one or more computers located in one or more locations. For example, such as those appropriately programmed according to this specification. Figure 1 The nucleating agent selection system 100 can execute process 300.

[0085] The system generates graph data (302) representing a graph characterizing the material. This graph includes a set of nodes and a set of edges. Each node in the node set represents a corresponding atom in a unit cell within the crystal structure of the material. Each edge in the edge set connects a corresponding pair of nodes in the node set. More specifically, each edge connection represents a corresponding pair of nodes that are separated by a pair of atoms in the unit cell of the material at a distance less than a threshold three-dimensional (3D) spatial distance.

[0086] The system associates each node in the node set with a set of node features that characterize the atom represented by the node. More specifically, the set of node features can include atom-specific features that characterize, for example, the element type and 3D spatial location of the atom, as well as global features including the lattice parameters of the unit cell, the lattice type of the unit cell, and so on.

[0087] For each node in the graph, the system uses an encoder subnetwork of the graph neural network to process the node's set of node features to generate the node embedding (304). The encoder subnetwork can have any suitable neural network architecture; for example, the encoder subnetwork can include a series of fully connected neural network layers.

[0088] The system uses a series of message-passing neural network layers of a graph neural network to process node embeddings of nodes in the graph (306). Each message-passing neural network layer is configured to process an input set of node embeddings through a set of message-passing neural network layer operations to generate an output set of node embeddings, the input set of node embeddings including the corresponding node embedding of each node in the graph, the message-passing neural network layer operations being topologically conditioned on the graph and parameterized by a set of layer parameters. The “topology” of the graph refers to how the nodes in the graph are connected by edges in the graph. The message-passing neural network layer can perform topologically conditioned operations on the graph, for example, updating the node embedding of each node in the graph based solely on the node embeddings of its neighboring nodes—that is, the nodes connected to it by edges in the graph.

[0089] The first message-passing neural network layer can receive a set of node embeddings generated by the encoder subnetwork, and each subsequent message-passing neural network layer can receive a set of node embeddings generated by the previous message-passing neural network layer.

[0090] The system generates material embeddings based on node embeddings generated by the final message-passing neural network layer in the sequence of message-passing neural network layers (308). For example, the system can generate material embeddings by aggregating (e.g., summing or averaging) the node embeddings generated by the final message-passing neural network layer.

[0091] Figure 4 This is a flowchart of an example process 400 for training an embedded machine learning model implemented as an embedded neural network. For convenience, process 400 will be described as being executed by a system of one or more computers located in one or more locations. For example, such as those appropriately programmed according to this specification. Figure 1 The nucleating agent selection system 100 can execute process 400.

[0092] The system obtains a set of training examples, each corresponding to a specific training material, and each training example includes: (i) training input comprising data defining the chemical composition and crystal structure of the training material, and (ii) a target output (402). The chemical composition and crystal structure of the training material can be obtained from many publicly accessible databases such as the Cambridge Structure Database (CSD) or the Inorganic Crystal Structure Database (ICSD).

[0093] In some implementations, each training example includes a target output that defines the energy (e.g., total energy) of the training material. The system can determine the energy of the material, for example, using density functional theory (DFT) calculations.

[0094] In some implementations, each training example includes a target output that defines the corresponding force acting on each atom in the unit cell of the training material (e.g., the force on each atom in a single unit cell of the training material, without considering interactions with atoms in other unit cells). The system can determine the forces acting on the atoms in the unit cell of the training material, for example, using DFT calculations. The forces acting on the atoms can be represented, for example, as 3D vectors defining the corresponding forces in the x, y, and z directions.

[0095] In some implementations, each training example includes some or all of the same target output as the training input of the training example. For example, the target output may include the crystal structure of the training material included in the training input.

[0096] The system performs joint training on the training example set for both the embedded neural network and the prediction neural network. The joint training of the embedded neural network and the prediction neural network on the training examples is described below with reference to steps 404-410.

[0097] The system uses an embedded neural network to process the training input of the training examples to generate an embedding of the training material represented by the training input (404).

[0098] The system uses a predictive neural network to process the embeddings of the training material to generate a predictive output (406). The predictive output defines a prediction of the target output specified by the training examples and may include one or more of the following: the predictive energy of the training material, the predictive force acting on atoms in the unit cell of the training material, or the predictive reconstruction of the training input processed by the embedded neural network.

[0099] A predictive neural network can have any suitable neural network architecture that enables it to perform the function it describes. Specifically, a predictive neural network can include any suitable type of neural network layers (e.g., fully connected layers, attention layers, convolutional layers, etc.) connected in any suitable number (e.g., 3 layers, 5 layers, or 10 layers) and in any suitable configuration (e.g., a directed graph as layers). For example, in a particular example, a predictive neural network can include a series of fully connected neural network layers.

[0100] The system determines the gradient of an objective function by considering the parameter sets of the embedded neural network and the prediction neural network, which measures the difference between (i) the predicted output generated by the prediction neural network and (ii) the target output specified by the training examples (408). The objective function may, for example, use the L1 norm or L2 norm or any other suitable method to measure the difference between the predicted output and the target output. The system may, for example, use backpropagation to determine the gradient of the objective function.

[0101] The system uses gradients to adjust the values ​​of the embedded neural network parameter set and the predictive neural network parameter set (410). For example, the system can use the update rules of any suitable gradient descent optimization algorithm, such as RMSprop or Adam, to adjust the values ​​of the embedded neural network parameter set and the predictive neural network parameter set.

[0102] Figure 5 This is a flowchart of an example process 500 for determining whether a candidate nucleating agent can stably coexist with a target material (or reactants to be used during the synthesis of the target material). For convenience, process 500 will be described as being executed by a system of one or more computers located in one or more locations. For example, such as those appropriately programmed according to this specification. Figure 1 The nucleating agent selection system 100 can execute process 500.

[0103] The system determines the thermodynamic convex hull (or phase diagram) of the chemical system comprising the target material and candidate nucleating agents (502). For example, if the target material has the chemical formula A2B2C7 and the candidate nucleating agent has the chemical formula X2C3, the system can determine the thermodynamic convex hull of the chemical system ABXC. The system can, for example, use a DFT or machine learning model to determine the thermodynamic formation energies of the components of the chemical system. Optionally, the chemical system can further include each reactant in the set of reactants used during the synthesis of the target material.

[0104] If the thermodynamic convex hull includes the tie-line between the candidate nucleating agent and the target material, then the system determines that the candidate nucleating agent can stably coexist with the target material (504). Of course, other methods can also be used to determine whether the target material can stably coexist with the target material.

[0105] Alternatively, if the thermodynamic convex hull includes the junction between the target material and the reactants, the system can determine that the candidate nucleating agent can stably coexist with the reactants to be used during the synthesis of the target material (506).

[0106] Figure 6An example is shown of determining the corresponding distances in the latent space between (i) the embedding 602 of the target material and (ii) the corresponding embedding 606-AE of each candidate nucleating agent in the candidate nucleating agent set. The distance 604 between the target material 602 and the nucleating agent embedding 606-A is smaller than the distance 608 between the target material 602 and the nucleating agent 606-B. Therefore, nucleating agent 606-A may have a higher compositional and structural similarity to the target material than nucleating agent 606-B, and the nucleating agent selection system is more likely to select nucleating agent 606-A rather than nucleating agent 606-B.

[0107] Figure 7 The nucleating agent Li3PO4 selected by the nucleating agent selection system for the target material Li2Si2O5 is shown.

[0108] Höland et al. reported in Phosphorus Research Bulletin 19 (2005) 36 that Li3PO4 was successfully used as a nucleating agent to induce the Li2Si2O5 phase in SiO2-Li2O-Al2O3-K2O-ZrO2-based functional glass ceramics. The nucleating agent selection system described in this specification identifies Li3PO4 as one of the top four recommended nucleating agents, alongside Li4P2O7, Li2SiO3, and Na2Si2O5.

[0109] DeCeanne et al.'s Journal of Non-Crystalline Solids, 591 (2022) 121714 reports various nucleating agents in glass-ceramic systems; for example, the addition of Nb2O5 to the Na2O-Al2O3-SiO2 glass system can induce the nucleation of NaNbO3. The nucleating agent selection system described in this specification ranks Nb2O5 first among all binary oxides for selectively nucleating the NaNbO3 phase.

[0110] This specification uses the term "configured" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform a specific operation or action, this means that the system has software, firmware, hardware, or a combination thereof installed thereon that causes the system to perform those operations or actions in operation. For one or more computer programs configured to perform a specific operation or action, this means that the one or more programs include instructions that, when executed by a data processing device, cause that device to perform that operation or action.

[0111] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their equivalents), or in a combination of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals—e.g., machine-generated electrical, optical, or electromagnetic signals—to generate artificially generated propagation signals to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0112] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The device may also be or further include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.

[0113] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages); and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not need to, correspond to a file in a file system. A program may be stored as a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites and interconnected via a data communication network.

[0114] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same one or more computers.

[0115] The processes and logic flows described in this specification can be executed by one or more programmable computers, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., an FPGA or ASIC) or by a combination of a dedicated logic circuit system and one or more programmable computers.

[0116] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or incorporated into a special-purpose logic circuit system. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to or from them, or both. However, a computer does not necessarily have to have such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.

[0117] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM discs.

[0118] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's device in response to a request received from a web browser. Additionally, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user in response.

[0119] Data processing devices used to implement machine learning models may also include, for example, dedicated hardware accelerator units for general and compute-intensive parts of machine learning training or production (i.e., inference, workloads).

[0120] Machine learning models can be implemented and deployed using machine learning frameworks (such as TensorFlow or Jax).

[0121] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface, web browser, or app that a user can interact with through an implementation of the subject matter described in this specification), or any combination of one or more such back-end components, middleware components, or front-end components. Components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0122] A computing system may include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is established by computer programs executed on respective computers that establish a client-server relationship between them. In some embodiments, the server transmits data (e.g., HTML pages) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from that user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.

[0123] While this specification contains numerous details of specific implementations, these details should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from that combination may be removed from the claimed combination, and the claimed combination may involve sub-combinations or variations thereof.

[0124] Similarly, although operations are depicted in the accompanying drawings and described in a specific order in the claims, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or requiring all shown operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0125] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions described in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require a specific order or sequence shown to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.

Claims

1. A method executed by one or more computers, the method comprising: Receive data to identify the target material; An embedding machine learning model is used to generate the embedding of the target material in the latent space; as well as One or more nucleating agents for the target material are selected from a set of candidate nucleating agents, based at least in part on the embedding of the target material in the potential space.

2. The method as described in claim 1, wherein, The data for identifying the target material includes: (i) data for identifying the chemical composition of the target material, and (ii) data for identifying the crystal structure of the target material.

3. The method as described in any of the preceding claims, wherein, Using the embedding machine learning model to generate the embedding of the target material in the latent space includes: The embedding machine learning model is used to process the model input, including the data that identifies the target material, based on the values ​​of the embedding machine learning model parameter set, to generate the embedding of the target material.

4. The method as described in any of the preceding claims, wherein, The embedded machine learning model is an embedded neural network that has been jointly trained with a predictive neural network, wherein the predictive neural network is configured as follows: Receive the embedding of input material generated by the embedded neural network; The embedding of the input material is processed based on the values ​​of the set of parameters of the predictive neural network to generate a prediction characterizing the input material.

5. The method of claim 4, wherein, The prediction characterizing the input material includes a prediction of the energy of the input material.

6. The method as described in any one of claims 4 to 5, wherein, The predictions characterizing the input material include corresponding predictions of the forces acting on each atom in the unit cell of the input material.

7. The method according to any one of claims 4 to 6, wherein, The predictions characterizing the input material include the predicted reconstruction of the chemical composition and crystal structure of the input material.

8. The method according to any one of claims 4 to 7, wherein, Joint training of the embedded neural network and the prediction neural network includes: Obtain a set of training examples, wherein each training example corresponds to a corresponding training material and includes: (i) a training input representing the training material, and (ii) the target output of the predictive neural network; and The embedded neural network and the prediction neural network are jointly trained on the training example set.

9. The method of claim 8, wherein, Joint training of the embedded neural network and the prediction neural network on the training example set includes, for each training example: The embedded neural network is used to process the training input of the training examples to generate an embedding of the training material represented by the training input; The prediction neural network is used to process the embeddings of the training material to generate a prediction output; The gradient of an objective function is determined with respect to the set of parameters of the embedded neural network and the set of parameters of the predictive neural network, the objective function measuring the difference between: (i) the predicted output generated by the predictive neural network, and (ii) the target output specified by the training examples; as well as The gradient is used to adjust the values ​​of the embedded neural network parameter set and the predictive neural network parameter set.

10. The method as claimed in any of the preceding claims, wherein, Based at least in part on the embedding of the target material in the latent space, selecting one or more nucleating agents for the target material from the set of candidate nucleating agents includes: For each candidate nucleating agent in the candidate nucleating agent set, obtain the corresponding embedding of the candidate nucleating agent generated using the embedding machine learning model; For each candidate nucleating agent in the set of candidate nucleating agents, a corresponding distance is determined between the following two in the potential space: (i) the embedding of the candidate nucleating agent and (ii) the embedding of the target material; The selection of one or more nucleating agents for the target material is based at least in part on the distance.

11. The method of claim 10, wherein, Selecting one or more nucleating agents for the target material based at least in part on the distance includes: The candidate nucleating agents in the candidate nucleating agent set are ranked based on their respective distances from the target material in the potential space. The candidate nucleating agent set is filtered based on the ranking to remove one or more candidate nucleating agents from the candidate nucleating agent set; and One or more of the remaining candidate nucleating agents in the candidate nucleating agent set are selected as nucleating agents for the target material.

12. The method of claim 11, wherein, Selecting one or more remaining candidate nucleating agents from the candidate nucleating agent set as nucleating agents for the target material includes: For each remaining candidate nucleating agent in the candidate nucleating agent set, determine whether the candidate nucleating agent can stably coexist with the target material; and The candidate nucleating agent set is filtered to remove any candidate nucleating agents that cannot stably coexist with the target material. After filtering the candidate nucleating agent set, one or more of the remaining candidate nucleating agents in the candidate nucleating agent set are selected as nucleating agents for the target material based on the following: (i) The ranking of each candidate nucleating agent based on its corresponding distance from the target material in the potential space, and (ii) Whether each candidate nucleating agent can coexist stably with the target material.

13. The method as described in any of the preceding claims, wherein, The embedded machine learning model includes neural networks.

14. The method of claim 13, wherein, The embedded machine learning model includes a graph neural network, which comprises multiple message-passing neural network layers.

15. The method of any one of claims 13 to 14, further comprising generating graph data representing a graph included in the network input to the graph neural network, wherein: The graph includes multiple nodes and multiple edges; Each node in the diagram represents a corresponding atom in the unit cell of the target material; and Each edge connection in the diagram represents a corresponding pair of nodes in the unit cell of the target material where a pair of atoms are separated by a three-dimensional spatial distance less than a threshold.

16. The method of claim 15, wherein, Each node in the graph is associated with a corresponding set of node features that characterize the atom represented by the node; For each node in the graph, the node feature set of the node includes features that identify the element type of the atom.

17. The method of any of the preceding claims, further comprising providing each of the selected nucleating agents for synthesizing the target material.

18. The method of any of the preceding claims, further comprising, for each of one or more of the selected nucleating agents, physically synthesizing the target material by means of a physical synthesis technique involving nucleating the target material using the selected nucleating agent.

19. A system comprising: One or more computers; as well as One or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operation of a corresponding method as described in any one of claims 1 to 17.

20. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operation of the corresponding method as claimed in any one of claims 1 to 17.