Establishment of reinforcement learning environment and interaction method
By simulating the phase transition process of crystal structures through a reinforcement learning environment and defining transformation rules, the number of experiments is reduced, the cost of crystal material synthesis is lowered, and economical and efficient material synthesis is achieved.
Patent Information
- Application Number
- CN202210348057.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-03
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-04-03
AI Technical Summary
In the process of synthesizing crystal materials, existing technologies require multiple experimental verifications to obtain the crystal structure with the best performance, which leads to excessively high costs.
This paper proposes a method for building a reinforcement learning environment. By simulating the phase transition process of crystal structures, the transformation rules are defined and intermediate structures are obtained through computer simulation, thereby reducing the number of experiments.
By reducing the number of experiments, the cost of material synthesis was lowered, and economic efficiency was improved.
Smart Images

Figure CN114936510B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of reinforcement learning and material synthesis, and specifically provides a reinforcement learning environment building and interaction method. BACKGROUND
[0002] Different structures of crystals have different characteristics. In the process of crystal material synthesis, it is usually necessary to prepare a material with a crystal structure having optimal performance. The method is to change the existing crystal structure into the target crystal structure through a phase change process, which needs to be verified through multiple experiments, and the cost of experiments is too high. SUMMARY
[0003] The application is made to solve the above problems, and aims to provide a reinforcement learning environment building and interaction method.
[0004] The application provides a reinforcement learning environment building method for simulating a phase change process of a crystal structure, which has the following characteristics: obtaining all possible structures of the crystal structure undergoing the phase change process to build a simulated phase change environment; and defining a transformation rule based on the built simulated phase change environment.
[0005] The method provided by the application also has the following characteristics: the transformation rule includes a first transformation rule, a second transformation rule and a third transformation rule, the first transformation rule is used to determine whether the crystal structure obtained by changing the element type of the atomic site of each possible structure after encoding by element type meets a first predetermined range; the second transformation rule is used to determine whether the distance between the dimension reduction information of the visualized each possible structure and the transformable structure after transformation meets a second predetermined range; and the third transformation rule is used to determine whether the available energy value of the transformable structure meeting the first predetermined range and the second predetermined range is a positive value.
[0006] The method provided by the application also has the following characteristics: if the available energy value of the transformable structure is a positive value, the transformation is successful; if the available energy value of the transformable structure is a negative value, the structure remains unchanged, and other transformable structures are searched within a predetermined range.
[0007] The method provided by the application also has the following characteristics: the predetermined range includes the first predetermined range and the second predetermined range, the first predetermined range is all possible structures after encoding, and the second predetermined range is whether the distance between the dimension reduction information of the visualized each possible structure and the transformable structure after transformation is less than a predetermined distance.
[0008] The method provided by the present application also has the characteristics that if the distance between the dimension reduction information of each possible structure after visualization and the transformable structure after transformation is less than a predetermined distance, the second predetermined range is met, otherwise, another transformable structure is selected for the predetermined distance determination.
[0009] The method provided by the present application also has the characteristics that the predetermined distance is used to determine the similarity between the possible structures before and after transformation.
[0010] The method provided by the present application also has the characteristics that the transformable structure that is successfully transformed is determined as the first transformable structure, and the next transformable structure is determined in turn.
[0011] The method provided by the present application also has the characteristics that it is determined whether the transformable structure is an end point structure or whether the maximum number of transformations is met; if so, the transformation is ended, otherwise, the transformation rule is repeatedly executed.
[0012] The present application provides an interaction method of a reinforcement learning environment, which has the characteristics that all possible structures corresponding to the reinforcement learning environment are mapped into state representations and interacted with a reinforcement learning agent; or all possible structures corresponding to the reinforcement learning environment are quantum state encoded and mapped into state representations and interacted with the reinforcement learning agent, wherein the reinforcement learning environment is built by the building method of the reinforcement learning environment.
[0013] The method provided by the present application also has the characteristics that the quantum state encoding includes: normalizing the high-dimensional data of each possible structure to obtain a normalized vector; processing the normalized vector to obtain a quantum state right vector; obtaining a corresponding quantum state left vector by conjugate transposition of the quantum state right vector; and performing an outer product of the quantum state right vector and the quantum state left vector to obtain a structure density matrix obtained after quantum state encoding.
[0014] Effects of the present application
[0015] According to the building method of the reinforcement learning environment provided by the present application, the phase transition process of the crystal structure is simulated by a computer to obtain the intermediate structure of the phase transition process, and a transformation rule is defined, so that the experiment can be effectively assisted, the number of experiments can be reduced, the cost can be reduced, and the material synthesis discipline has high economic value. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is the overall flowchart of the building method of the reinforcement learning environment of the embodiment of the present application;
[0017] Figure 2 is a formation schematic diagram of the first predetermined range and the second predetermined range of the embodiment of the present application;
[0018] Figure 3 is a range diagram of the first predetermined range and the second predetermined range of the embodiment of the present application;
[0019] Figure 4 is a diagram of the first predetermined range of the embodiment of the present application;
[0020] Figure 5 is a diagram of the second predetermined range of the embodiment of the present application;
[0021] Figure 6 is a diagram of determining the first transformable structure of the embodiment of the present application;
[0022] Figure 7 is a quantum state encoding process diagram of the crystal structure. DETAILED DESCRIPTION
[0023] In order to make the technical means, creative features, purposes and effects of the present application easy to understand, the following will be specifically described in combination with embodiments and drawings.
[0024] <EMBODIMENT>
[0025] Figure 1 is a general flowchart of the method for building the reinforcement learning environment of the embodiment of the present application.
[0026] The general idea of the method for building the reinforcement learning environment provided by the embodiment is to build a simulated phase transition environment by obtaining all possible structures of the crystal structure undergoing phase transition; and define a transformation rule based on the built simulated phase transition environment, the transformation rule in the embodiment mainly includes three transformation rules, which are a first transformation rule, a second transformation rule and a third transformation rule.
[0027] As shown in Figure 1 , the method for building the reinforcement learning environment in the embodiment of the present application includes the following steps:
[0028] Step S1, obtaining all possible structures of the crystal undergoing phase transition and corresponding atomic three-dimensional space coordinate information and energy initial value information.
[0029] Specifically, the embodiment performs simulation calculation based on the DFT theory (density functional theory) to obtain all possible structure information of the crystal, at least including atomic three-dimensional space coordinate information and energy initial value information of each possible structure.
[0030] Figure 2 is a formation diagram of the first predetermined range and the second predetermined range of the embodiment of the present application; Figure 3 is a range diagram of the first predetermined range and the second predetermined range of the embodiment of the present application. Referring to Figures 2-3The formation and range of the first predetermined range and the second predetermined range in the embodiment are described.
[0031] In step S2, all possible structures are encoded by element type to form the first predetermined range.
[0032] Further, the first transformation rule in the embodiment is used to encode each possible structure by element type, and to determine whether the encoded crystal structure satisfies the first predetermined range, which is all possible structures corresponding to the crystal structure. Specifically, all encoded crystal structures have their own independent category representation. For example, if there are two variable elements A and B in the crystal and there are m atomic sites in common, a certain structure can be represented by the following encoding: (1, 1, -1…1), where the A element is encoded as -1 and the B element is encoded as 1. All possible encoded structures are saved as a two-dimensional array, where each row represents a possible structure representation and each column represents a variable position.
[0033] The structure transformation method of the first transformation rule in the embodiment is that the element type of some atomic sites of each possible structure after encoding is different, so that the structure transformation between different possible structures can be realized by changing the element type of the atomic sites of the possible structure. The first predetermined range in the embodiment is that based on the encoded each possible structure and the structure transformation method of the first transformation rule, that is, the allowed element type change of the atomic sites of each transformable structure, and the possible structure after changing the element type also belongs to the range of all possible structures specified.
[0034] In the embodiment, based on the obtained energy initial value information of each possible structure, the maximum difference between the energy initial value information of all possible structures is calculated as a reference value for simulating the energy supply of the phase change environment, and based on the predetermined heating method for synthesizing the experimental material, the single-step energy supply value of the simulated phase change environment is regulated.
[0035] In step S3, based on the formed first predetermined range, the atomic three-dimensional space coordinate information of each possible structure is processed by dimension reduction to calculate the similarity and form the second predetermined range.
[0036] Further, the second transformation rule in the embodiment is to calculate the similarity of each possible structure by setting a similarity radius. Specifically, in the embodiment, the atomic three-dimensional coordinate information of each possible structure is processed to obtain high-dimensional information corresponding to the atomic three-dimensional coordinate information, and then each group of high-dimensional information is processed by a manifold dimension reduction algorithm to obtain a two-dimensional vector corresponding to the atomic three-dimensional coordinate information. The above dimension reduction processing method is used to maintain a certain invariant feature quantity of high-dimensional data and low-dimensional data to find a reasonable low-dimensional feature representation, which is beneficial to visualization and reduces the calculation requirement. The corresponding two-dimensional data points (dimension reduction coordinates) of each possible structure are visualized, and a transformable range of mutual transformation between possible materials is specified. Referring to Figure 5 , the similarity radius r is determined according to the size of the two-dimensional plane after visualization, the density of points on the graph, and the heating mode of material synthesis (three aspects), and the scaling method of the r value can be adjusted according to the accuracy requirement through multiple simulations. Figure 5 In the formula, X1 and Xn are the minimum value and the maximum value of the dimension reduction coordinates of all possible structures, and the range of the transformation radius is set based on the maximum difference X (not shown in the figure) between the two, so that Figure 5 In the formula, the radius The specific transformation radius can be randomly set according to specific conditions, and the embodiment includes other transformation radii set based on the maximum difference X.
[0037] As shown in Figure 3 , in the embodiment, all possible structures that meet the conditions are first screened through the first predetermined range, and then the second predetermined range is formed by calculating the similarity within the first predetermined range.
[0038] In step S4, a crystal structure is randomly selected based on the formed first predetermined range and the second predetermined range for structure transformation.
[0039] Referring to Figures 4-6 , in the formed first predetermined range and the second range, a crystal structure is randomly selected as an initial structure, and the m atomic sites of each possible structure can be transformed by changing the element type of the atomic site. The embodiment exemplarily describes the atomic site 1. The B element of the atomic site 1 of the initial structure is transformed into the A element, and it is determined that the transformed structure 1 meets the first predetermined range. Then, based on the first predetermined range, it is judged whether the structure 1 is within the similarity radius r (r0) of the initial structure, and it is determined that the structure 1 is within the similarity radius r of the initial structure after judgment. The structure 1 is determined as the first transformable structure. Figure 5
[0040] In step S5, it is judged whether the transformed crystal structure meets the first predetermined range based on step S4, and if it meets, step S6 is executed; otherwise, step S4 is returned.
[0041] Step S6, it is judged whether the crystal structure satisfying the first predetermined range satisfies the second predetermined range, if yes, step S7 is executed; otherwise, step S4 is returned.
[0042] Step S7, it is judged whether the available energy value of the first to-be-transformable structure satisfying the first predetermined range and the second predetermined range is a positive value.
[0043] Step S8, if the available energy value of the crystal structure is a positive value, the structure transformation is successful this time, the structure is determined as the first transformable structure, the transformation structure range and the search radius area are switched, and the next transformable structure is determined in turn; otherwise, step S4 is returned.
[0044] Specifically, it is judged whether the available energy value of the first transformable structure is a positive value, if yes, the first transformable structure is transformed successfully. In the embodiment, the definition of energy is defined as follows: the material structure phase transition needs to absorb or release energy, and each structure has an initial energy value in the calculated generated data. At the same time, in the process of simulating the phase transition, an energy supplement value is also given to the environment according to the heating mode at the beginning of the experiment. The higher the available energy value of the structure at each moment, the larger the transformable range. The calculation method of the available energy is: the energy supplement value of the environment + the difference between the initial energy values of the structures before and after transformation. If the available energy value is negative, the transformation is not successful, and the original structure is maintained. At the same time, the constraint degree of the transformation rule of the energy supplement value is determined according to the type of the environment.
[0045] Step S9, steps S4-S8 are repeated, and it is judged whether the transformed crystal structure is the end point structure or whether the maximum transformation number is satisfied, if one of the two is satisfied, the structure transformation is ended; otherwise, step S4 is returned.
[0046] In addition, the embodiment also defines an environment reward function R for evaluating the transformation of the crystal structure, a good transformation will give an agent a positive feedback, and a bad transformation will give an agent a negative feedback. The good and bad of the transformation is defined according to the tendency of the phase transition of the crystal material in the specific experiment. According to the defined reward function, the algorithm iterates to affect how the agent selects the transformation path. After several iterations, the algorithm converges, and the selected transformation path often conforms to the process of the phase transition of the crystal material in the real environment. The specific setting method is divided into different cases according to whether the specific known starting structure and end point structure.
[0047] If the starting and end point structures are known, according to the general reinforcement learning environment setting, the reward for the successful structure transformation can be set to 1, and the reward for the remaining intermediate over-structure can be set to 0, so as to facilitate the algorithm to quickly iterate to find a reasonable phase transition path.
[0048] If the starting point or the end point is unknown, the energy absorption or release value in the phase transition process can be used as a reward setting reference value, and the structure satisfying the maximum number of transformations is defined as the end point structure.
[0049] The above two reward functions are only specific examples, and the specific reward function can be modified and adjusted according to the definition of the reward function. The present patent covers these modifications and adjustments.
[0050] The embodiment also provides a method for interacting the above built reinforcement learning environment with the quantum reinforcement learning algorithm, and the specific process is as follows: the built simulation phase change environment can judge whether the material can transform between different structures. Each state in the quantum reinforcement learning algorithm is defined as an encoding representation of a structure, and all structures form a state set of the simulation phase change system. A structure is obtained from the simulation environment at a certain moment, which is mapped to a representation, and the state representation is input to the reinforcement learning agent. The agent selects an output action from all available actions, the material transforms into another state, i.e. another structure, and obtains a reward feedback for iterative updating of the algorithm.
[0051] Referring to Figure 7 The embodiment also provides a quantum encoding method for encoding and representing atomic sites, and the specific process is as follows:
[0052] Step S1, normalizing the structure encoding representation with a length of n.
[0053] Step S2, converting the normalized vector into a complex number form, i.e. a quantum state right vector.
[0054] Step S3, taking the conjugate transpose of the nx1 right vector to obtain a 1xn left vector.
[0055] Step S4, performing outer product on the right vector and the left vector to obtain a nxn quantum state density matrix, and completing the quantum state encoding.
[0056] Effects of the embodiment
[0057] According to the method for building the reinforcement learning environment provided by the embodiment, because the method uses computer simulation of the phase change process of the crystal structure to obtain the intermediate structure of the phase change process and defines the transformation rule, the method can effectively assist the experiment, so that the number of experiments can be reduced and the cost can be reduced, and the method has high economic value for the material synthesis discipline.
[0058] Further, the embodiment encodes each possible structure according to the element type first, thereby determining the range of transformable structures, and then performs dimension reduction on the atomic three-dimensional space coordinates of each possible structure, and performs similarity judgment after visualization to further limit, so that after double constraints, the method can not only reduce the calculation amount required by the training iteration of the reinforcement learning algorithm, but also increase the reliability of the trajectory of the computer simulation material phase change process result, and has a high application prospect in the field of materials.
[0059] The above description and drawings give a typical embodiment of the specific structure of the specific implementation, and the above application proposes the preferred embodiment of the prior art, but these contents are not as limitations. For those skilled in the art, various changes and modifications will undoubtedly be apparent after reading the above description. Therefore, the appended claims should be considered as covering all changes and modifications within the true intent and scope of the application. Any and all equivalent ranges and contents within the scope of the claims should be considered as still within the intent and scope of the application.
Claims
1. A method for constructing a reinforcement learning environment to simulate the phase transition process of a crystal structure, characterized in that, include: To obtain all possible structures for a phase transition process in a crystal structure and thus construct a simulated phase transition environment; Transformation rules are defined based on the constructed simulated phase transition environment; The transformation rules include a first transformation rule, a second transformation rule, and a third transformation rule; The first transformation rule is used to determine whether the crystal structure obtained by encoding each possible structure by element type and changing the element type of the atomic sites meets the first predetermined range. The second transformation rule is used to determine whether the distance between the dimensionality reduction information after visualization of each possible structure and its transformed transformable structure satisfies a second predetermined range; The third transformation rule is used to determine whether the available energy value of the transformable structure that satisfies the first predetermined range and the second predetermined range is positive. If the available energy value of the transformable structure is positive, then the transformation is successful; If the available energy value of the transformable structure is negative, the structure remains unchanged, and other transformable structures are searched within a predetermined range. The predetermined range includes the first predetermined range and the second predetermined range. The first predetermined range is all possible structures after encoding. The second predetermined range is whether the distance between the dimensionality reduction information after visualization of each possible structure and its transformed transformable structure is less than a predetermined distance; If the distance between the dimensionality reduction information of each possible structure after visualization and its transformed transformable structure is less than a predetermined distance, then the second predetermined range is satisfied; otherwise, another transformable structure is selected again for the predetermined distance judgment. The predetermined distance is used to determine the similarity between possible structures before and after the transformation; Specifically, the transformable structure that is successfully transformed is determined as the first transformable structure, and the next transformable structure is determined sequentially. Among them, it is determined whether the transformable structure is an end point structure or whether it satisfies the maximum number of transformations; If yes, the transformation ends; otherwise, the transformation rule is repeated.
2. An interactive method for a reinforcement learning environment, characterized in that: After mapping all possible structures corresponding to the reinforcement learning environment into state representations, they interact with the reinforcement learning agent; or All possible structures corresponding to the reinforcement learning environment are encoded into quantum states, mapped to state representations, and then interacted with the reinforcement learning agent. The reinforcement learning environment is constructed using the reinforcement learning environment construction method described in claim 1.
3. The method according to claim 2, Its features are: The quantum state encoding includes: Normalize the high-dimensional data for each possible structure to obtain a normalized vector; The normalized vector is processed to obtain the quantum state right vector; The corresponding left quantum state is obtained by conjugating and transposing the right quantum state vector. The structure density matrix obtained by encoding the quantum state is obtained by taking the outer product of the right vector of the quantum state and the left vector of the quantum state.
Citation Information
Patent Citations
Systems and methods for training generative machine learning models
US20200401916A1
Generative design shape optimization using build material strength model for computer aided design and manufacturing
US20220004679A1