Polymer formula design method and device based on strategy learning

Through a deep reinforcement learning algorithm, the Stack-RNN formula generator is optimized, combined with the LSTM performance prediction model, and the problem of difficult polymer formulation design in the existing technology that meets practical applications is solved, and the generation of formulas that meet multiple target performances is achieved, and the application field is expanded.

CN120108557APending Publication Date: 2025-06-06烟台国工智能科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510171091.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing polymer formulation design methods based on optimization algorithms are difficult to generate formulas that meet practical applications, and there are too many types of formulas produced.

Method used

Using a strategy learning method based on deep reinforcement learning algorithm, the Stack-RNN formula generator and LSTM performance prediction model is used to optimize the recipe generation to generate formulas that meet multiple target performances.

Benefits of technology

Optimizing the Stack-RNN formula generator through deep reinforcement learning algorithms can generate polymer formulas that are more in line with practical applications, expanding the application field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108557A_ABST
    Figure CN120108557A_ABST
Patent Text Reader

Abstract

The invention discloses a macromolecular formula design method and device based on strategy learning. The method comprises the following steps: setting a data center to sort and obtain an initial data set; merging and converting a plurality of raw material matching fields in the initial data set into formula character string fields; arranging the formula character string field and the plurality of polymer performance fields to generate a preprocessed data set; the method comprises the following steps: constructing a Stack-RNN (Recurrent Neural Network) formula generator, and training the Stack-RNN formula generator by setting a training strategy to obtain a trained Stack-RNN formula generator; respectively constructing and training LSTM regression prediction models for the plurality of polymer properties through the preprocessing data set and an LSTM algorithm, and obtaining a plurality of performance predictors; combining the trained Stack-RNN formula generator with a plurality of performance predictors, and putting the combined Stack-RNN formula generator and the plurality of performance predictors into a reinforcement learning system; and enabling the Stack-RNN formula generator to generate a formula meeting a plurality of target performances through a reinforcement learning system. A deep reinforcement learning algorithm is used for optimization, so that the generated formula is more in line with practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of polymer formula design, and in particular to a polymer formula design method and device based on strategy learning. Background Art

[0002] Although conventional formulation design methods based on optimization algorithms can quickly converge to a better solution in a huge search space, the resulting formulations are often too numerous to be suitable for practical applications.

[0003] At present, after converting structured recipe data into string format, recurrent neural networks can be used to learn serialized data such as recipe strings, and combined with reinforcement learning methods to enable recurrent neural networks to generate recipes that meet multiple target performances, and the generated recipes are more suitable for practical applications.

[0004] Therefore, how to invent a polymer formula design method that can be optimized using a deep reinforcement learning algorithm to make the generated formula more suitable for practical applications has become an urgent problem to be solved. Summary of the invention

[0005] To this end, the present invention provides a polymer formula design method and device based on strategy learning, which optimizes the Stack-RNN formula generator through a deep reinforcement learning algorithm so that it can generate formulas that simultaneously meet multiple target performances, expand the application field and make the generated formula more suitable for practical applications.

[0006] In order to achieve the above object, the present invention provides the following technical solution: a polymer formulation design method based on strategy learning, comprising:

[0007] An initial data set is obtained by setting a data center for sorting; the initial data set includes several raw material ratio fields and several polymer performance fields;

[0008] Merge the several raw material ratio fields in the initial data set into a recipe string field according to a set rule, wherein the set rule includes separating different raw materials and ratios with a comma and separating the raw material type and ratio with a colon;

[0009] Arrange the formula string field and several polymer property fields to generate a preprocessing data set;

[0010] Constructing a Stack-RNN recipe generator, the Stack-RNN recipe generator comprising: an input layer, a GRU layer, a Stack layer and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeated characters in the recipe string field; and training the Stack-RNN recipe generator by setting a training strategy to obtain the trained Stack-RNN recipe generator;

[0011] By using the preprocessed data set and the LSTM algorithm, LSTM regression prediction models are constructed and trained for several polymer properties, each of which includes several LSTM layers, several fully connected hidden layers and one output layer, to obtain several performance predictors;

[0012] Merging the trained Stack-RNN recipe generator with a plurality of the performance predictors and placing them into a reinforcement learning system;

[0013] The Stack-RNN recipe generator is enabled by the reinforcement learning system to generate recipes that meet several target performances.

[0014] As a preferred solution for the polymer formula design method based on strategy learning, when constructing the Stack-RNN formula generator, the GRU layer is set with 1500 neurons, the Stack layer is set with 512 neurons, and the Stack layer has POP operation, PUSH operation and NO-OP operation. The POP operation is to delete the top element of the Stack layer, the PUSH operation is to insert a new element at the top of the Stack layer, and the NO-OP operation means not to perform any operation.

[0015] As a preferred solution of the polymer formula design method based on policy learning, in the process of merging the trained Stack-RNN formula generator with several performance predictors and placing them into the reinforcement learning system, the trained Stack-RNN formula generator is used as the agent of the reinforcement learning system; the action space of the agent is a character table composed of non-repeating characters in the formula string field; the state space of the agent is all possible character strings composed of the character table; several performance predictors are used as reviews of the reinforcement learning system; the output values ​​of several performance predictors are weighted summed by setting a reward function to obtain a sum value; and the sum value is used as a reward of the reinforcement learning system.

[0016] As a preferred solution for the polymer formulation design method based on policy learning, when the Stack-RNN formulation generator is trained using a reinforcement learning system, the policy gradient algorithm is used to update the parameters of the Stack-RNN formulation generator. The policy gradient loss function is:

[0017]

[0018] In the formula, R i is the reward at time step i; γ is the discount factor; p(s i |s 0 …s i-1; θ) is the probability distribution of the i-th character calculated based on the prefix sequence; θ is the parameter of the Stack-RNN recipe generator.

[0019] As a preferred solution of the polymer formulation design method based on strategy learning, the reward expression of the reinforcement learning system is:

[0020]

[0021] Where pred i is the output value of the i-th performance predictor; y i is the target value of the ith performance; n is the number of performances.

[0022] The present invention also provides a polymer formula design device based on strategy learning, which adopts the above-mentioned polymer formula design method based on strategy learning, comprising:

[0023] An initial data set acquisition module is used to obtain an initial data set by setting a data center; the initial data set includes several raw material ratio fields and several polymer performance fields;

[0024] A recipe string field conversion module, used for merging several raw material ratio fields in the initial data set into recipe string fields according to set rules, wherein the set rules include separating different raw materials and ratios with commas and separating raw material types and ratios with colons;

[0025] A preprocessing data set generation module, used for arranging the formula string field and several polymer property fields to generate a preprocessing data set;

[0026] A Stack-RNN recipe generator construction and training module is used to construct a Stack-RNN recipe generator, wherein the Stack-RNN recipe generator includes: an input layer, a GRU layer, a Stack layer, and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeated characters in the recipe string field; and the Stack-RNN recipe generator is trained by setting a training strategy to obtain the trained Stack-RNN recipe generator;

[0027] A performance predictor acquisition module is used to construct and train LSTM regression prediction models for several polymer properties respectively through the preprocessed data set and the LSTM algorithm, each of which includes several LSTM layers, several fully connected hidden layers and one output layer, to obtain several performance predictors;

[0028] A reinforcement learning system placement module, used to merge the trained Stack-RNN recipe generator and a plurality of the performance predictors, and place them into a reinforcement learning system;

[0029] A recipe generation module is used to enable the Stack-RNN recipe generator to generate recipes that meet several target performances through the reinforcement learning system.

[0030] As a preferred solution for the polymer formula design device based on strategy learning, in the Stack-RNN formula generator construction and training module, when constructing the Stack-RNN formula generator, the GRU layer is set with 1500 neurons, the Stack layer is set with 512 neurons, and the Stack layer has POP operation, PUSH operation and NO-OP operation. The POP operation is to delete the top element of the Stack layer, the PUSH operation is to insert a new element at the top of the Stack layer, and the NO-OP operation means not to perform any operation.

[0031] As a preferred solution of the polymer formula design device based on strategy learning, the reinforcement learning system is placed in a module. In the process of merging the trained Stack-RNN formula generator with several performance predictors and placing them in the reinforcement learning system, the trained Stack-RNN formula generator is used as the agent of the reinforcement learning system; the action space of the agent is a character table composed of non-repeating characters in the formula string field; the state space of the agent is all possible character strings composed of the character table; several performance predictors are used as reviews of the reinforcement learning system; the output values ​​of several performance predictors are weighted summed by setting a reward function to obtain a sum value; and the sum value is used as a reward for the reinforcement learning system.

[0032] As a preferred solution for the polymer formula design device based on policy learning, the reinforcement learning system is placed in the module. When the Stack-RNN formula generator is trained using the reinforcement learning system, the policy gradient algorithm is used to update the parameters of the Stack-RNN formula generator. The policy gradient loss function is:

[0033]

[0034] In the formula, R i is the reward at time step i; γ is the discount factor; p(s i |s 0 …s i-1 ; θ) is the probability distribution of the i-th character calculated based on the prefix sequence; θ is the parameter of the Stack-RNN recipe generator.

[0035] As a preferred solution of the polymer formula design device based on strategy learning, the reinforcement learning system is placed in the module, and the reward expression of the reinforcement learning system is:

[0036]

[0037] Where pred i is the output value of the i-th performance predictor; y i is the target value of the ith performance; n is the number of performances.

[0038] The present invention has the following advantages: the present invention obtains an initial data set by setting a data center; the initial data set includes several raw material ratio fields and several polymer performance fields; several raw material ratio fields in the initial data set are merged and converted into a recipe string field; the recipe string field and several polymer performance fields are sorted to generate a preprocessing data set; a Stack-RNN recipe generator is constructed, and the Stack-RNN recipe generator is trained by setting a training strategy to obtain the trained Stack-RNN recipe generator; LSTM regression prediction models are constructed and trained for several polymer properties by the preprocessing data set and LSTM algorithm to obtain several performance predictors; the trained Stack-RNN recipe generator is merged with several performance predictors and placed in a reinforcement learning system; the Stack-RNN recipe generator is enabled to generate recipes that meet several target performances by the reinforcement learning system. The present invention optimizes the Stack-RNN recipe generator through a deep reinforcement learning algorithm so that it can generate recipes that meet multiple target performances at the same time, which can expand the application field and make the generated recipes more in line with practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0040] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0041] Figure 1 A schematic flow chart of a polymer formulation design method based on strategy learning provided in Example 1 of the present invention;

[0042] Figure 2 A schematic diagram of a learning framework for a formulation generation strategy in a polymer formulation design method based on strategy learning provided in Example 1 of the present invention;

[0043] Figure 3 This is a schematic diagram of the architecture of a polymer formula design device based on strategy learning provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0044] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] Example 1

[0046] See also Figure 1 Embodiment 1 of the present invention provides a polymer formulation design method based on strategy learning, comprising the following steps:

[0047] S1. Obtaining an initial data set by setting a data center; the initial data set includes several raw material ratio fields and several polymer performance fields;

[0048] S2. Merge the several raw material ratio fields in the initial data set into a recipe string field according to a set rule, wherein the set rule includes separating different raw materials and ratios with a comma and separating the raw material type and ratio with a colon;

[0049] S3, arranging the formula character string field and several polymer property fields to generate a preprocessing data set;

[0050] S4, constructing a Stack-RNN recipe generator, the Stack-RNN recipe generator comprising: an input layer, a GRU layer, a Stack layer and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeated characters in the recipe string field; and training the Stack-RNN recipe generator by setting a training strategy to obtain the trained Stack-RNN recipe generator;

[0051] S5. Construct and train LSTM regression prediction models for several polymer properties respectively through the preprocessed data set and LSTM algorithm, each of the trained LSTM regression prediction models comprises several LSTM layers, several fully connected hidden layers and one output layer, and obtain several performance predictors;

[0052] S6, merging the trained Stack-RNN recipe generator with the plurality of performance predictors and placing them into a reinforcement learning system;

[0053] S7. Using the reinforcement learning system, the Stack-RNN recipe generator generates a recipe that meets several target performances.

[0054] In this embodiment, in step S1, an initial data set is obtained by setting a data center for sorting; the initial data set includes several raw material ratio fields and several polymer performance fields;

[0055] Specifically, an initial data set including several raw material ratio fields and several polymer performance fields is obtained by setting a data center.

[0056] In this embodiment, in step S2, several raw material ratio fields in the initial data set are merged and converted into a recipe character string field;

[0057] Specifically, all the raw material ratio fields in the initial data set are merged and converted into a formula string field. Formula string example: "R18:36.0,R6:25.0,F4:41.0,A4:0.2,A2:0.1,A3:0.2", different raw materials and ratios are separated by commas, and the raw material type and its ratio are separated by colons.

[0058] In this embodiment, in step S3, the formula string field and several polymer property fields are collated to generate a preprocessed data set;

[0059] Specifically, the formula character string field and several polymer property fields are combined into a preprocessing data set.

[0060] In this embodiment, in step S4, a Stack-RNN recipe generator is constructed, and the Stack-RNN recipe generator is trained by setting a training strategy to obtain the trained Stack-RNN recipe generator;

[0061] The Stack-RNN recipe generator includes: an input layer, a GRU layer, a Stack layer and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeating characters in the recipe string field.

[0062] Specifically, the input layer and the output layer each have 21 neurons, the GRU layer has 1500 neurons, and the Stack layer has 512 neurons. The Stack layer has three operations: POP, PUSH, and NO-OP. The POP operation refers to deleting the top element of the Stack layer, the PUSH operation refers to inserting a new element to the top of the Stack layer, and the NO-OP operation refers to not performing any operation.

[0063] In this embodiment, in step S5, LSTM regression prediction models are constructed and trained for several polymer properties respectively through the preprocessed data set and LSTM algorithm to obtain several performance predictors;

[0064] The LSTM regression prediction model includes: several LSTM layers, several fully connected hidden layers and an output layer.

[0065] Specifically, the LSTM regression prediction model has two LSTM layers and two fully connected hidden layers.

[0066] In this embodiment, in step S6, the trained Stack-RNN recipe generator is merged with a plurality of the performance predictors and placed into a reinforcement learning system;

[0067] Specifically, Figure 2 As shown in Figure 2, the trained Stack-RNN recipe generator is combined with several different performance predictors into a reinforcement learning system.

[0068] The trained Stack-RNN recipe generator is used as the agent of the reinforcement learning system; the action space of the agent is a character table consisting of non-repeated characters in the recipe string field;

[0069] Specifically, the character table is: {',','-','.','0','1','2','3','4','5','6','7','8','9',':','A','F','H','Q','R','T','W'}.

[0070] The state space of the agent is all possible character strings composed of the character table; a number of the performance predictors are used as the evaluation of the reinforcement learning system; the output values ​​of the several performance predictors are weighted summed by setting a reward function to obtain a sum value; and the sum value is used as the reward of the reinforcement learning system.

[0071] Specifically, the reward expression of the reinforcement learning system is:

[0072]

[0073] Where pred i is the output value of the i-th performance predictor; y i is the target value of the ith performance; n is the number of performances.

[0074] In this embodiment, the policy gradient algorithm is used to update the parameters of the Stack-RNN recipe generator, and the policy gradient loss function is:

[0075]

[0076] In the formula, R i is the reward at time step i; γ is the discount factor; p(s i |s 0 …s i-1 ; θ) is the probability distribution of the i-th character calculated based on the prefix sequence; θ is the parameter of the Stack-RNN recipe generator.

[0077] In this embodiment, in step S7, the Stack-RNN recipe generator is enabled by the reinforcement learning system to generate recipes that meet several target performances.

[0078] Specifically, Figure 2 As shown, the Stack-RNN recipe generator is enabled to generate recipes that meet multiple target performances by training the reinforcement learning system.

[0079] In a possible embodiment, a specific formulation design example is provided as follows:

[0080] Given the target values ​​of four product properties, namely tensile strength = 35.6, flexural strength = 42.5, MFR = 18.5, and density = 1.016, the recipe generator obtains a series of recipes after iterative optimization using the reinforcement learning algorithm. One of the recipe design solutions is shown in Table 1. Its tensile strength prediction value is 34.6 with a relative error of 2.7%, its flexural strength prediction value is 43.7 with a relative error of 2.7%, its MFR prediction value is 18.7 with a relative error of 1.1%, and its density prediction value is 1.015 with a relative error of 0.13%.

[0081]

[0082] Table 1 Product performance and formulation design parameters

[0083] In summary, the present invention obtains an initial data set by setting a data center; the initial data set includes several raw material ratio fields and several polymer performance fields; several raw material ratio fields in the initial data set are merged and converted into a recipe string field; the recipe string field and several polymer performance fields are sorted to generate a preprocessed data set; a Stack-RNN recipe generator is constructed, and the Stack-RNN recipe generator is trained by setting a training strategy to obtain the trained Stack-RNN recipe generator; LSTM regression prediction models are constructed and trained for several polymer properties respectively through the preprocessed data set and LSTM algorithm to obtain several performance predictors; the trained Stack-RNN recipe generator is merged with several performance predictors and placed in a reinforcement learning system; the Stack-RNN recipe generator is enabled to generate recipes that meet several target performances through the reinforcement learning system. The present invention optimizes the Stack-RNN recipe generator through a deep reinforcement learning algorithm so that it can generate recipes that meet multiple target performances at the same time, which can expand the application field and make the generated recipes more in line with practical applications.

[0084] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.

[0085] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0086] Example 2

[0087] See also Figure 3 Embodiment 2 of the present invention further provides a polymer formulation design device based on strategy learning, comprising:

[0088] An initial data set acquisition module 001 is used to obtain an initial data set by setting a data center; the initial data set includes several raw material ratio fields and several polymer performance fields;

[0089] A recipe string field conversion module 002 is used to merge the several raw material ratio fields in the initial data set into a recipe string field according to a set rule, wherein the set rule includes separating different raw materials and ratios with a comma, and separating the raw material type and ratio with a colon;

[0090] A preprocessing data set generation module 003 is used to sort the formula string field and several polymer property fields to generate a preprocessing data set;

[0091] The Stack-RNN recipe generator construction and training module 004 is used to construct a Stack-RNN recipe generator, wherein the Stack-RNN recipe generator includes: an input layer, a GRU layer, a Stack layer, and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeated characters in the recipe string field; and the Stack-RNN recipe generator is trained by setting a training strategy to obtain the trained Stack-RNN recipe generator;

[0092] The performance predictor acquisition module 005 is used to construct and train LSTM regression prediction models for several polymer properties respectively through the preprocessed data set and the LSTM algorithm, each of which includes several LSTM layers, several fully connected hidden layers and one output layer, to obtain several performance predictors;

[0093] A reinforcement learning system placement module 006 is used to merge the trained Stack-RNN recipe generator and a plurality of the performance predictors and place them into a reinforcement learning system;

[0094] The recipe generation module 007 is used to enable the Stack-RNN recipe generator to generate recipes that meet several target performances through the reinforcement learning system.

[0095] In this embodiment, in the Stack-RNN recipe generator construction and training module 004, when constructing the Stack-RNN recipe generator, the GRU layer is set with 1500 neurons, the Stack layer is set with 512 neurons, and the Stack layer has POP operation, PUSH operation and NO-OP operation. The POP operation is to delete the top element of the Stack layer, the PUSH operation is to insert a new element at the top of the Stack layer, and the NO-OP operation means not to perform any operation.

[0096] In this embodiment, the reinforcement learning system is placed in module 006. In the process of merging the trained Stack-RNN recipe generator with the several performance predictors and placing them in the reinforcement learning system, the trained Stack-RNN recipe generator is used as the agent of the reinforcement learning system; the action space of the agent is a character table composed of non-repeating characters in the recipe string field; the state space of the agent is all possible strings composed of the character table; the several performance predictors are used as the evaluation of the reinforcement learning system; the output values ​​of the several performance predictors are weighted summed by setting a reward function to obtain a sum value; and the sum value is used as the reward of the reinforcement learning system.

[0097] In this embodiment, the reinforcement learning system is placed in module 006. When the Stack-RNN recipe generator is trained using the reinforcement learning system, the policy gradient algorithm is used to update the parameters of the Stack-RNN recipe generator. The policy gradient loss function is:

[0098]

[0099] In the formula, R i is the reward at time step i; γ is the discount factor; p(s i |s 0 …s i-1 ; θ) is the probability distribution of the i-th character calculated based on the prefix sequence; θ is the parameter of the Stack-RNN recipe generator.

[0100] In this embodiment, the reinforcement learning system is placed in module 006, and the reward expression of the reinforcement learning system is:

[0101]

[0102] Where pred i is the output value of the i-th performance predictor; y i is the target value of the ith performance; n is the number of performances.

[0103] It should be noted that the information interaction, execution process and other contents between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.

[0104] Example 3

[0105] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code for a polymer formulation design method based on strategy learning is stored, and the program code includes instructions for executing a polymer formulation design method based on strategy learning of embodiment 1 or any possible implementation thereof.

[0106] The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0107] Example 4

[0108] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0109] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute a polymer formula design method based on strategy learning in Example 1 or any possible implementation thereof.

[0110] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor implemented by reading software codes stored in a memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.

[0111] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0112] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing system, they can be concentrated on a single computing system, or distributed on a network composed of multiple computing systems, and optionally, they can be implemented by a program code executable by a computing system, so that they can be stored in a storage system and executed by the computing system, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0113] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.

Claims

1. A polymer formulation design method based on strategy learning, characterized in that: include: An initial data set is obtained by setting a data center for sorting; the initial data set includes several raw material ratio fields and several polymer performance fields; Merge the several raw material ratio fields in the initial data set into a recipe string field according to a set rule, wherein the set rule includes separating different raw materials and ratios with a comma and separating the raw material type and ratio with a colon; Arrange the formula string field and several polymer property fields to generate a preprocessing data set; Constructing a Stack-RNN recipe generator, the Stack-RNN recipe generator comprising: an input layer, a GRU layer, a Stack layer and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeated characters in the recipe string field; and training the Stack-RNN recipe generator by setting a training strategy to obtain the trained Stack-RNN recipe generator; By using the preprocessed data set and the LSTM algorithm, LSTM regression prediction models are constructed and trained for several polymer properties, each of which includes several LSTM layers, several fully connected hidden layers and one output layer, to obtain several performance predictors; Merging the trained Stack-RNN recipe generator with a plurality of the performance predictors and placing them into a reinforcement learning system; The Stack-RNN recipe generator is enabled by the reinforcement learning system to generate recipes that meet several target performances.

2. A polymer formulation design method based on strategy learning according to claim 1, characterized in that: When constructing the Stack-RNN recipe generator, the GRU layer is set with 1500 neurons, the Stack layer is set with 512 neurons, and the Stack layer has POP operation, PUSH operation and NO-OP operation. The POP operation is to delete the top element of the Stack layer, the PUSH operation is to insert a new element at the top of the Stack layer, and the NO-OP operation means not to perform any operation.

3. A polymer formulation design method based on strategy learning according to claim 2, characterized in that: In the process of merging the trained Stack-RNN recipe generator with the several performance predictors and placing them into the reinforcement learning system, the trained Stack-RNN recipe generator is used as an agent of the reinforcement learning system; the action space of the agent is a character table composed of non-repeating characters in the recipe string field; the state space of the agent is all possible strings composed of the character table; the several performance predictors are used as evaluations of the reinforcement learning system; the output values ​​of the several performance predictors are weighted summed by setting a reward function to obtain a sum value; The summed value is used as a reward of the reinforcement learning system.

4. A polymer formulation design method based on strategy learning according to claim 3, characterized in that: When using the reinforcement learning system to train the Stack-RNN recipe generator, the policy gradient algorithm is used to update the parameters of the Stack-RNN recipe generator. The policy gradient loss function is: In the formula, R i is the reward at time step i; γ is the discount factor; p(s i |s0…s i-1 ; θ) is the probability distribution of the i-th character calculated based on the prefix sequence; θ is the parameter of the Stack-RNN recipe generator.

5. The polymer formulation design method based on strategy learning according to claim 3, characterized in that: The reward expression of the reinforcement learning system is: Where pred i is the output value of the i-th performance predictor; y i is the target value of the ith performance; n is the number of performances.

6. A polymer formula design device based on strategy learning, using a polymer formula design method based on strategy learning according to any one of claims 1 to 5, characterized in that: include: An initial data set acquisition module is used to obtain an initial data set by setting a data center; The initial data set includes several raw material ratio fields and several polymer property fields; A recipe string field conversion module, used for merging several raw material ratio fields in the initial data set into recipe string fields according to set rules, wherein the set rules include separating different raw materials and ratios with commas and separating raw material types and ratios with colons; A preprocessing data set generation module, used for arranging the formula string field and several polymer property fields to generate a preprocessing data set; A Stack-RNN recipe generator construction and training module is used to construct a Stack-RNN recipe generator, wherein the Stack-RNN recipe generator includes: an input layer, a GRU layer, a Stack layer, and an output layer; the number of neurons in the input layer and the output layer is the number of non-repeated characters in the recipe string field; and the Stack-RNN recipe generator is trained by setting a training strategy to obtain the trained Stack-RNN recipe generator; A performance predictor acquisition module is used to construct and train LSTM regression prediction models for several polymer properties respectively through the preprocessed data set and the LSTM algorithm, each of which includes several LSTM layers, several fully connected hidden layers and one output layer, to obtain several performance predictors; A reinforcement learning system placement module, used to merge the trained Stack-RNN recipe generator and a plurality of the performance predictors, and place them into a reinforcement learning system; A recipe generation module is used to enable the Stack-RNN recipe generator to generate recipes that meet several target performances through the reinforcement learning system.

7. The polymer formulation design device based on strategy learning according to claim 6 is characterized in that: In the Stack-RNN recipe generator construction and training module, when constructing the Stack-RNN recipe generator, the GRU layer is set with 1500 neurons, the Stack layer is set with 512 neurons, and the Stack layer has POP operation, PUSH operation and NO-OP operation. The POP operation is to delete the top element of the Stack layer, the PUSH operation is to insert a new element at the top of the Stack layer, and the NO-OP operation means not to perform any operation.

8. The polymer formulation design device based on strategy learning according to claim 7, characterized in that: The reinforcement learning system is placed in a module. In the process of merging the trained Stack-RNN recipe generator with the several performance predictors and placing them in the reinforcement learning system, the trained Stack-RNN recipe generator is used as an agent of the reinforcement learning system; the action space of the agent is a character table composed of non-repeating characters in the recipe string field; the state space of the agent is all possible strings composed of the character table; the several performance predictors are used as evaluations of the reinforcement learning system; the output values ​​of the several performance predictors are weighted summed by setting a reward function to obtain a sum value; The summed value is used as a reward of the reinforcement learning system.

9. The polymer formulation design device based on strategy learning according to claim 8, characterized in that: The reinforcement learning system is placed in the module. When the Stack-RNN recipe generator is trained using the reinforcement learning system, the policy gradient algorithm is used to update the parameters of the Stack-RNN recipe generator. The policy gradient loss function is: In the formula, R i is the reward at time step i; γ is the discount factor; p(s i |s0…s i-1 ; θ) is the probability distribution of the i-th character calculated based on the prefix sequence; θ is the parameter of the Stack-RNN recipe generator.

10. The polymer formulation design device based on strategy learning according to claim 9, characterized in that: The reinforcement learning system is placed in the module, and the reward expression of the reinforcement learning system is: Where pred i is the output value of the i-th performance predictor; y i is the target value of the ith performance; n is the number of performances.

Citation Information

Patent Citations

  • Drug molecule generator training method based on domain knowledge and deep reinforcement learning

    CN113223637A

  • Multi-target attribute molecule generation method and system based on strategy learning

    CN114974461A

  • Charging pile sales volume model training method and device, equipment and storage medium

    CN118096244A

  • Systems and methods for training and leveraging a multi-headed machine learning model for predictive actions in a complex prediction domain

    US20240256988A1