A Reinforcement Learning-Based Method and System for Optimizing Manufacturing Process Parameters
By optimizing process parameters using the Q-learning algorithm based on reinforcement learning, the problem of time-consuming and labor-intensive process parameter adjustment in traditional methods is solved, realizing automated and real-time adjustment, and improving product quality and production efficiency.
Patent Information
- Application Number
- CN202311346848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-10-18
AI Technical Summary
In industrial manufacturing, the adjustment of process parameters relies on engineers' experience, which is time-consuming and labor-intensive, resulting in inconsistent product quality and an inability to respond to dynamic and complex changes in the production environment in real time.
A reinforcement learning-based approach is adopted, using the Q-learning algorithm to establish a Markov policy, and then using the state space, action space and reward function to optimize process parameters, thereby achieving automated and real-time adjustment.
It improves product quality consistency and production efficiency, reduces manual intervention, and can respond to changes in complex production environments in real time, meeting the needs of high-precision products.
Smart Images

Figure CN117420800B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of parameter optimization technology, specifically to a method and system for optimizing manufacturing process parameters based on reinforcement learning. Background Technology
[0002] In the industrial manufacturing sector, high-quality manufactured products heavily rely on optimal process parameter settings. Process parameter optimization is a crucial task that can help improve production efficiency, reduce costs, and enhance product quality. In actual industrial scenarios, the impact of process parameter values on part quality is subject to various uncertainties and randomness. Multiple key production process parameters need to be pre-set during part manufacturing, and the rationality of these settings directly affects the final product quality. Currently, process parameters in actual industrial production are mainly determined based on engineers' experience and manually adjusted through trial and error, which is time-consuming and labor-intensive, resulting in inconsistent product quality. The relationship between product quality and process parameters is often ambiguous and influenced by complex factors from various variables, making the determination of process parameters for high-quality products challenging. This is especially true for products with extremely high dimensional accuracy requirements, where the process window is often narrow and irregular, making it more difficult to obtain appropriate process parameter settings than for general products. Therefore, it is crucial to quickly find effective process parameter optimization methods to solve the challenges of producing high-precision products.
[0003] Traditional methods for adjusting process parameters rely on engineers' experience, selecting optimal parameters through repeated trial and error by on-site engineers. However, these methods have several problems. First, they depend on prior knowledge, which may limit their applicability in dynamic environments. Second, they are insufficient for automation, requiring manual intervention and multiple trials. Furthermore, these methods are often static, optimized based on fixed datasets, and cannot adapt to the changes and challenges that may arise in dynamic and complex manufacturing processes.
[0004] To address this problem, existing technologies have proposed several solutions; some of these technologies are based on experimental design and statistical process mapping, seeking optimal or feasible parameters through finite sampling and analysis of the parameter space. However, these methods have some limitations. Existing technologies require a large amount of initial data to establish the relationship between parameters and performance, which becomes a limiting factor in some fields such as injection molding and metal additive manufacturing where data generation costs are high. Secondly, most existing technologies tend to use static models, which cannot adapt well to dynamic and complex production environments.
[0005] In addition, existing technologies often lack real-time or online feedback mechanisms; they cannot adjust parameters based on real-time process monitoring or quality inspection results, resulting in an inability to respond promptly to changes and fluctuations in the manufacturing process. Summary of the Invention
[0006] In view of the shortcomings of the prior art described above, the present invention provides a method and system for optimizing manufacturing process parameters based on reinforcement learning, which is used to replace traditional manual parameter tuning to improve product quality and production efficiency, better adapt to dynamic and complex production environments, adjust parameters based on real-time process monitoring or quality inspection results, and promptly respond to changes and fluctuations in the manufacturing process.
[0007] To achieve the above effects, the technical solution of the present invention is as follows:
[0008] In a first aspect, the present invention provides a method for optimizing manufacturing process parameters based on reinforcement learning, comprising the following steps:
[0009] Step 1: Collect process parameter data of the production system;
[0010] Step 2: Establish the Markov policy for the reinforcement learning system, specifically:
[0011] Establish a Markov strategy for reasoning and decision-making based on manufacturing process parameters. The Markov strategy represents the state of the product and the actions to be taken based on the process parameters during the manufacturing process. a The reward function R is defined under the product transfer state and process parameter input; the relationship between process parameters and product quality is estimated, and the optimization problem of manufacturing process parameters is defined as maximizing reward; the agent in the Markov policy is a machine that performs the task.
[0012] Step 3: Select the Q-learning reinforcement learning algorithm, specifically:
[0013] In the Q-learning reinforcement learning algorithm, a Q-table is created, with the dimension being the state space. S Action space A Used to store the Q value;
[0014] Step 4: Training and Learning: Train the Q-table using a reinforcement learning algorithm;
[0015] Training process: At time t, the agent selects an action. a Get the reward and the state at the next moment, and update the Q-table; keep iterating until the maximum number of iterations is reached or a converged Q-table is obtained;
[0016] Step 5: After training, query the Q-table to select the optimal action sequence, use the Q-table to select the action with the maximum Q value, and obtain the optimal process parameters for output.
[0017] Furthermore, step 2, which involves establishing a Markov strategy for reasoning and decision-making based on manufacturing process parameters, specifically includes:
[0018] Define state space SThe state space represents the current state of the reinforcement learning system, including the machine's temperature, material type, injection speed, and holding time. S This represents the current state of the reinforcement learning system; in the injection molding process parameter optimization problem, the state includes the injection speed. v Holding time t and injection pressure P ;
[0019] Define action space A Action space A represents the actions taken further during the optimization of process parameters, i.e., defining how to adjust the process parameters; in the injection molding process, actions include the changes in injection speed, holding time, and injection pressure after adjustment, i.e., Δ. v Δ t and Δ P The action space is the set of all possible actions, and consists of continuous or discretized values; the state space is the injection speed. v Holding time t and injection pressure P The range of values that can change;
[0020] Define reward function R The reward function metric measures the immediate reward an agent receives after performing an action in a given state; it defines the reward function and calculates the reward value based on product quality indicators.
[0021] If the product quality is lower than the preset target quality, the reward will be negative; if the product quality is within the preset target quality range, the reward will be zero or a small positive value.
[0022] Determine the state transition probability: simulate the impact of process parameter adjustments on product quality;
[0023] Define the objective function J: The objective function J is used to represent maximizing the cumulative total reward.
[0024] Furthermore, step 2 defines the optimization problem of manufacturing process parameters as maximizing the reward, specifically as follows:
[0025] Define the optimization objective, environment, state, action, and reward for the manufacturing process parameter optimization problem; where the manufacturing process is the injection molding process of part processing; the optimization objective represents optimizing the process parameters in the part processing manufacturing process; the environment represents the part processing production environment, including the setting range of process parameters and their impact on part quality; the state represents the current setting value of the process parameters; the action represents adjusting the process parameters; and the reward represents the maximum reward obtained in the manufacturing process by measuring the product quality indicators.
[0026] Furthermore, the state space is the injection speed.v Holding time t and injection pressure P The range of values for the change is expressed as follows:
[0027]
[0028] In the formula, {“{ min}”}、{“{ max}”} represent the minimum value and the maximum value, respectively.
[0029] Furthermore, the action space A is represented as:
[0030] .
[0031] Furthermore, in the injection molding problem, the reward function... R The difference between the actual shrinkage rate and the expected shrinkage rate is expressed as:
[0032]
[0033] In the formula, s This represents the rate of contraction in the current state. s target This represents the expected shrinkage rate. a Indicates the current action.
[0034] Furthermore, in the process problem of injection molding manufacturing, the objective function J is to maximize the cumulative total reward, that is:
[0035]
[0036] In the formula, γ represents the discount factor.
[0037] Furthermore, the update rule for the Q table is: update the Q value based on the Bellman equation, that is:
[0038]
[0039] In the formula, α represents the learning rate. s ′ represents the next state. a ′ represents the action for the next state.
[0040] Secondly, this invention provides a reinforcement learning-based system for optimizing manufacturing process parameters, comprising:
[0041] The data acquisition module is used to collect process parameter data from the production system.
[0042] The Markov policy building module is used to build Markov policies for reinforcement learning systems, specifically:
[0043] Establish a Markov strategy for reasoning and decision-making based on manufacturing process parameters. The Markov strategy represents the state of the product and the actions to be taken based on the process parameters during the manufacturing process. a The reward function R is defined under the product transfer state and process parameter input; the relationship between process parameters and product quality is estimated, and the optimization problem of manufacturing process parameters is defined as maximizing reward; the agent in the Markov policy is a machine that performs the task.
[0044] The reinforcement learning algorithm setting module is used to select the Q-learning reinforcement learning algorithm, specifically:
[0045] In the Q-learning reinforcement learning algorithm, a Q-table is created, with the dimension being the state space. S Action space A Used to store the Q value;
[0046] The training module is used for training and learning: it trains the Q-table using a reinforcement learning algorithm;
[0047] Training process: At time t, the agent selects an action. a Get the reward and the state at the next moment, and update the Q-table; keep iterating until the maximum number of iterations is reached or a converged Q-table is obtained;
[0048] The process parameter optimization module is used to query the Q-table after training to select the optimal action sequence, use the Q-table to select the action with the maximum Q value, and obtain the optimal process parameters for output.
[0049] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0050] 1. Improve product quality: In existing technologies, process parameters are adjusted based on engineers' experience or manual trial and error, resulting in inconsistent product quality. In contrast, this invention establishes a Markov strategy for a reinforcement learning system and introduces Q-learning reinforcement learning methods to automatically optimize the input of optimal process parameters based on the designed product quality goals, making product quality more stable and controllable.
[0051] 2. Reduced manual intervention: In traditional methods, adjusting process parameters requires a lot of manual intervention and monitoring, which is time-consuming and resource-intensive; however, the present invention introduces reinforcement learning to automatically optimize process parameters, reducing the risk of human error and improving production efficiency.
[0052] 3. Addressing Complexity: In industrial production, process parameters are multidimensional and complexly interconnected. This invention introduces reinforcement learning, which can effectively address the multidimensionality and complexity of process parameters and select the optimal parameters to meet product quality requirements.
[0053] 4. Real-time optimization: Reinforcement learning can optimize process parameters under real-time monitoring and make adjustments based on constantly changing production conditions and demands to maintain product quality. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the manufacturing process parameter optimization method based on reinforcement learning according to the present invention. Detailed Implementation
[0055] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0056] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0057] Example
[0058] This embodiment proposes a method for optimizing manufacturing process parameters based on reinforcement learning. Please refer to [link / reference]. Figure 1 This includes the following steps:
[0059] Step 1: Collect process parameter data of the production system;
[0060] Step 2: Establish the Markov policy for the reinforcement learning system, specifically:
[0061] To facilitate subsequent use of reinforcement learning for process parameter optimization decisions, a Markov policy for reasoning and decision-making regarding manufacturing process parameters is established. This Markov policy represents the state of the product and the proposed actions to be taken based on the process parameters during the manufacturing process. a The reward function R is defined under the product transfer state and process parameter input; the relationship between process parameters and product quality is estimated, and the optimization problem of manufacturing process parameters is defined as maximizing reward; the agent in the Markov policy is a machine that performs the task.
[0062] Step 3: Select the Q-learning reinforcement learning algorithm, specifically:
[0063] In the Q-learning reinforcement learning algorithm, a Q-table is created, with the dimension being the state space. S Action space A Used to store the Q value;
[0064] Step 4: Training and Learning: Use reinforcement learning algorithms to train the Q-table to learn to take appropriate actions in different states to maximize rewards;
[0065] Training process: At time t, the agent selects an action. a Get the reward and the state at the next moment, and update the Q-table; keep iterating until the maximum number of iterations is reached or a converged Q-table is obtained;
[0066] Step 5: Output the optimal injection molding process parameters:
[0067] After training, the optimal action sequence is selected by querying the Q-table, and the action with the maximum Q-value is chosen from the Q-table to obtain the optimal process parameters for output. The system inputs the optimized parameters and outputs the desired product quality indicators, such as the optimal process parameters for maximizing shrinkage during injection molding.
[0068] This invention employs a Q-learning reinforcement learning algorithm to optimize manufacturing process parameters. Q-learning is a value-iterative reinforcement learning algorithm used to learn a Q-table, which stores the estimated value of each action in each state. The Q-learning algorithm continuously updates the Q-table to select the optimal action to maximize cumulative reward. The reinforcement learning algorithm in this invention is not limited to Q-learning; it can also use deep reinforcement learning algorithms such as DQN, DDPG, PPO, SARSA, and Actor-Critic, or policy gradient-based methods to optimize the decision model established in the early stages of this invention. Addressing the need for producing high-precision products, this invention overcomes the problems of data intensity, static models, and lack of real-time feedback, enabling reliable manufacturing of high-quality products. It better adapts to dynamic and complex production environments, adjusting parameters based on real-time process monitoring or quality inspection results to promptly respond to changes and fluctuations in the manufacturing process.
[0069] As a preferred technical solution, in this embodiment, step 2, establishing a Markov strategy for reasoning and decision-making based on manufacturing process parameters, specifically involves:
[0070] Define state space S The state space represents the current state of the reinforcement learning system, including the machine's temperature, material type, injection speed, and holding time. S This represents the current state of the reinforcement learning system; in the injection molding process parameter optimization problem, the state includes the injection speed.v Holding time t and injection pressure P ;
[0071] Define action space A Action space A represents the actions taken further during the optimization of process parameters, defining how to adjust the process parameters, such as increasing or decreasing each parameter by a certain percentage. In the injection molding process, actions include the changes in injection speed, holding time, and injection pressure, i.e., Δ. v Δ t and Δ P The action space is the set of all possible actions, and consists of continuous or discretized values; the state space is the injection speed. v Holding time t and injection pressure P The range of values that can change;
[0072] Define reward function R The reward function metric measures the immediate reward an agent receives after performing an action in a given state; it defines the reward function and calculates the reward value based on product quality indicators.
[0073] The reward function calculates the reward based on the difference between the part size and the target size; if the size is close to the target, the reward is higher; if the size deviates from the target, the reward is lower; during the injection molding process, if the product quality meets the target quality requirements, the reward is positive, indicating that it is close to the target quality; if the product quality is lower than the preset target quality, the reward is negative; if the product quality is within the preset target quality range, the reward is zero or a small positive value.
[0074] Determine the state transition probability: The transition probability depends on the action taken; simulate the impact of process parameter adjustments on product quality, using physical models or historical data for estimation; in the injection molding process, increasing the temperature setting may lead to easier product molding and higher quality; increasing the pressure setting may lead to higher product density, but may also cause deformation or defects; increasing the injection speed setting will affect the surface smoothness of the product; adjusting the cooling time setting may affect the hardness and cooling shape of the product;
[0075] Define the objective function J: The objective function J is used to represent maximizing the cumulative total reward.
[0076] The Markov model employs agent representation to optimize decision-making.
[0077] As a preferred technical solution, in this embodiment, step 2, which defines the optimization problem of manufacturing process parameters as maximizing the reward, specifically means:
[0078] This section defines the optimization objective, environment, state, action, and reward for optimizing manufacturing process parameters. The manufacturing process refers to the injection molding process for part processing. The optimization objective represents optimizing the process parameters during part manufacturing to achieve the required product quality standards. The environment refers to the production environment for part processing, including the setting range of process parameters and their impact on part quality. The state represents the current set value of the process parameters. Actions represent adjustments to the process parameters, such as increasing or decreasing the value of a parameter. The reward represents the reward obtained by measuring product quality indicators, such as dimensional accuracy and surface finish. The final optimization objective is set as the maximum reward obtained during the manufacturing process to achieve the goal of optimizing product quality.
[0079] As a preferred technical solution, in this embodiment, the state space is the injection speed. v Holding time t and injection pressure P The range of values for the change is expressed as follows:
[0080]
[0081] In the formula, {“{ min}”}、{“{ max}”} represent the minimum value and the maximum value, respectively.
[0082] The state space is the set of all possible states and is a combination of continuous or discrete variables.
[0083] As a preferred technical solution, in this embodiment, the action space A is represented as:
[0084] .
[0085] As a preferred technical solution, in this embodiment, in the injection molding problem, the reward function... R The difference between the actual shrinkage rate and the expected shrinkage rate is expressed as:
[0086]
[0087] In the formula, s This represents the rate of contraction in the current state. s target This represents the expected shrinkage rate. a Indicates the current action.
[0088] As a preferred technical solution, in this embodiment, in the process problem of injection molding manufacturing, the objective function J is to maximize the accumulation of total reward, that is:
[0089]
[0090] In the formula, γ represents the discount factor.
[0091] As a preferred technical solution, in this embodiment, the update rule for the Q table is: to update the Q value based on the Bellman equation, that is:
[0092]
[0093] In the formula, α represents the learning rate. s ′ represents the next state. a ′ represents the action for the next state.
[0094] It should be noted that, in one specific embodiment, at the injection pressure P When the process parameters remain constant, the optimization problem of the manufacturing process is defined as an injection molding problem, namely, how to select the optimal injection speed and holding time to minimize the shrinkage of the molded product. The optimization problem is modeled as a reinforcement learning problem, and the specific steps are as follows:
[0095] 1) Define the state space: using binary tuples ( v, t () indicates the current state. v Indicates the injection speed. t Indicates the pressure holding time; limits are imposed according to actual conditions. v and t The range of values for , v =[0.1,1.0] m / s ,t=[5,15] s Set a certain distance from the step size, such as 0.1. m / s and 1 s ;
[0096] 2) Define the action space: using a binary tuple (Δ) v, Δ t () indicates the current action a Δ v Δ represents the change in injection rate. t This indicates the change in holding time; Δ is limited according to the actual situation. v and Δ t The range of values for Δ v =[-0.1,0.1] m / s ,Δt=[-1,1] s, Set a certain distance from the step size, 0.01 m / s and 0.1 s ;
[0097] 3) Define the reward function: Use the shrinkage rate of the molded product as the negative value of the reward function, i.e. r =- s ( v, t ),in s ( v, t) indicates the injection speed v The shrinkage rate is determined by the holding time t; s ( v, t The expression for ) is, s ( v, t =0.01 + 0.05e {-v*t} ;
[0098] 4) Define the transition probability: Assume that the state transition is deterministic, i.e., the execution of the action (Δ) v, Δ t After that, the state changes from ( v, t ) becomes ( v+ Δ v, t + Δ t );
[0099] 5) Select reinforcement learning algorithm: Select Q-learning reinforcement learning algorithm;
[0100] 6) Interactive learning with the environment: By interacting with the environment (i.e., the injection molding machine), data is collected, the value function or strategy function is updated, and continuous exploration and utilization are carried out until an optimal or suboptimal value function or strategy function is converged.
[0101] 7) Output optimal or suboptimal injection molding process parameters: Based on the obtained value function or strategy function, output an optimal or suboptimal injection molding process parameter, which yields the injection speed with the minimum shrinkage rate. v′ and holding time t′ ;
[0102] As a preferred technical solution, in this embodiment, based on the state space, action space, reward function, and transition probability defined above, a reinforcement learning algorithm based on Q-learning is designed to optimize injection molding process parameters; the specific steps are as follows:
[0103] Step 4.1, Initialize the Q-table: Create a Q-table to store the Q-value of each action in each state; the size of the Q-table is the product of the state space and the action space; initially, initialize all Q-values to 0 or a very small random value;
[0104] Step 4.2, Select Action: Based on the current state ( v, t Find the Q value of all corresponding actions in the Q table, and select an action according to the preset strategy; the preset strategy includes the greedy strategy (select the action with the largest Q value) and the ε greedy strategy (randomly select an action with a certain probability ε, and select the action with the largest Q value with a probability of 1-ε).
[0105] Step 4.3, Execute the action: Based on the selected action (Δ) v, Δ t), calculate the state at the next time step ( v+ Δ v, t + Δ t );
[0106] Step 4.4, Update the Q table: Based on the reward function r =- s ( v, t ) and transition probabilities, to calculate the reward obtained at the current time. r and the next state ( v+ Δ v, t + Δ t Maximum Q value maxQ Update the current state according to the update formula of the Q-learning algorithm. v, t ) Execute action (Δ) v, Δ t The Q value of ) is expressed as:
[0107]
[0108] In the formula, α It's the learning rate. γ It is a discount factor;
[0109] Step 4.5: Repeat steps 2-4 above until the number of iterations reaches the preset value or the Q-table converges.
[0110] The following describes the reinforcement learning-based manufacturing process parameter optimization system provided by the embodiments of the present invention. The reinforcement learning-based manufacturing process parameter optimization system described below can be referred to in conjunction with the reinforcement learning-based manufacturing process parameter optimization method described above.
[0111] The manufacturing process parameter optimization system based on reinforcement learning provided in this embodiment of the invention includes:
[0112] The data acquisition module is used to collect process parameter data from the production system.
[0113] The Markov policy building module is used to build Markov policies for reinforcement learning systems, specifically:
[0114] Establish a Markov strategy for reasoning and decision-making based on manufacturing process parameters. The Markov strategy represents the state of the product and the actions to be taken based on the process parameters during the manufacturing process. a The reward function R is defined under the product transfer state and process parameter input; the relationship between process parameters and product quality is estimated, and the optimization problem of manufacturing process parameters is defined as maximizing reward; the agent in the Markov policy is a machine that performs the task.
[0115] The reinforcement learning algorithm setting module is used to select the Q-learning reinforcement learning algorithm, specifically:
[0116] In the Q-learning reinforcement learning algorithm, a Q-table is created, with the dimension being the state space. S Action space A Used to store the Q value;
[0117] The training module is used for training and learning: it trains the Q-table using a reinforcement learning algorithm;
[0118] Training process: At time t, the agent selects an action. a Get the reward and the state at the next moment, and update the Q-table; keep iterating until the maximum number of iterations is reached or a converged Q-table is obtained;
[0119] The process parameter optimization module is used to query the Q-table after training to select the optimal action sequence, use the Q-table to select the action with the maximum Q value, and obtain the optimal process parameters for output.
[0120] Compared with the prior art, the method of the present invention has the following advantages:
[0121] Resource and cost savings: Automatic optimization of process parameters can save resources and reduce costs by reducing scrap rates, improving product quality and consistency, and enhancing the economics of production.
[0122] Data-driven decision-making: Reinforcement learning algorithms can make decisions based on real data and feedback, which is data-driven and helps to better understand changes and trends in the manufacturing process.
[0123] Reinforcement learning-based methods for optimizing manufacturing process parameters improve the efficiency and quality of the manufacturing process and reduce the need for human intervention; they can optimize process parameters in real time or offline to meet production needs and quality standards.
[0124] The process parameter optimization method of this invention has broad application prospects in the manufacturing industry, especially in highly customized production environments; it can automatically optimize processes in fields such as semiconductor manufacturing, 3D printing, metal processing, chemical and food production to improve production efficiency and quality.
[0125] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for optimizing manufacturing process parameters based on reinforcement learning, characterized in that, The following steps are involved: Step 1: Collect process parameter data of the production system; Step 2: Establish the Markov policy for the reinforcement learning system, specifically: Establish a Markov strategy for reasoning and decision-making based on process parameters in the manufacturing process. The Markov strategy represents the state of the product, the action to be taken by the process parameters a, the product transfer state, and the reward function R under the input of the process parameters in the manufacturing process. The relationship between process parameters and product quality is estimated, and the optimization problem of manufacturing process parameters is defined as maximizing reward; the agent in the Markov policy is a machine that performs the task. Step 2, which involves establishing a Markov strategy for reasoning and decision-making based on manufacturing process parameters, specifically includes: Define the state space S: The state space represents the current state of the reinforcement learning system, including the machine temperature, material type, injection speed, and holding time; the state space S represents the current state of the reinforcement learning system; in the injection molding process parameter optimization problem, the state includes the injection speed v, holding time t, and injection pressure P; Define action space A: Action space A represents the actions taken further in the process of optimizing process parameters, that is, defining how to adjust process parameters; in the injection molding process, actions include the changes after adjusting injection speed, holding time and injection pressure, namely Δv, Δt and ΔP; The action space is the set of all possible actions, which are continuous or discretized values; the state space is the range of values for the injection speed v, holding time t, and injection pressure P. Define a reward function R: The reward function metric is used to measure the immediate reward obtained by an agent after performing an action in a certain state; define a reward function to calculate the reward value based on product quality indicators; the product quality indicators include the shrinkage rate of the molded product; If the product quality is lower than the preset target quality, the reward will be negative; if the product quality is within the preset target quality range, the reward will be zero or a small positive value. Determine the state transition probability: simulate the impact of process parameter adjustments on product quality; Define the objective function J: The objective function J is used to represent maximizing the cumulative total reward; Step 3: Select the Q-learning reinforcement learning algorithm, specifically: In the Q-learning reinforcement learning algorithm, a Q-table is created, with the dimension of the state space S and the action space A used to store the Q-values. Step 4: Training and Learning: Train the Q-table using a reinforcement learning algorithm; Training process: At time t, the agent selects an action a, obtains the reward and the state for the next time step, and updates the Q-table; Continue iterating until the maximum number of iterations is reached or a converged Q-table is obtained; Step 5: After training, query the Q-table to select the optimal action sequence, use the Q-table to select the action with the maximum Q value, and obtain the optimal process parameters for output.
2. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 1, characterized in that, Step 2 defines the optimization problem of manufacturing process parameters as maximizing the reward, specifically as follows: Define the optimization objective, environment, state, action, and reward for the manufacturing process parameter optimization problem; where the manufacturing process is the injection molding process of part processing; the optimization objective represents optimizing the process parameters in the part processing manufacturing process; the environment represents the part processing production environment, including the setting range of process parameters and their impact on part quality; the state represents the current setting value of the process parameters; the action represents adjusting the process parameters; and the reward represents the maximum reward obtained in the manufacturing process by measuring the product quality indicators.
3. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 2, characterized in that, The state space is the range of values for the injection rate v, holding time t, and injection pressure P, and is expressed as follows: (S={(v,t,P)|v∈[v_{"{{min}},v_{"{max}}],t∈[t_{"{min}},t_{"{max}}],P∈[P_{"{min}},P_{"{max}}]}) In the formula, {"{min}"} and {"{max}"} represent the minimum value and the maximum value, respectively.
4. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 3, characterized in that, The action space A is represented as follows: (A={(△v,△t,△P)|△v∈[△v_{{min}},△v_{{max}}],△t∈[△t_{{min}},△t_{{max}}],△P∈[ΔP {{min}},△P_{{mas}} ]}).
5. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 4, characterized in that, In the injection molding problem, the reward function R, based on the difference between the actual shrinkage rate and the expected shrinkage rate, is expressed as: (R(s,a)=-|s-s target |) In the formula, s represents the contraction rate of the current state. target Let 'a' represent the expected rate of contraction, and 'a' represent the current action.
6. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 5, characterized in that, In the process problem of injection molding manufacturing, the objective function J is to maximize the cumulative total reward, that is: In the formula, γ represents the discount factor.
7. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 6, characterized in that, The update rule for the Q table is: update the Q value based on the Bellman equation, that is: In the formula, α represents the learning rate, s′ is the next state, and a′ is the action of the next state.
8. The method for optimizing manufacturing process parameters based on reinforcement learning according to claim 7, characterized in that, Step 4 specifically involves: Step 4.1, Initialize the Q-table: Create a Q-table to store the Q-value of each action in each state; the size of the Q-table is the product of the state space and the action space; initially, initialize all Q-values to 0 or a very small random value; Step 4.2, Select Action: Based on the current state (v,t), find the Q value of all corresponding actions in the Q table, and select an action according to the preset strategy; the preset strategy includes a greedy strategy and an ε-greedy strategy; Step 4.3, Execute the action: Represent the current action a with a tuple (Δv, Δt), where Δv represents the change in injection rate and Δt represents the change in holding time; Calculate the state (v+Δv, t+Δt) at the next moment based on the selected action (Δv, Δt); Step 4.4, Update the Q-table: Based on the reward r = -s(v,t) and the transition probability, calculate the reward r obtained at the current time and the maximum Q-value maxQ of the next state (v+Δv,t+Δt); update the Q-value of the action (Δv,Δt) performed in the current state (v,t) according to the Q-table update rules, expressed as: Q(v,t,Δv,Δt)=(1-α)*Q(v,t,Δv,Δt)+α*(r+γ*maxQ) Step 4.5: Repeat steps 2-4 above until the number of iterations reaches the preset value or the Q-table converges.
9. A manufacturing process parameter optimization system based on reinforcement learning, using the manufacturing process parameter optimization method based on reinforcement learning as described in any one of claims 1 to 8, characterized in that, include: The data acquisition module is used to collect process parameter data from the production system. The Markov policy building module is used to build Markov policies for reinforcement learning systems, specifically: Establish a Markov strategy for reasoning and decision-making based on process parameters in the manufacturing process. The Markov strategy represents the state of the product, the action to be taken by the process parameters a, the product transfer state, and the reward function R under the input of the process parameters in the manufacturing process. The relationship between process parameters and product quality is estimated, and the optimization problem of manufacturing process parameters is defined as maximizing reward; the agent in the Markov policy is a machine that performs the task. The reinforcement learning algorithm setting module is used to select the Q-learning reinforcement learning algorithm, specifically: In the Q-learning reinforcement learning algorithm, a Q-table is created, with the dimension of the state space S and the action space A used to store the Q-values. The training module is used for training and learning: it trains the Q-table using a reinforcement learning algorithm; Training process: At time t, the agent selects an action a, obtains the reward and the state for the next time step, and updates the Q-table; Continue iterating until the maximum number of iterations is reached or a converged Q-table is obtained; The process parameter optimization module is used to query the Q-table after training to select the optimal action sequence, use the Q-table to select the action with the maximum Q value, and obtain the optimal process parameters for output.
Citation Information
Patent Citations
Policy estimation device, policy estimation method, and program
WO2022244260A1
Policy estimation device, policy estimation method, and program
WO2022244263A1