Grinding parameter optimization method for numerical control grinding machine based on improved reinforcement learning algorithm
By improving the reinforcement learning algorithm and multi-agent learning mechanism to optimize the grinding parameters of CNC grinding machines, the problem of high energy consumption of CNC grinding machines was solved, and the dual optimization of grinding energy consumption and processing efficiency was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, CNC grinding machines consume a lot of energy during the grinding process, and it is difficult to optimize processing efficiency without changing the surface roughness of the material.
An improved reinforcement learning algorithm is adopted to establish a mathematical model for grinding parameter optimization through multiple linear regression and least squares method. Combined with a multi-agent learning mechanism, the workpiece feed rate, grinding wheel speed and grinding depth are optimized to reduce grinding energy consumption and improve processing efficiency.
Without altering the surface roughness of the material, it significantly reduces the energy consumption of CNC grinding machines and improves the processing efficiency of the grinding process.
Smart Images

Figure CN115600504B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mechanical processing, in particular to a numerical control grinding machine grinding parameter optimization method based on an improved reinforcement learning algorithm. BACKGROUND
[0002] It is well known that as energy is increasingly tight, the traditional manufacturing industry is in urgent need of transformation, and energy-saving improvement is needed for high-energy-consuming equipment. In manufacturing industry, machine tools are one of the main production equipment applied in manufacturing. In the process of machine tool machining, material surface roughness requirement is essential. As a kind of finishing equipment, the installation proportion of numerical control grinding machine in numerical control machine tool reaches 10%. In the whole machining process of machine tool, the grinding stage is an important stage of forming material surface roughness, and the energy consumed and the quality of material in this stage are determined by the combination of grinding parameters. Based on the development of current intelligent algorithm, many scholars optimize the grinding parameters of numerical control grinding machine for different optimization objectives and optimization methods, so as to improve the material surface roughness and grinding efficiency. Reinforcement learning algorithm is a machine learning method in the field of artificial intelligence, which can simplify the learning task to three factors of state, action and reward. Intelligent agent obtains reward by interacting with the environment. Today, reinforcement learning has been applied in various fields, but it is less used in machine tool energy optimization problem. The improved reinforcement learning method is used to solve the energy optimization problem of numerical control grinding machine. SUMMARY
[0003] To solve the above technical problems, the present application provides a numerical control grinding machine grinding parameter optimization method based on an improved reinforcement learning algorithm, which optimizes the selection of processing parameters, effectively reduces the energy consumption of machine tool without changing the material surface roughness, and improves the processing efficiency.
[0004] The technical scheme adopted by the present application is that the numerical control grinding machine grinding parameter optimization method based on the improved reinforcement learning algorithm,
[0005] S1, first, the grinding specific energy formula and the material surface roughness formula are obtained, the coefficients affecting the grinding specific energy and the material roughness are fitted by using the multivariate linear regression and the least square method principle, and the mathematical model of the numerical control grinding machine grinding stage processing parameter optimization is established;
[0006] S2, the grinding parameters are optimized by using the way of mutual learning of multiple intelligent agents, and the specific steps are as follows:
[0007] S2.1, n intelligent agents are set, the iteration number is T, and each intelligent agent is initialized and observed;
[0008] S2.2, after the initialization of each intelligent agent, the intelligent agent will select action according to the observation, and it is set that the intelligent agent has t maxEach time t, the agent will update a step, until t = t max The iteration step updates end at t = t
[0009] S2.3, the agent interacts with the environment according to the behavior at the current time t, and obtains a new state S t+1 And reward R t Reward R t Is set as: R t = -(f j (S t+1 ) - f j (S t )), the agent obtains a new state S t+1 And reward R t From the environment, take new measures; if R t ≤ 0, it means that the current behavior of the agent will not make the result tend to be better, so as to make the agent return to the original state; the agent observes according to the new state S t+1 , so as to obtain a new observation value f j (S t+1 );
[0010] S2.4, the agent updates;
[0011] S2.5, when the agent iteratively updates the state to the maximum iteration number T, the optimal agent of the last iteration is the optimal solution f best , and the corresponding state is the optimal state It is the optimal grinding parameter required.
[0012] As a preferred solution, the step S2.1, the specific process is as follows:
[0013] There are n agents, the iteration number is T, each agent learns in its own environment, and the feed speed v w , the grinding wheel speed v s , the grinding depth a p Is expressed as x i (i = 1, 2, 3), then the current state S of the jth agent can be expressed as:
[0014] S j = (x1, x2, x3) j = (v w , v s , a p ) j
[0015] The formula, j = 1, 2,..., n;
[0016] The value range of x i is [ai , b i ], the three parameters of the jth agent are initialized as:
[0017] x i j = a i + [b i -a i ] *rand(0, 1)
[0018] According to the current state S of the agent, the observation value is set as the function value to be optimized, that is:
[0019] O j = f(S j )
[0020] By initializing the state of each agent and performing observation, the observation value f(S j ) of the optimal agent is recorded as the minimum observation value, and the observation value of the optimal agent is f best .
[0021] As a preferred solution, the step S2.2 has the following specific process:
[0022] At the beginning of the update, the agent first performs observation according to the current state S t at time t, and then plans to take the next action a t according to the current state S t at time t;
[0023] Before the agent takes an action, the action to be taken needs to be evaluated, that is,
[0024]
[0025] The formula, Δf k represents the difference between the observation value obtained after taking the kth action a t,k and the action f(S t );
[0026] After evaluating each action, the action a t is randomly selected using the probability ε, that is:
[0027]
[0028] As a preferred solution, the agent in the current state corresponds to k actions, and since the grinding parameters to be optimized are three variables, it is set that k = 6, and for each action, the following formula is used:
[0029]
[0030] where Δx = (b i-a i )*δ*β j δ is the strengthening factor, β j The basis is the current state S of agent j. t The obtained observation value f j With the optimal intelligent agent f best The value of β is determined by the difference; the smaller the difference, the lower the β. j The smaller.
[0031] As a preferred embodiment, step S2.4 is specifically performed as follows:
[0032] When each agent reaches its maximum step size t max At that time, based on the observation value f of each agent j Sort the data and select the agent with the best observations as the optimal agent f. best The corresponding state is the optimal state S. best Based on the optimal agent, m agents are regenerated, and observations are performed to obtain observation values f. The agent states are updated as follows:
[0033]
[0034] Sort the n+m observations, reselect the optimal agent, delete the m agents that are at the bottom of the list, and then use the updated n agents for the next iteration.
[0035] As a preferred embodiment, step S1 is specifically performed as follows:
[0036] Establish the grinding specific energy model SEG:
[0037]
[0038] Where P represents net grinding power, F t Indicates the tangential grinding force, v s The grinding wheel linear velocity (v) is represented by MRR, which represents the material removal rate. During the grinding process, the grinding wheel linear velocity (v) is... s It is much greater than the workpiece feed rate v. w At this time, the workpiece feed speed v w Negligible. MRR is represented as:
[0039] MRR = v w a p b
[0040] Among them, v w (mm / s) represents the workpiece feed rate; a p (μm) represents the grinding depth; b (mm) represents the wheel width.
[0041] And the change of normal grinding force is related to grinding parameters, and the empirical formula of normal grinding force is expressed as:
[0042]
[0043] Wherein, K1, a1, b1, c1 are coefficients related to machine tool and machining material;
[0044] The grinding specific energy formula is established by combining the above formula:
[0045]
[0046] According to the empirical formula, the material surface roughness is often expressed as:
[0047]
[0048] Similarly, the parameters K2, a2, b2 and c2 in the mathematical expression of the surface roughness Ra can be obtained;
[0049] For the optimization of grinding specific energy and material surface roughness of numerical control grinding machine at the same time, different weight coefficients need to be set to convert multiple objectives into a single objective for solving, and the two are normalized to convert multiple objectives into a single objective function, and the function expression is:
[0050]
[0051] Wherein c is the weight coefficient of grinding specific energy and material surface roughness, c∈[0,1].
[0052] In actual processing, the workpiece feed speed v w , the grinding wheel speed v s , the grinding depth a p are constrained, and the final mathematical model of grinding stage processing parameter optimization is obtained:
[0053]
[0054] The beneficial effects of the present application are:
[0055] The present application optimizes the design, analyzes the machining energy consumption and material surface roughness during the grinding process, and optimizes the grinding parameters, so as to realize the double target optimization of machining energy consumption and machining quality of numerical control grinding machine in the grinding stage; The present application effectively reduces the energy consumption of the machine tool while not changing the material surface roughness, and improves the processing efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Fig. 1 The overall flowchart of the present application is shown in the figure.
[0058] Fig. 2 The flowchart of the algorithm in the present application is shown in the figure. DETAILED DESCRIPTION
[0059] The present application will be described in detail below through exemplary embodiments. However, it should be understood that the elements, structures and features in one embodiment can be beneficially combined into other embodiments without further description.
[0060] Embodiment 1,
[0061] The accompanying drawings will be described below. Figs. 1-2 The optimization process of the present embodiment will be described in detail:
[0062] Since the grinding energy consumption value cannot be directly obtained, the specific grinding energy is commonly used for calculation, which represents the energy consumed for removing unit volume of metal material, and the specific grinding energy model SEG is:
[0063]
[0064] wherein P represents the net grinding power, F t represents the tangential grinding force, v s represents the wheel linear speed, and MRR represents the material removal rate, which represents the size of the material volume removed per second. Since the wheel linear speed v s is much larger than the workpiece feed speed v w in the grinding process, the workpiece feed speed v w can be ignored. For plane grinding, MRR is represented as:
[0065] MRR = v w a p b
[0066] wherein v w (mm / s) is the workpiece feed speed; a p (μm) is the grinding depth; and b (mm) is the wheel width
[0067] The change of the normal grinding force is related to the grinding parameters, and the empirical formula of the normal grinding force is represented as:
[0068]
[0069] wherein K1, a1, b1, c1 are coefficients related to machine tool and machining material
[0070] The grinding specific energy formula is established by combining the above formula:
[0071]
[0072] Since the actual machining will be affected by different factors such as machine tool type, cutting tool and machining material, etc., so when solving the correlation coefficient, different experimental data need to be fitted. The above coefficients are solved by using multiple linear regression.
[0073] Since SEG is a nonlinear equation, the logarithm of both sides is taken:
[0074]
[0075] Let y = lnSEG, β1 = a1-1, β2 = b1+1, β3 = c1-1, x1 = lnv w , x2 = lnv s , x3 = lna p , then it can be converted into a linear equation:
[0076] y = β0 + β1x1 + β2x2 + β3x3
[0077] According to the experimental data, a multiple linear regression equation is established, assuming that there are n groups of data in total
[0078]
[0079] The above formula can be expressed in matrix form as:
[0080] y = Xβ
[0081] wherein
[0082] The least squares method is used to estimate the parameters β, and b0, b1, b2, b3 are the least squares estimates of the parameters β0, β1, β2, β3, respectively, is the regression value of the test result y i , then the regression equation is:
[0083]
[0084] According to the principle of least squares, the sum of the squares of the random errors ε i should be minimized, and ε i is the difference between the test result y i and the regression value The difference. Then the sum of squared random errors of the n groups of trials is represented by Q as:
[0085]
[0086] According to the principle of extrema, take the first derivative of Q with respect to b and set the derivative to zero. The simplified equation is then obtained.
[0087]
[0088] The parameters of grinding specific energy are obtained from the above formula, and thus the grinding specific energy formula is established.
[0089] According to empirical formulas, the surface roughness of a material is often expressed as:
[0090]
[0091] Similarly, the parameters K2, a2, b2, and c2 in the mathematical expression for surface roughness Ra can be obtained.
[0092] To simultaneously optimize the grinding specific energy and material surface roughness of a CNC grinding machine, it is necessary to transform the multi-objective problem into a single objective by setting different weighting coefficients. Since the grinding specific energy and surface roughness have different dimensions, they also need to be normalized. After normalization, the multi-objective problem is transformed into a single-objective function, with the following expression:
[0093]
[0094] Where c is the weighting coefficient between grinding specific energy and material surface roughness, c∈[0,1].
[0095] In actual machining, due to differences in environment, materials, and machine tools, it is necessary to select a reasonable optimization range suitable for the specific machining process. Therefore, before starting optimization, it is necessary to adjust the workpiece feed rate v. w Grinding wheel speed v s Grinding depth a p Constraints were applied. The final mathematical model for energy-saving optimization during the grinding stage is then obtained as follows:
[0096]
[0097] Reinforcement learning (RL) is a type of unsupervised learning where the agent learns through trial and error. The learning task in reinforcement learning can be described as an agent interacting with its environment. The environment can be considered a system, and the agent learns by observing the current state s. t Based on each observation o t Generate action a t , with action a tA new state s is generated in the environment t+1 With the reward r t , the degree of good or bad of the action made by the agent is represented. t
[0098] The present application improves the reinforcement learning for numerical control grinding, adopts a multi-agent learning mode to solve the problem of processing parameter selection in energy consumption optimization of numerical control grinding machine, and the specific method is as follows:
[0099] 1) Initialize the state of the agent and observe
[0100] n agents are set, the iteration number is T, each agent learns in the respective environment, and the feed speed v w , the grinding wheel speed v s , the grinding depth a p is represented by x i (i=1, 2, 3), then the current state S of the jth agent can be represented as:
[0101] S j =(x1,x2,x3) j =(v w ,v s ,a p ) j
[0102] The formula, j=1, 2,..., n.
[0103] The value range of x i is [a i , b i ], then the three parameters of the jth agent are initialized as:
[0104] x i j =a i +[b i -a i ]*rand(0,1)
[0105] According to the current state S of the agent, the observation value is set as the function value to be optimized, that is:
[0106] O j =f(S j )
[0107] By initializing the state of each agent and observing, the observation value f(S j ) of the optimal agent is recorded as the minimum, and the observation value is f best ;
[0108] 2) Behavior evaluation and execution
[0109] After initializing each agent, the agent will select actions according to the observation. It is set that the agent has t max steps in each iteration, and the agent will update the step once at each time t, until t = t max , which indicates that the step update is over at this iteration;
[0110] When the update starts, the agent first observes the state S t at the current time t, and then plans to take the next action a t according to the state S t at the current time. It is set that the agent has k actions corresponding to the current state, because the optimization of the grinding parameters is three variables, so it is set that k = 6, and for each action, the following formula is shown:
[0111]
[0112] Where Δx = (b i -a i )*δ*β j , δ is the reinforcement factor, and β j is determined according to the difference between the observation value f t obtained by the current agent j in the state S j and the optimal agent f best , the smaller the difference, the smaller the β j ;
[0113] Before the agent takes the action, the action to be taken needs to be evaluated, that is,
[0114]
[0115] The formula is Δf k , which represents the difference between the observation value obtained after taking the kth action a t,k and the action f(S t );
[0116] After evaluating each action, the action a t is randomly selected using the probability ε, that is:
[0117]
[0118] 3) State update
[0119] The agent interacts with the environment according to the action at the current time t, thereby obtaining a new state S t+1 and a reward R t , where the reward R t is set as:
[0120] Rt = -(f j (S t+1 )-f j (S t ))
[0121] The agent gets a new state S t+1 and a reward R t , takes a new action, and if R t ≤ 0, it means that the current agent's action does not tend to make the result better, so the agent returns to the original state,
[0122] The agent observes the new state S t+1 and gets a new observation f j (S t+1 );
[0123] 4) Agent update
[0124] When each agent reaches the maximum step t max , sort according to the observation value f j of each agent, and select the best observation value as the optimal agent f best , and the corresponding state is the optimal state S best , and generate m agents according to the optimal agent and get the observation value f. The agent updates the state as follows:
[0125]
[0126] Sort according to the observation value of n+m, reselect the optimal agent, and delete the last m agents, then update the n agents for the next iteration;
[0127] 5) Optimal solution selection
[0128] When the agent iteratively updates the state to the maximum iteration T, the optimal agent of the last iteration is the optimal solution f best , and the corresponding state is the optimal state , which is the optimal grinding parameter required.
[0129] Specific implementation method:
[0130] First, calculate the grinding specific energy formula SEG and the material surface roughness formula Ra according to the reference manual and formula derivation;
[0131] Use multiple linear regression and least squares method to fit the coefficients affecting grinding specific energy and material roughness, and establish a mathematical model for optimizing the machining parameters of the grinding stage of the numerical control grinding machine;
[0132] An improved reinforcement learning algorithm for grinding parameter optimization is used to optimize the grinding parameters:
[0133] According to 1) n initial solutions are randomly generated in the domain, and the solution is updated through 2) ~ 3);
[0134] Determine whether the maximum step length is reached. If not, return to 2); if yes, go to 4);
[0135] Determine whether the maximum iteration number is reached. If not, return to 2); if yes, output the optimal agent state S best , which contains three variables x i , the optimal workpiece feed speed v w , the wheel speed v s , and the grinding depth a p , and further obtain the grinding specific energy SEG and the surface roughness Ra of the processed material in the optimal grinding stage.
[0136] The multiple linear regression and least squares method in the above method are used to fit the grinding parameter optimization formula, and a multi-objective mathematical model of grinding parameter optimization of numerical control grinding machine is established, and the results are as follows:
[0137]
[0138]
[0139]
[0140]
[0141] The improved reinforcement learning algorithm is used to optimize the mathematical model, and the coefficient c is taken from 0.1 to 0.9 for calculation, and the optimal grinding parameters under different values are output.
[0142] To further verify the performance of the algorithm, the results obtained by the algorithm are compared with those obtained by other algorithms.
[0143] Comparative Example 1:
[0144] The comparative literature used in this comparative example is: Tian Xiao. Experimental study on energy consumption and surface quality in ultra-fine hard alloy grinding process [D]. Fujian University of Technology, 2020;
[0145] The results obtained by the orthogonal optimization method used by Tian Xiao show that the optimal grinding parameters in the original text are: grinding parameters v w = 48 mm / s, v s = 30 m / s, a p = 5 μm, and the result obtained is SEG = 70.13 J / mm 3, Ra = 0.6329 pm. The result obtained by using the improved reinforcement learning algorithm is: v w = 29.48 mm / s, v s = 30 m / s, a p = 20 pm, the obtained result SEG = 30.0024 J / mm 3 , Ra = 0.6324 pm. The result SEG is reduced by 57.21% and Ra is reduced by 0.079% compared with the above result. Although the optimization result of surface roughness is not significant, the grinding specific energy is greatly reduced, which means that the improved reinforcement learning algorithm can further reduce the grinding specific energy while achieving the same processing quality.
[0146] Comparative Example 2:
[0147] The comparative document used in this comparative example is: Zhang Y, Li B, Yang J, et al. Modeling and optimization of alloy steel 20CrMnTi grinding process parameters based on experiment investigation [J]. International Journal of Advanced Manufacturing Technology, 2018, 95(5-8): 1859-1873;
[0148] According to the results obtained by Zhang et al. using the Pareto optimization algorithm, the optimal grinding parameter results in the original text are: v w = 0.836 m / s, v s = 120 m / s, a p = 13.5 pm, the obtained results are material surface roughness Ra = 0.686 pm, and material removal rate Q w = 11.286 mm 2 / s. The result obtained by using the improved reinforcement algorithm is: v w = 0.72 m / s, v s = 150 m / s, a p = 19 pm, the obtained results are roughness Ra = 0.685 pm, and material removal rate Q w = 13.68 mm 2 / s. Compared with the above results, Ra is reduced by 0.15%, and Q w is increased by 21.21%. It can be seen that: while the surface roughness has a certain optimization effect, the material removal rate is greatly improved.
[0149] It should be noted that, although the application has been described by means of the above embodiments, the application can also be implemented in other ways. Those skilled in the art will obviously make various corresponding changes and modifications to the application without departing from the spirit and scope of the application, and these changes and modifications should all belong to the scope of protection of the claims of the application and their equivalents.
Claims
1. A numerical control grinding machine grinding parameter optimization method based on an improved reinforcement learning algorithm, characterized by: S1, first obtain the grinding specific energy formula and the material surface roughness formula, use multiple linear regression and least square method to fit the coefficients affecting the grinding specific energy and material roughness, and establish a mathematical model for optimizing the machining parameters in the grinding stage of the numerical control grinding machine; S2, use the way of multiple agents learning from each other to optimize the grinding parameters, the specific steps are as follows: S2.1, set n agents, the iteration number is T, initialize the state of each agent and observe; S2.2, after initializing each agent, the agent will select action according to observation, set the agent in each iteration a total of step, at each time t, the agent will update a step, until time, the iteration step update is over; S2.3, the agent interacts with the environment according to the current behavior at time t, and obtains a new state and the reward , the reward is set as: , the agent obtains a new state from the environment and the reward , a new measure is taken; if , it indicates that the current behavior of the agent will not make the result tend to be better, so the agent is returned to the original state; the agent observes according to the new state , so as to obtain a new observation value ; S2.4, agent update; S2.5, when the intelligent agent iteratively updates the state to reach the maximum number of iterations T, the optimal intelligent agent of the last iteration is the optimal solution , and the corresponding state is the optimal state is the optimal grinding parameter required The step S1, the specific process is as follows: Establish the grinding specific energy model SEG: wherein, represents net grinding power, represents tangential grinding force, represents wheel linear speed, represents material removal rate, since in the grinding process, the wheel linear speed is much greater than the workpiece feed speed , at which time the workpiece feed speed is negligible; MRR is represented as: wherein, is the workpiece feed speed; is the grinding depth; b is the wheel width; The change of normal grinding force is related to grinding parameters, and the empirical formula of normal grinding force is expressed as: K1, 1, b1, c1 are coefficients related to the machine tool and the machining material; The grinding specific energy formula is established by combining the above formula: ; According to the empirical formula, the material surface roughness is often expressed as: The same applies to the parameters in the mathematical expression of the surface roughness , , , , ; For the optimization of grinding specific energy and material surface roughness of numerical control grinding machine at the same time, different weight coefficients need to be set to convert multiple objectives into single objective for solving, and the two are normalized to convert multiple objectives into single objective function, and the function expression is: where c is a weight coefficient of the grinding specific energy and the surface roughness of the material, ; In actual processing, the workpiece feed speed , the grinding wheel speed , the grinding depth are constrained, and the grinding stage processing parameter optimization final mathematical model is obtained as follows: 。 2. The improved reinforcement learning algorithm-based grinding parameter optimization method for a CNC grinding machine according to claim 1, characterized in that: The step S2.1, the specific process is as follows: There are n agents, and the number of iterations is T. Each agent learns in its own environment and adjusts the feed rate. Grinding wheel speed Grinding depth by If we express this as follows, then the current state S of the j-th agent can be represented as: Equation, ; corresponding the value range of then the three parameters of the jth agent are initialized as: According to the current state S of the agent, set the function value to be optimized as the observation value, that is: By initializing the state of each agent and making observations, let the observation value of the agent with the minimum observation value be the optimal agent recorded as the optimal agent, and the observation value of the optimal agent is .
3. The improved reinforcement learning algorithm-based grinding parameter optimization method for a CNC grinding machine according to claim 1, characterized in that: The step S2.2, the specific process is as follows: At the beginning of the update, the agent first observes the state at the current time t performs an observation, and then plans the next action based on the state at the current time ; The agent needs to evaluate the behavior to be taken before taking the behavior, that is Equation, represents taking the kth action the difference between the observation obtained after taking the action the kth action; After valuing each action, use probability Randomly select action That is: 。 4. The improved reinforcement learning algorithm-based grinding parameter optimization method for a CNC grinding machine according to claim 3, characterized in that: The agent is set to correspond to k behaviors in the current state, and the optimization of the grinding parameters is three variables. Here, set For each behavior, the following formula is shown: wherein , is a reinforcement factor, According to is according to the current state of the agent j State The resulting observation The difference from the optimal agent is determined, the smaller the difference the smaller.
5. The CNC grinding machine grinding parameter optimization method based on improved reinforcement learning algorithm according to claim 1, characterized in that: The step S2.4, the specific process of agent update is as follows: When each agent reaches the maximum step length , the observation value of each agent is sorted according to the observation value of each agent , and the best observation value is selected as the optimal agent , and the corresponding state is the optimal state ; according to the optimal agent, m agents are regenerated, and the observation value f is obtained by observation, and the state of the agent is updated as follows: According to the observation value of n+m, sort and select the optimal agent, delete the last m agents, and then update the n agents for the next iteration.
Citation Information
Patent Citations
Grinding quality estimation model generating device, grinding quality estimating device, poor quality factor estimating device, grinding machine operation command data adjustment model generating device, and grinding machine operation command data updating device
US20200033842A1
Multi-agent deep reinforcement learning proxy method based on intelligent grid
WO2020000399A1