Ball screw feeding system servo parameter optimization method and device based on flexible action-evaluation algorithm

By introducing the SAC algorithm and maximum entropy concept based on the Actor-Critic framework, and combining the simulation model to train the agent network, the problem of bionic algorithms being easily trapped in local optimality is solved, and the efficient optimization of the servo parameters of the ball screw feed system is realized, which improves the system's motion accuracy and response speed.

CN119987285AActive Publication Date: 2025-05-13HUAZHONG UNIV OF SCI & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510092964.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The existing bionic algorithms are prone to local optimization in the optimization of servo parameters of ball screw feeding systems, and have high calculation costs, making it difficult to adapt to complex environments and uncertainties, resulting in low optimization efficiency.

Method used

The flexible action-evaluation algorithm (SAC algorithm) based on the Actor-Critic framework is adopted, and the maximum entropy concept and playback buffer are introduced, and the agent network is trained in combination with the simulation model. The servo parameters are optimized through soft update strategies to avoid local optimization and improve training stability.

Benefits of technology

It improves the exploration and stability of servo parameter optimization, reduces calculation costs, shortens training time, enhances the algorithm's anti-interference ability and cross-device migration ability, and improves the motion accuracy and response speed of the ball screw system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987285A_ABST
    Figure CN119987285A_ABST
Patent Text Reader

Abstract

The invention belongs to the related technical field of servo control optimization, and discloses a ball screw feeding system servo parameter optimization method and device based on a flexible action-evaluation algorithm, and the method comprises the following steps: (1) constructing an intelligent agent network of an SAC algorithm based on an Actor-Critic framework, and introducing the maximum entropy into the intelligent agent network; and (2) associating the simulation model of the ball screw feeding system with the intelligent agent network of the SAC algorithm through an environment search interface, training the intelligent agent network of the SAC algorithm, and then obtaining optimized servo parameters by adopting the trained intelligent agent network. The maximum entropy concept is introduced into the used SAC algorithm, the maximum entropy of the random strategy is explored while the maximum accumulated reward value is explored, the exploration space of the intelligent agent network is widened, the exploration randomness is improved, and compared with a heuristic algorithm and other deep reinforcement learning algorithms, local optimum is not prone to occurring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to servo control optimization, and more specifically, relates to a servo parameter optimization method and device for a ball screw feed system based on a flexible motion-evaluation algorithm. Background Art

[0002] As a common drive system for general CNC machine tools, the ball screw feed system plays a vital role in ensuring and improving the processing accuracy, work efficiency and comprehensive performance of the entire CNC machine tool. The dynamic performance optimization of the ball screw system is a hot topic of research. By designing servo control strategies and control algorithms, the dynamic performance of the feed drive system can be estimated and optimized, thereby effectively improving the manufacturing accuracy. The performance of the control algorithm is largely affected by the key servo parameters. Only by selecting appropriate parameter values ​​can the improvement of the dynamic performance of the feed drive system be guaranteed. Therefore, it is of great significance to study the motion control parameter optimization technology of the ball screw feed system.

[0003] According to the different principles of servo parameter optimization methods, they can generally be divided into parameter optimization methods based on control models and intelligent parameter optimization methods. The parameter optimization method based on the control model first needs to determine the system model of the control object. Specific methods include logarithmic frequency characteristic method, root locus method, attenuation curve method, etc. Then, based on the transfer function or state space equation of the control object, the theoretical value of the servo parameter can be calculated according to the control principle. This parameter optimization method is suitable for scenarios with small interference in the working environment and simple control objects, but it usually requires manual fine-tuning after obtaining the theoretical value. Therefore, in practical applications, there may be problems of long time consumption and low efficiency, which cannot meet the requirements of high precision and high efficiency in industrial production. With the in-depth study of optimization problems by scholars, a variety of intelligent parameter optimization methods have been proposed, among which heuristic parameter optimization algorithms have been widely used in industrial production. At present, heuristic algorithms are mainly based on natural body algorithms, such as particle swarm optimization algorithm and genetic optimization algorithm. The particle swarm optimization algorithm has few hyperparameters and fast convergence, and is suitable for continuous problems; the genetic optimization algorithm has a large amount of calculation and is suitable for continuous or discrete problems; and the Bayesian optimization algorithm is suitable for medium and low dimensional optimization problems. The goal of the heuristic parameter optimization algorithm is to find the global optimal solution of the objective function, which can effectively adapt to the control optimization problems of complex systems. However, this type of algorithm has high computational cost, strong dependence of algorithm performance on hyperparameters, and is prone to fall into the dilemma of local optimality during the parameter search process.

[0004] Recently, deep learning and reinforcement learning have become research hotspots for scholars. Introducing deep reinforcement learning algorithms into the optimization of control system servo parameters can help improve the intelligence of the control system tuning process. At the same time, the intelligent agent constructed by this type of algorithm can realize the migration of optimization strategies across devices and working conditions, effectively reducing the time for repeated training. Among common deep reinforcement learning models, the Actor-Critic algorithm has fast parameter adjustment speed and strong anti-interference ability. Compared with heuristic algorithms, it can adapt to more complex environments and uncertainties, and has stronger generalization ability. However, the Actor-Critic algorithm training process is computationally intensive, the network converges slowly, and the optimization results are prone to fall into local optimality. Summary of the invention

[0005] In view of the above defects or improvement needs of the prior art, the present invention provides a servo parameter optimization method and device for a ball screw feed system based on a flexible motion-evaluation algorithm, which aims to solve the problem that the existing bionics are prone to fall into local optimality.

[0006] To achieve the above object, according to one aspect of the present invention, a method for optimizing servo parameters of a ball screw feed system based on a flexible motion-evaluation algorithm is provided, the method comprising the following steps:

[0007] (1) An agent network of the SAC algorithm based on the Actor-Critic framework is constructed. The agent network introduces maximum entropy, and the formula of its Actor strategy network is:

[0008]

[0009] Among them, R(s t ,a t ) is the reward value, H(π(·|s t )) is the entropy value, γ represents the discount factor, and α is the temperature coefficient;

[0010] (2) The simulation model of the ball screw feed system is associated with the intelligent agent network of the SAC algorithm through the environment search interface, and the intelligent agent network of the SAC algorithm is trained. Then, the trained intelligent agent network is used to obtain the optimized servo parameters.

[0011] Furthermore, the SAC algorithm agent network simultaneously learns two Critic networks and one Actor policy network. Each Critic network also has its corresponding target network. The soft update strategy is used to update the parameters of the target network during training.

[0012] Furthermore, the state value function is:

[0013]

[0014] The action value function of the Critic network is:

[0015]

[0016] In the formula, s t 、a t and r t are the system state value, output action and reward value at the current moment, i.e., time t; γ is the discount factor, which is used to control the weight of future rewards; R(s t ,a t ) is state s t Take action a t The reward value of 0 is the initial state; a 0 is the initial action; α is the temperature coefficient.

[0017] Furthermore, the search environment of the SAC algorithm agent network includes an agent-environment interaction function and a search environment initialization function. During the agent training process, the done value is used to determine whether the current training round is completed, that is, whether the target convergence value is reached. The default value of done is 0.

[0018] The agent-environment interaction function calculates the current action a based on the objective function of the reward value t The reward value r t , compare the reward value with the target value goal to get the done value, the calculation formula is:

[0019]

[0020] The agent performs action a t After the state is updated, s t+1 , the state update function is:

[0021] s t+1 ~P(s t+1 ∣s t ,a t )

[0022] The search environment initialization function is:

[0023] s t =s 0

[0024] In the formula, s 0 is the initial state.

[0025] Furthermore, the agent network needs to initialize the environment for each round of search. In each round of search, the agent network is used to calculate the reward value of the current action and the state of the next step, and to determine whether the current round has reached the target convergence value.

[0026] Furthermore, the training parameters of the intelligent agent network include the number of training rounds, learning rate and maximum entropy value; the preset servo parameter values ​​are used as input to train the intelligent agent network, the next action is selected according to the current state, and the reward value is calculated after each round of search is completed, and the training is stopped after the set number of rounds is completed.

[0027] Furthermore, the update formula of the critic network is:

[0028]

[0029] The update formula of the Actor network is:

[0030]

[0031] The update formula of α value is:

[0032]

[0033] In the formula, Q θ (s t ,a t ) is the Critic network in state s t Take action a t The expected cumulative reward; D is the data in the replay buffer; r(s t ,a t ) is state s t Take action a t The reward value of For the Actor network in state s t+1 The expected cumulative reward under φ is the current strategy sample; f φ (ε t ;s t ) uses the reparameterization technique, and the actions are sampled from the policy Gaussian distribution, that is, a t =f φ (ε t ;s t ), where f φ is a policy neural network with parameters φ, ε t is the initialization parameter of the network; Q θ (s t ,f φ (ε t ;s t )) is the target Critic network state s t Next take action f φ (ε t ;s t )’s expected cumulative reward; π t For state s tTake action a t probability; is the target value of entropy, which is used to automatically adjust the entropy regularization term.

[0034] Furthermore, the objective function reward of the reward value is:

[0035] reward=-(tr+overshot+ts+td).

[0036] The present invention also provides a servo parameter optimization system for a ball screw feed system based on a flexible motion-evaluation algorithm, the system comprising a memory and a processor, the memory storing a computer program, and the processor executing the servo parameter optimization method for a ball screw feed system based on a flexible motion-evaluation algorithm as described above when executing the computer program.

[0037] The present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the servo parameter optimization method of the ball screw feed system based on the flexible motion-evaluation algorithm as described above.

[0038] In general, compared with the prior art, the above technical solution conceived by the present invention, the servo parameter optimization method and device of the ball screw feed system based on the flexible action-evaluation algorithm provided by the present invention mainly have the following beneficial effects:

[0039] 1. The SAC algorithm used in the present invention introduces the concept of maximum entropy. While exploring the maximum cumulative reward value, it also explores the maximum entropy of random strategies, broadens the exploration space of the agent network, improves the randomness of the exploration, and is less likely to fall into local optimality compared to heuristic algorithms and other deep reinforcement learning algorithms.

[0040] 2. The SAC algorithm used in the present invention introduces a playback buffer and a target network. The experience playback area stores historical samples and randomly extracts small batches of samples for training and updating, which reduces the correlation of samples and improves the utilization rate of samples. The target network and the estimated network have the same structure, and their parameters are regularly copied from the estimated network, making the training process of the intelligent agent more stable and avoiding the oscillation problem caused by the rapid update of the estimated network parameters.

[0041] 3. The SAC algorithm used in the present invention is generalizable, and the trained intelligent agent can be migrated to different devices and different working conditions for servo parameter tuning, which reduces the model repetitive training time, can effectively help improve the motion accuracy and response speed of the ball screw system, and has good industrial application value.

[0042] 4. The SAC algorithm uses two critic networks to alleviate the problem of overestimation of Q values. The structure of the algorithm's target network is the same as that of the critic network, which is used to stabilize the Q value update. The target network maintains the stability of training through soft updates (i.e., only some parameters are updated after each training).

[0043] 5. Considering that the interaction time required to train the intelligent agent using a real ball-wire feed system is long and there are many interferences, a simulation model is chosen to improve the interaction efficiency. By analyzing the system's motion control process and obtaining the system's attribute parameters, a simulation model is built to simulate the real ball-wire feed system, thereby shortening the interaction time between the environment and the intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flow chart of a servo parameter optimization method of a ball screw feed system based on a flexible action-evaluation algorithm provided by the present invention;

[0045] Figure 2 It is a graph showing the change in the reward function value during the agent training process of the SAC algorithm of the present invention;

[0046] Figure 3 It is a graph showing the change in the average reward function value during the agent training process of the SAC algorithm in the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0048] The present invention provides a servo parameter optimization method for a ball screw feed system based on a flexible action-evaluation algorithm, and the method mainly comprises the following steps:

[0049] Step 1: construct an agent network of the SAC algorithm based on the Actor-Critic framework. The agent network introduces maximum entropy, and the formula of its Actor strategy network is:

[0050]

[0051] Among them, R(s t ,a t ) is the reward value, H(π(·|s t )) is the entropy value, γ represents the discount factor, and α is the temperature coefficient.

[0052] Specifically, the SAC algorithm is an improved algorithm based on the Actor-Critic framework. The SAC algorithm simultaneously learns two Critic networks and one Actor policy network. Each Critic network also has its corresponding target network. The parameters of the target network are updated using a soft update strategy during training. At the same time, the SAC algorithm introduces entropy into the reinforcement learning algorithm to prevent the strategy from falling into the local optimal point, and explores multiple feasible solutions to complete the specified task, thereby improving the algorithm's anti-interference ability.

[0053] The goal of a deep reinforcement learning algorithm is to learn a strategy with the highest expected cumulative reward value. In order to randomize the strategy, that is, to make the probability of each action output as dispersed as possible, rather than falling into a local optimum, the SAC algorithm introduces the concept of maximum entropy, requiring that the entropy of each action output by the strategy is the largest. The formula for the Actor strategy network is:

[0054]

[0055] Among them, R(s t ,a t ) is the reward value, H(π(·|s t )) is the entropy value, γ represents the discount factor, and α is the temperature coefficient.

[0056] The state value function is changed to:

[0057]

[0058] The action value function of the Critic network is changed to:

[0059]

[0060] The SAC algorithm uses two critic networks to alleviate the problem of overestimation of Q values. The structure of the algorithm's target network is the same as that of the critic network, which is used to stabilize the Q value update. The target network maintains the stability of training through soft updates (i.e., only some parameters are updated after each training).

[0061] Due to the different goals of industrial production, the performance requirements of the ball screw feed system are also different, so it is necessary to set an objective function with different reward values ​​for different performance requirements. The objective function can be set in a single indicator, a weighted sum of multiple indicators, or a weighted sum of multiple indicators with constraints.

[0062] Build the search environment of the SAC algorithm agent network, including the agent-environment interaction function and the search environment initialization function. During the agent training process, the done value is used to determine whether the current training round is completed, that is, whether the target convergence value is reached. The default value of done is 0.

[0063] The agent-environment interaction function calculates the current action a based on the objective function of the reward value t The reward value r t , compare the reward value with the target value goal to get the done value, the calculation formula is:

[0064]

[0065] The agent performs action a t After the state is updated, s t+1 , the state update function is:

[0066] s t+1 ~P(s t+1 ∣s t ,a t )

[0067] The search environment initialization function is:

[0068] s t =s 0

[0069] The agent network needs to initialize the environment for each round of search. In each round of search, the agent network is used to calculate the reward value of the current step action and the state of the next step, and to determine whether the current round has reached the target convergence value.

[0070] Step 2: Associating the simulation model of the ball screw feed system with the intelligent agent network of the SAC algorithm through the environment search interface, training the intelligent agent network of the SAC algorithm, and then using the trained intelligent agent network to obtain the optimized servo parameters.

[0071] Among them, considering that the interaction time required for training the intelligent agent using a real ball-bearing wire feed system is long and there are many interferences, we chose to use a simulation model to improve the interaction efficiency. By analyzing the system's motion control process and obtaining the system's attribute parameters, a simulation model was built to simulate the real ball-bearing wire feed system, thereby shortening the interaction time between the environment and the intelligent agent.

[0072] Set the training parameters of the agent network, including the number of training rounds, learning rate, and maximum entropy value, etc., use the preset servo parameter values ​​as input to train the agent network, select the next action based on the current state, and calculate the reward value after each round of search. Stop training after completing the set number of rounds.

[0073] The agent training process of the SAC algorithm is as follows: (1) The Actor network is trained based on the current state s t Select action a t , and sends it to the environment to execute the action; (2) the environment executes the action and obtains the reward r t and the new state st+1 ; (3) The Actor network transfers the state (the current step state s t 、Action a t and reward r t , the next state s t+1 ) is stored in the playback buffer as the data set for training the estimation network; (4) N state transition process data are randomly sampled from the playback buffer as a mini-batch training data set for the estimation network and the target network; (5) The average value of the two critic target networks is calculated and The minimum value of; (6) Update the two Critic networks and (7) Update the Actor network (8) Update the value of α. Repeat the above process until convergence.

[0074] Among them, the update formula of the critic network is:

[0075]

[0076] The update formula of the Actor network is:

[0077]

[0078] The update formula of α value is:

[0079]

[0080] The present invention is further described in detail below with reference to specific embodiments.

[0081] This embodiment provides a method for optimizing servo parameters of a ball screw feed system based on a flexible motion-evaluation algorithm. The specific implementation steps are as follows:

[0082] Step 1: Build a simulation model of the ball wire feed system.

[0083] By using Python to call the GUI of the FMPy library, setting reasonable ball screw structural parameters, and adopting the basic PID control method, a simple ball screw motion speed control system was constructed. The input of the simulation model is the servo control parameters KP, KI, and KD of the equipment, that is, the parameters that need to be tuned. The returned system performance evaluation indicators include the rise time tr, overshoot overshot, adjustment time ts, and setting time td of the system speed response curve. According to the relevant knowledge of control theory, it can be known that the rise time tr is antagonistic to the other three indicators, and when the four outputs are small, the speed following performance of the system is better. At the same time, the simulation model provides a set of self-tuning optimal parameters (also initial parameters) and corresponding evaluation indicators to test the performance of the SAC parameter optimization algorithm.

[0084] Step 2: Build a SAC algorithm agent network based on the Actor-Critic framework.

[0085] The SAC algorithm consists of five networks. The agent network simultaneously learns two Critic networks and one Actor policy network. Each Critic network also has its corresponding target network. During the training process, the soft update strategy is used to update the parameters of the target network. The training flow chart of the SAC algorithm is as follows: Figure 1 shown.

[0086] Figure 1 Medium t 、a t and r t are the system state value, output action and reward value at the current moment, i.e., time t; s t+1 、a t+1 are the next step, i.e., the system state value and output action at time (t+1); batch_size is the size of samples extracted from the experience pool, i.e., the playback buffer; is the strategy network, i.e., the Actor network; the two Critic networks are and The target networks of the two Critic networks are and

[0087] The formula for the Actor policy network is:

[0088]

[0089] Among them, R(s t ,a t ) is the reward value, H(π(·|s t )) is the entropy value, γ represents the discount factor, and α is the temperature coefficient.

[0090] The state value function is changed to:

[0091]

[0092] The action value function of the Critic network is changed to:

[0093]

[0094] The SAC algorithm uses two critic networks to alleviate the problem of overestimation of Q values. The structure of the target network is the same as that of the critic network, which is used to stabilize the Q value update. The target network maintains the stability of training through soft updates.

[0095] Step 3: Based on the performance requirements of the ball screw feed system, set the objective function of the SAC algorithm reward value.

[0096] The objective function of the reward value is used to evaluate the quality of the agent's action selection in the current round, so as to train the agent to obtain the optimal strategy network. In this embodiment, the objective function reward is:

[0097] reward=-(tr+overshot+ts+td)

[0098] Step 4: Build the search environment for the SAC algorithm agent network.

[0099] The search environment of the SAC algorithm agent includes the agent state update function, the agent action selection function and the search environment initialization function.

[0100] During the agent training process, the done value is used to determine whether the current training round is completed, that is, whether the target convergence value is reached. The default value of done is 0.

[0101] The agent-environment interaction function calculates the current action a based on the objective function of the reward value t The reward value r t , compare the reward value with the target value goal to get the done value, the calculation formula is:

[0102]

[0103] The agent performs action a t After the state is updated, s t+1 , the state update function in this embodiment is:

[0104] s t+1 =a t

[0105] The search environment initialization function is:

[0106] s t =s0

[0107] In the formula, s 0 is the initial state, that is, the initial parameters.

[0108] The agent needs to initialize the environment for each round of search. In each round of search, the agent network is used to calculate the next state and reward value, and to determine whether the current round has reached the target convergence value. In addition, the range of PID parameter optimization is set to ±10% of the initial PID parameter.

[0109] Step 5: Train the agent network of the SAC algorithm.

[0110] The simulation model and the agent are associated through the environment search interface, and the training parameters of the agent are set, including the number of training rounds, learning rate, and maximum entropy value. The preset initial servo parameter values ​​are used as input to train the agent of the SAC algorithm. The training steps are as follows: (1) The Actor network is trained according to the current step state s t Select an action t , issued to the environment to perform the action. In this embodiment, the action is selected according to the state of the ball screw motion speed control system at the current step, that is, the PID parameters of the current step are selected according to the PID parameters of the previous step and input into the ball screw motion speed control system; (2) The environment performs the action to obtain the reward r t and the new state s t+1 , that is, the ball screw motion speed control system generates the corresponding speed response curve according to the PID parameters input in the current step, calculates the reward value of the current step according to the response curve, and uses the PID parameters of the current step as the state of the next step; (3) The Actor network transfers this state process (the current step state s t 、Action a t and reward r t , next state s t+1 ) is stored in the playback buffer as the data set for training the estimation network; (4) 64 state transition process data samples are randomly extracted from the playback buffer as a mini-batch training data set for the estimation network and the target network; (5) The average value of the two critic target networks is calculated and The minimum value of; (6) Update the two Critic networks and (7) Update the Actor network (8) Update the value of α. Repeat the above process until convergence.

[0111] Among them, the update formula of the Critic network is:

[0112]

[0113] The update formula of the Actor network is:

[0114]

[0115] The update formula of α value is:

[0116]

[0117] The reward value and average reward value of the agent training process are as follows: Figure 2 and Figure 3 shown.

[0118] Step 6: Test the trained SAC algorithm agent network, obtain the optimized servo parameters, and verify the effectiveness of the algorithm.

[0119] After training the SAC algorithm agent network, a PID parameter tuning test was carried out on the same ball screw motion speed control system simulation model to compare and analyze the changes in performance indicators before and after the control system optimization, thereby verifying the effectiveness of the algorithm.

[0120] The test results are shown in Table 1. It can be seen that after the servo parameters of the system are optimized by the SAC algorithm, although the rise time tr is slightly increased, the overshoot overshot, adjustment time ts and setting time td are significantly reduced, and the overall performance of the system is significantly improved.

[0121] Table 1 Comparison of parameter tuning results of the SAC algorithm agent on the training device

[0122] KP KI KD tr overshot ts td Initial parameters / indicators 32000 500 5e-05 0.068 0.055 0.103 0.104 Optimized parameters / indicators 30484.99 513.84 5.19e-05 0.07 0.021 0.071 0.092 Optimization percentage — — — -2.94% 61.82% 31.07% 11.54%

[0123] In addition, in order to verify that the intelligent agent network trained by the SAC algorithm has good migration and application capabilities, the structural parameters of the ball screw in the simulation model are modified, and other conditions are kept unchanged. The trained intelligent agent network is used to directly perform cross-device PID parameter tuning tests. The test results are shown in Table 2. The comparative analysis shows that the servo parameters optimized by the intelligent agent have a certain effect on improving the system performance of the new device, successfully verifying the possibility of cross-device application of the intelligent agent.

[0124] Table 2 Comparison of cross-device parameter tuning results of the SAC algorithm

[0125] KP KI KD tr overshot ts td Initial parameters / indicators 32000 500 5e-05 0.081 0.178 0.161 0.162 Optimized parameters / indicators 29502.03 524.57 5.01e-05 0.082 0.14 0.155 0.157 Optimization percentage — — — -1.23% 21.35% 3.73% 3.09%

[0126] The present invention also provides a servo parameter optimization system for a ball screw feed system based on a flexible motion-evaluation algorithm, the system comprising a memory and a processor, the memory storing a computer program, and the processor executing the servo parameter optimization method for a ball screw feed system based on a flexible motion-evaluation algorithm as described above when executing the computer program.

[0127] The present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the servo parameter optimization method of the ball screw feed system based on the flexible motion-evaluation algorithm as described above.

[0128] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A servo parameter optimization method for a ball screw feed system based on a flexible motion-evaluation algorithm, characterized in that: The method comprises the following steps: (1) An agent network of the SAC algorithm based on the Actor-Critic framework is constructed. The agent network introduces maximum entropy, and the formula of its Actor strategy network is: Among them, R(s t ,a t ) is the reward value, H(π(·|s t )) is the entropy value, γ represents the discount factor, and α is the temperature coefficient; (2) The simulation model of the ball screw feed system is associated with the intelligent agent network of the SAC algorithm through the environment search interface, and the intelligent agent network of the SAC algorithm is trained. Then, the trained intelligent agent network is used to obtain the optimized servo parameters.

2. The servo parameter optimization method of a ball screw feed system based on a flexible motion-evaluation algorithm according to claim 1, characterized in that: The SAC algorithm agent network simultaneously learns two Critic networks and one Actor policy network. Each Critic network also has its corresponding target network. During the training process, a soft update strategy is used to update the parameters of the target network.

3. The servo parameter optimization method of a ball screw feed system based on a flexible motion-evaluation algorithm according to claim 2, characterized in that: The state value function is: The action value function of the Critic network is: In the formula, s t 、a t and r t are the system state value, output action and reward value at the current moment, i.e., time t; γ is the discount factor, which is used to control the weight of future rewards; R(s t ,a t ) is state s t Take action a t The reward value; s0 is the initial state; a0 is the initial action; α is the temperature coefficient.

4. The servo parameter optimization method of a ball screw feed system based on a flexible motion-evaluation algorithm according to claim 1, characterized in that: The search environment of the SAC algorithm agent network includes the agent-environment interaction function and the search environment initialization function.

5. The servo parameter optimization method of a ball screw feed system based on a flexible motion-evaluation algorithm according to claim 4, characterized in that: The agent network needs to initialize the environment for each round of search. In each round of search, the agent network is used to calculate the reward value of the current action and the state of the next step, and to determine whether the current round has reached the target convergence value.

6. The method for optimizing servo parameters of a ball screw feed system based on a flexible motion-evaluation algorithm according to claim 1, characterized in that: The training parameters of the agent network include the number of training rounds, learning rate and maximum entropy value; the preset servo parameter values ​​are used as input to train the agent network, the next action is selected according to the current state, and the reward value is calculated after each round of search is completed until the training is stopped after the set number of rounds.

7. The method for optimizing servo parameters of a ball screw feed system based on a flexible motion-evaluation algorithm according to claim 1, characterized in that: The update formula of the critic network is: The update formula of the Actor network is: The update formula of α value is: In the formula, Q θ (s t ,a t ) is the Critic network in state s t Take action a t The expected cumulative reward; D is the data in the replay buffer; r(s t ,a t ) is state s t Take action a t The reward value of For the Actor network in state s t+1 The expected cumulative reward under φ is the current strategy sample; f φ (ε t ;s t ) uses the reparameterization technique, and the actions are sampled from the policy Gaussian distribution, that is, a t =f φ (ε t ;s t ), where f φ is a policy neural network with parameters φ, ε t is the initialization parameter of the network; Q θ (s t ,f φ (ε t ;s t )) is the target Critic network state s t Next take action f φ (ε t ;s t )’s expected cumulative reward; π t For state s t Take action a t probability; is the target value of entropy, which is used to automatically adjust the entropy regularization term.

8. The servo parameter optimization method of a ball screw feed system based on a flexible motion-evaluation algorithm according to any one of claims 1 to 7, characterized in that: The objective function reward of the reward value is: reward=-(tr+overshot+ts+td).

9. A servo parameter optimization system for a ball screw feed system based on a flexible motion-evaluation algorithm, characterized in that: The system includes a memory and a processor, the memory stores a computer program, and the processor executes the servo parameter optimization method of the ball screw feed system based on the flexible motion-evaluation algorithm as described in any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by the processor, the machine-executable instructions prompt the processor to implement the servo parameter optimization method of the ball screw feed system based on the flexible motion-evaluation algorithm as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Numerical-control machine tool feeding control compensation method based on Actor-Critic algorithm

    CN110488759A

  • Multi-agent deep reinforcement learning strategy optimization method based on attention mechanism

    CN113392935A

  • Mobility load balancing method based on reinforcement learning

    CN114598655A

  • Unmanned ship automatic berthing control method based on reinforcement learning

    CN115903474A

  • Controlling robots using entropy constraints

    US20220019866A1