Traveling wave rotating ultrasonic motor speed control method and system based on deep reinforcement learning

By building an LSTM dynamic speed model and a proximal strategy optimization algorithm based on deep reinforcement learning, an intelligent controller is designed to solve the modeling accuracy and control accuracy problems of the traveling wave rotary ultrasonic motor, and achieve high-precision and fast-response speed control.

CN120785209APending Publication Date: 2025-10-14NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510752817.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

In the existing technology, the traveling wave rotary ultrasonic motor has insufficient modeling accuracy, low control accuracy, slow response speed and poor adaptability to dynamic characteristics, which makes it difficult to achieve high-precision control under complex working conditions.

Method used

A deep reinforcement learning method is used to construct a dynamic velocity model based on LSTM, and an intelligent controller is designed in combination with the proximal policy optimization algorithm. By optimizing the reward function and hyperparameters, high-precision prediction and rapid response of TRUM are achieved.

Benefits of technology

In simulation and actual environments, the accuracy of the dynamic speed model reached 95%, the maximum steady-state error of the PPO controller was 0.21rpm, and the steady-state error was always kept below 0.93%, demonstrating good dynamic response capability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120785209A_ABST
    Figure CN120785209A_ABST
Patent Text Reader

Abstract

The invention provides a traveling wave rotating ultrasonic motor speed control method and system based on deep reinforcement learning. Firstly, a virtual dynamic model based on a long short-term memory (LSTM) network is constructed, the time-varying behavior of TRUM can be accurately captured, and the average prediction error is within 0.5%. Secondly, designing an intelligent controller based on a near-end strategy optimization PPO algorithm, and optimizing a reward function and structure to enhance the adaptability of a control strategy; simulation and experiment results show that the method realizes the quick response of 0.034 seconds and the tracking error of less than 0.21 rpm, and the error is maximally reduced by 0.42 rpm compared with the traditional controller. According to the method, the problems of insufficient modeling precision, low control precision, slow response speed, poor adaptive capacity to dynamic characteristics and the like in the prior art are effectively solved, and a new path is provided for application of the TRUM in high-precision complex tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of speed control of traveling wave rotary ultrasonic motor (TRUM), and particularly relates to a speed control method of traveling wave rotary ultrasonic motor based on deep reinforcement learning. BACKGROUND

[0002] Traveling wave rotary ultrasonic motor (TRUM) is widely used in aerospace, precision instruments and medical equipment due to its advantages of no electromagnetic interference, fast response, high driving precision, etc. However, during the operation of TRUM, the micro high-frequency vibration of piezoelectric ceramic is converted into the macro rotation of the output rotor, and the energy conversion, friction driving and wear are coupled with each other, resulting in the significant nonlinearity of the output characteristics, which makes the modeling and control of TRUM extremely complex.

[0003] Traditional modeling methods such as theoretical model, equivalent circuit model and identification model cannot fully capture the dynamic nonlinear characteristics of TRUM, and it is difficult to achieve high-precision modeling and control under complex working conditions. In addition, traditional speed controllers such as PID controller and fuzzy controller have poor adaptability when facing changing environment or interference, which limits their robustness in practical application.

[0004] Although deep learning has shown strong nonlinear fitting ability in TRUM modeling, however, TRUM data has significant time series characteristics, that is, the current state is not only affected by the current input, but also closely related to the previous input. Long short-term memory (LSTM) recurrent neural network (RNN) can not only solve the problem of gradient explosion, but also effectively capture the long-term and short-term dynamic changes of TRUM system. However, existing researches mostly train models based on stable state or limited dynamic data, which will lead to the loss of the ability of the model to accurately estimate when the dynamic characteristics are missing. Reinforcement learning (RL) has been successfully applied in many fields, including motor speed control. However, in TRUM control, the existing reinforcement learning method based on stable state model still has the shortcomings of slow response speed and large steady-state error. Therefore, further optimization of reinforcement learning method to adapt to the dynamic characteristics of TRUM is an important direction for future research. SUMMARY

[0005] In order to solve the problems of the prior art, the present application provides a speed control method of traveling wave rotary ultrasonic motor based on deep reinforcement learning, to solve the problems of insufficient modeling precision, low control precision, slow response speed and poor adaptability to dynamic characteristics in the prior art.

[0006] In order to achieve the above purpose, the speed control method of traveling wave rotary ultrasonic motor based on deep reinforcement learning can adopt the following technical solutions:

[0007] A traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning, comprising:

[0008] ultrasonic motor data collection;

[0009] design and train a dynamic speed model based on an LSTM model; the input of the dynamic speed model is the frequency value at consecutive time points, and the output of the dynamic speed model is the speed predicted according to historical information;

[0010] design and train an intelligent controller based on a proximal policy optimization algorithm; the design of the intelligent controller includes a reward function and input and output, and the intelligent controller adjusts the policy parameters according to the feedback signal provided by the reward function to pursue the maximum cumulative reward;

[0011] experimental verification.

[0012] Further, the ultrasonic motor data collection is based on three frequency change modes of linear change, step change and random jump.

[0013] Further, the dynamic speed model is designed according to the characteristics of the ultrasonic motor and the modeling and control requirements; the input is the frequency value at 10 consecutive time points, and the formula is:

[0014] i t =[f k ,f t ,f t ...f t ]

[0015] wherein i k is an array of 10 consecutive characteristic values starting from time step k, and f k+1 is the frequency value at the kth time point;

[0016] The output of the dynamic speed model is the speed predicted according to historical information, and the formula is:

[0017] o k+2 =[V k+9 ]

[0018] The input and output constitute an input-output pair, and the formula is:

[0019] i t =[f t ,f t ,f δ ...f s ]→o p =[V δ ]

[0020] The optimal hyperparameter combination is found by training the dynamic speed model, including learning rate, number of layers, number of hidden layer nodes, batch size and training period.

[0021] Further, the reward function formula is:

[0022] R t = R δ + R s - R p

[0023] Wherein, R δ is the tracking reward, R s is the stability reward, and R p is the penalty;

[0024] The tracking reward R δ gives the intelligent controller a certain reward according to the difference between the current traveling wave rotary ultrasonic motor speed and the target speed, and the formula is:

[0025] R δ = max[(1-δ), 0]

[0026] Wherein, δ = |V t -V p | represents the absolute value of the difference between the current speed and the target speed, and the smaller the absolute value of the speed difference, the greater the tracking reward;

[0027] The stability reward R s is given by the formula:

[0028] R s = 0.1

[0029] The penalty term R p gives the intelligent controller a certain penalty according to the change amount of the action at the current time and the previous time, and the formula is:

[0030] R p = 0.3log(1+Δa t )

[0031] Wherein, Δa t = |a t -a t-1 | represents the change amount of the action at the current time and the previous time;

[0032] The intelligent controller is composed of an actor and a critic, and the input includes the input of the actor and the input of the critic, which are the same, that is, the current t time state S t obtained by the ultrasonic motor from the environment, including speed v t and desired speed v r , and the formula is:

[0033] St = [v r ,v t ]

[0034] The output includes the output of the actor and the output of the critic; the actor decides the action a during training t The output of the actor is the average value of the action a t,μ and the variance of the action a t,σ The formula is:

[0035] a t,μ = [μ f ], a t,σ = [σ f ]

[0036] The input and output of the actor form an input-output pair, and the formula is:

[0037] S t = [v r ,v t ]→a t = [a t,μ ,a t,σ ]

[0038] The output of the critic is the state value V(S t ) at the current time t, and the input and output of the critic form an input-output pair, and the formula is:

[0039] S t = [v r ,v t ]→V(S t )

[0040] The intelligent controller based on the proximal policy optimization algorithm interacts with the environment, adjusts the policy according to the obtained feedback, maximizes the expected cumulative reward, and obtains the optimal policy π * The formula is:

[0041]

[0042] Wherein, the argmax function is used to find the policy π that maximizes the expected cumulative reward, E represents the expectation, that is, the average of all possible trajectories τ, τ represents the trajectory from the initial state to the terminal state, γ is the discount factor, which balances the importance of immediate reward and future reward, r t is the reward obtained at time t, and T is the length of the trajectory;

[0043] The proximal policy optimization algorithm uses a policy gradient method to guide policy update, and the formula is:

[0044]

[0045] where π θ is the current policy, parameterized by θ, a t is the action taken at time step t, s t is the state at time step t, is the advantage function, representing the relative advantage of taking action a t in state s t The advantage function is typically estimated by:

[0046]

[0047] where Q(s t ,a t ) is the Q-value of taking action a t in state s t is typically estimated by:

[0048] Q(s t ,a t ) = r t + γV(s t+1 )

[0049] where r t is the reward obtained at time step t, γ is the discount factor, and V(s t+1 ) is the value of the next state s t+1 estimated by the Critic network, the step size of policy update is limited by introducing a clipping function clip, the formula is:

[0050]

[0051] where π is the new policy, is the old policy, and ε is a hyperparameter, usually set to 0.2; the clipping function clip ensures that the probability ratio is within the range [1-ε, 1+ε], and the loss function of the Critic network is used to update the state value V(s), the formula is:

[0052]

[0053] where V is the state value estimate of the new policy, is the state value estimate of the old policy.

[0054] Further, the experimental verification includes simulation experimental verification and actual environment experimental verification.

[0055] The application also provides a traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning.

[0056] A traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning, comprising:

[0057] A data collection module for ultrasonic motor data collection;

[0058] A speed model building module for designing and training a dynamic speed model based on LSTM; the input of the dynamic speed model is the frequency value at consecutive time points, and the output of the dynamic speed model is the speed predicted according to historical information;

[0059] An intelligent control module for designing and training an intelligent controller based on a proximal policy optimization algorithm; the design of the intelligent controller includes a reward function and input and output, and the intelligent controller adjusts the policy parameters according to the feedback signal provided by the reward function to pursue the maximum cumulative reward;

[0060] An experimental verification module.

[0061] Further, in the data collection module, the ultrasonic motor data collection is based on three frequency change modes of linear change, step change and random jump.

[0062] Further, in the speed model building module, the dynamic speed model is designed for input and output according to the characteristics of the ultrasonic motor and the modeling and control requirements; the input is the frequency value at 10 consecutive time points, and the formula is:

[0063]

[0064] Where, i t is an array of 10 consecutive characteristic values starting from time step k, f k is the frequency value at the kth time point;

[0065] The output of the dynamic speed model is the speed predicted according to historical information, and the formula is:

[0066]

[0067] The input and output form an input-output pair, and the formula is:

[0068]

[0069] The optimal hyperparameter combination is found through the dynamic speed model training, including learning rate, number of layers, number of hidden layer nodes, batch size and training period.

[0070] Further, in the intelligent control module, the reward function formula is:

[0071] R t = R δ + R s - R p

[0072] wherein R δ is a tracking reward, R s is a stability reward, and R p is a penalty;

[0073] The tracking reward R δ is given to the intelligent controller according to the difference between the current traveling wave rotary ultrasonic motor speed and the target speed, and the formula is:

[0074] R δ = max[(1- δ), 0]

[0075] wherein δ = |V t -V p | represents the absolute value of the difference between the current speed and the target speed, and the tracking reward is greater when the absolute value of the speed difference is smaller;

[0076] The stability reward R s is given to the intelligent controller according to the difference between the current traveling wave rotary ultrasonic motor speed and the target speed, and the formula is:

[0077] R s = 0.1

[0078] The penalty term R p is given to the intelligent controller according to the change amount of the action at the current time and the previous time, and the formula is:

[0079] R p = 0.3 log(1 + Δa t )

[0080] wherein Δa t = |a t -a t-1 | represents the change amount of the action at the current time and the previous time;

[0081] The intelligent controller is composed of an actor and a critic, and the input includes the input of the actor and the input of the critic, which are the same, and is the current state S t of the ultrasonic motor obtained from the environment at time t, including the speed v t and the desired speed v r , and the formula is:

[0082] S t = [v r , v t ]

[0083] The output includes the output of the actor and the output of the critic; during training, the actor decides the action a t , and the actor output is the average value of the action a t,μ and the variance of the action a t,σ , and the formula is:

[0084] a t,μ =[μ f ], a t,σ =[σ f ]

[0085] The input and output of the actor form an input-output pair, and the formula is:

[0086] S t =[v r ,v t ]→a t =[a t,μ ,a t,σ ]

[0087] The output of the reviewer is the state value V(S t ), the commentator’s input and output form an input-output pair, and the formula is:

[0088] S t =[v r ,v t ]→V(S t )

[0089] The intelligent controller based on the proximal policy optimization algorithm interacts with the environment and continuously adjusts the strategy according to the feedback obtained to maximize the expected cumulative reward and obtain the optimal strategy π * , the formula is:

[0090]

[0091] Among them, the argmax function is used to find the strategy π that maximizes the expected cumulative reward, E represents the expectation, that is, the average of all possible trajectories τ, τ represents the trajectory from the initial state to the terminal state, γ is the discount factor used to balance the importance of immediate rewards and future rewards, r t is the reward obtained at time t, T is the length of the trajectory;

[0092] The proximal policy optimization algorithm uses the policy gradient method Guiding strategy update, the formula is:

[0093]

[0094] Among them, π θ is the current policy, with parameters θ, a t is the action taken at time step t, s t is the state at time step t, is the advantage function, which means that in state s t Next take action a t The relative advantage, advantage function is typically estimated by:

[0095]

[0096] where Q(s t ,a t ) is the Q-value of taking action a t in state s t , which is typically estimated by:

[0097] Q(s t ,a t ) = r t + γV(s t+1 )

[0098] where r t is the reward obtained at time step t, γ is the discount factor, and V(s t+1 ) is the value of the next state s t+1 estimated by the Critic network. To limit the step size of the policy update, a clipping function clip is introduced, which is defined as:

[0099]

[0100] where μ is the new policy, μ is the old policy, and ε is a hyperparameter typically set to 0.2. The clipping function clip ensures that the probability ratio is within the range [1-ε, 1+ε]. The loss function of the Critic network is used to update the state value V(s), which is defined as:

[0101]

[0102] where V is the state value estimate of the new policy, and V is the state value estimate of the old policy.

[0103] Further, in the experimental verification module, simulation experiment verification and actual environment experiment verification are included.

[0104] The application provides a TRUM speed control method and system based on deep reinforcement learning, aiming to solve the problems of insufficient modeling accuracy, low control accuracy, slow response speed and poor dynamic characteristic adaptation capability in the prior art. The solution includes constructing a dynamic speed model based on LSTM to accurately capture the time-varying behavior of the TRUM; in the model training stage, the hyperparameters are optimized to ensure high prediction accuracy of the model; an intelligent controller designed in combination with the PPO algorithm is used to optimize the reward function, realize fast response and high-precision speed tracking; finally, the simulation results show that the accuracy of the dynamic speed model is 95%, in the actual TRUM system, the maximum steady-state error of the PPO controller is 0.21 rpm, the minimum error achieves zero error control, and in the disturbance test, the maximum rise time is 80 ms, and the steady-state error is always kept below 0.93%, showing good dynamic response capability and robustness. BRIEF DESCRIPTION OF DRAWINGS

[0105] Figure 1 The flowchart of the TRUM speed control method and system based on deep reinforcement learning of the application.

[0106] Figure 2 The simulation experimental result graph of the application, (a) is a nonlinear speed response result graph, and (b) is the average error between experimental data and model prediction data and the maximum error e max Comparison chart.

[0107] Figure 3 The result comparison chart of different control methods under dynamic working conditions, (a) is the speed response of the TRUM in the starting, accelerating, stabilizing and decelerating stages, (b) is the driving frequency, (c) (d) (e) is Figure 3 The local enlarged view in (a).

[0108] Figure 4 The disturbance test result graph of the application. DETAILED DESCRIPTION

[0109] The application will be further illustrated below in combination with the drawings and specific embodiments, and it should be understood that the following specific embodiments are only used to illustrate the application and not to limit the scope of the application, and after reading the application, those skilled in the art can make various equivalent modifications of the application, which all fall within the scope defined by the claims attached hereto.

[0110] Please refer to Figure 1 The application provides a TRUM speed control method and system based on deep reinforcement learning, aiming to solve the problems of insufficient modeling accuracy, low control accuracy, slow response speed and poor dynamic characteristic adaptation capability in the prior art. The solution includes constructing a dynamic speed model based on LSTM to accurately capture the time-varying behavior of the TRUM; in the model training stage, the hyperparameters are optimized to ensure high prediction accuracy of the model; an intelligent controller designed in combination with the PPO algorithm is used to optimize the reward function, realize fast response and high-precision speed tracking; finally, the simulation results show that the accuracy of the dynamic speed model is 95%, in the actual TRUM system, the maximum steady-state error of the PPO controller is 0.21 rpm, the minimum error achieves zero error control, and in the disturbance test, the maximum rise time is 80 ms, and the steady-state error is always kept below 0.93%, showing good dynamic response capability and robustness.

[0111] Ultrasonic motor data collection;

[0112] The dynamic speed data is normalized to the range of [-1, 1] to obtain a dynamic speed data set, so as to accelerate the training process and improve the model performance;

[0113] The dynamic speed model is designed and trained based on the LSTM model; the input of the dynamic speed model is the frequency value at a continuous time point, and the output of the dynamic speed model is the predicted speed according to historical information; in the embodiment, the dynamic speed data set is divided into a training set, a validation set and a test set, the training set is used for training the dynamic speed model, the validation set is used for hyperparameter optimization, and the test set is used for simulating the experiment to evaluate the performance of the model, so as to ensure accurate prediction of the speed change of the TRUM; the work in this step not only optimizes the prediction accuracy of the model, but also provides reliable model support for the design of the subsequent intelligent controller;

[0114] The intelligent controller is designed and trained based on the proximal policy optimization algorithm PPO; the design of the intelligent controller includes a reward function and input and output, the intelligent controller adjusts the policy parameters according to the feedback signal provided by the reward function, and pursues the maximum cumulative reward;

[0115] Experimental verification.

[0116] In the step of collecting the ultrasonic motor data, the ultrasonic motor data is collected by using an experimental device based on the designed three frequency change modes, namely linear change, step change and random jump, and 30000 groups of dynamic speed data of the TRUM under different working conditions are collected;

[0117] In the step of designing and training the dynamic speed model based on the LSTM model, the LSTM model has excellent time series data processing capability, can effectively capture the dynamic change and time-varying behavior of the TRUM, and solves the deficiency of the traditional model in dynamic characteristic modeling, therefore, when the dynamic speed model is constructed, the ultrasonic motor driving frequency values at 10 continuous time points are used to effectively capture the main dynamic characteristics of the system, and the calculation efficiency of the model is maintained, and the input formula is:

[0118] i t =[f k ,f k+1 ,f k+2 ...f k+9 ]

[0119] Wherein, i t is an array of 10 continuous feature values starting from time step k, and f k is the frequency value at the kth time point;

[0120] The output of the dynamic speed model is the predicted speed according to historical information, and the formula is:

[0121] ot = [V t ]

[0122] The input and output constitute an input-output pair, and the formula is:

[0123] i t = [f k , f k+1 , f k+2 ... f k+9 ] → o t = [V t ]

[0124] Hyperparameters are important parameters of network structure and training process, including learning rate, number of layers, number of hidden layer nodes, batch size and training period, which have a significant impact on the accuracy of estimation, so it is necessary to find the best combination of hyperparameters for training. As shown in Table 1, 7 groups of parameters are set, group A selects hyperparameters according to experience as the control group, groups B-F change only one hyperparameter each time compared with group A, and group G adjusts multiple hyperparameters considering the experimental results of groups B-F to optimize performance. In this way, the influence of each hyperparameter on model performance can be systematically evaluated to determine the optimal combination of hyperparameters and improve prediction accuracy.

[0125] Table 1. Hyperparameter combinations of network

[0126]

[0127]

[0128] The above hyperparameter combinations of network are trained using the same data set, making them fair. In addition, all tests are conducted in the same environment with the same test standards, and the test results are shown in Table 2. The error indicators for evaluating the performance of these networks consist of two parts, one part is single operating point performance indicators, and the other part is full envelope performance indicators. Single operating point performance indicators are used to evaluate the performance of the model at a certain specific operating point, including the maximum average estimation error e max , the minimum average estimation error e min , and the maximum transient estimation error e Emax . Operating point refers to the state of a device or system under a certain specific operating condition, for example, the operating state of a motor at a certain specific frequency. Full envelope performance indicators are used to evaluate the comprehensive performance of the model in the entire operating range, including full envelope average estimation error e env , full envelope estimation error variance e Var , and all working point average maximum transient estimation error e amaxThe full envelope range refers to a collection of all possible operating conditions of a device or system, covering all working points.

[0129] Table 2. Training results of different hyperparameter combinations

[0130]

[0131] As can be seen from Table 2, each parameter has a certain degree of influence on the model performance. The C group uses a lower learning rate, and due to insufficient training of the network, the variance of the error increases, and the performance of each working point is inconsistent, resulting in a significant increase in error. The D group uses more layers, making it more difficult for the network to converge under the same training period, resulting in a larger error variance. The E group increases the batch size, which helps to regularize and enhance the generalization ability of the model, reducing the error. The G group considers the performance of the network in groups B-F, and selects a higher learning rate, appropriate number of layers, number of hidden layer nodes, batch size and training period to balance estimation accuracy and training cost, significantly reducing the error. The performance of the single working point and the full envelope performance are both excellent, significantly better than other combinations. In this embodiment, the G group of hyperparameter combinations is finally selected to ensure that the model can achieve high precision, fast response and high robustness of speed control in actual application.

[0132] In the steps of designing and training the intelligent controller based on the proximal policy optimization algorithm, the design of the intelligent controller includes the reward function and the input and output. The intelligent controller adjusts the policy parameters according to the feedback signal provided by the reward function and pursues the maximum cumulative reward. The reward function includes tracking reward, stability reward and penalty term. The tracking reward is used to reduce the speed error, the stability reward ensures the stability of the system under complex working conditions, and the penalty term limits the action change amount to avoid system instability. Through the design of the three parts, the controller based on the proximal policy optimization algorithm PPO can realize fast response and high precision speed tracking, significantly improving the control performance of the TRUM. The reward function formula is:

[0133] R t = R δ + R s - R p

[0134] wherein R δ is the tracking reward, R s is the stability reward, and R p is the penalty term.

[0135] The most important task in the speed control problem of ultrasonic motor is to track the speed reference with small tracking error δ. However, the contribution of error to the total reward should be different in different stages of control. For example, when the error is large, the speed is far from the target speed, and the intelligent controller should not be given any reward; when the error becomes small, the ultrasonic motor is closer to the target speed, and more reward should be given to encourage the intelligent controller. Therefore, the design of tracking reward is first considered, and the formula is:

[0136] R δ = max[(1- δ), 0]

[0137] where δ = |V t -V p | represents the absolute value of the difference between the current speed of the ultrasonic motor and the target speed, and the tracking reward becomes larger when the absolute value of the speed difference becomes smaller;

[0138] The second important task in the speed control problem of ultrasonic motor is to keep the system working continuously. In the initial stage of training, the intelligent controller sometimes cannot find the appropriate convergence direction, and incorrect controller parameters will cause the system to malfunction. In order to avoid the collapse of the system and ensure the stability of the speed and action, a stability reward needs to be designed, and the formula is:

[0139] R s = 0.1

[0140] The penalty term R p gives the intelligent controller a certain penalty according to the change amount of the action at the current time and the previous time, avoids large action changes that cause the system to be unstable, and encourages to take smaller actions to maintain stability, and the formula is:

[0141] R p = 0.3 log(1+ Δa t )

[0142] where Δa t = |a t -a t-1 | represents the change amount of the action at the current time and the previous time, and this penalty term makes small changes receive less punishment and large changes receive more punishment;

[0143] The intelligent controller is composed of an actor and a critic, and the input includes the input of the actor and the input of the critic, which are the same, the current speed v t and the desired speed v r of the ultrasonic motor obtained from the environment at time t, and the formula is:

[0144] S t = [v r , v t ]

[0145] The output includes the output of the actor and the output of the critic; the actor determines the action a during training t The actor output is the average value of the action a t,μ And the variance of the action a t,σ The formula is:

[0146] a t,μ = [μ f ], a t,σ = [σ f ]

[0147] The input and output of the actor form an input-output pair, and the formula is:

[0148] S t = [v r ,v t ]→a t = [a t,μ ,a t,σ ]

[0149] The output of the critic is the state value V(S t ) at the current time t, and the input and output of the critic form an input-output pair, and the formula is:

[0150] S t = [v r ,v t ]→V(S t )

[0151] The intelligent controller based on the proximal policy optimization algorithm interacts with the environment, adjusts the policy according to the obtained feedback, maximizes the expected cumulative reward, and obtains the optimal policy π * The formula is:

[0152]

[0153] Wherein, the argmax function is used to find the policy π that maximizes the expected cumulative reward, E represents the expectation, i.e. the average of all possible trajectories τ, τ represents the trajectory from the initial state to the terminal state, γ is the discount factor, which balances the importance of immediate rewards and future rewards, r t is the reward obtained at time t, and T is the length of the trajectory;

[0154] The proximal policy optimization algorithm uses a policy gradient method to guide policy update, and the formula is:

[0155]

[0156] Wherein, π θ is the current policy, and the parameter is θ, a tis the action taken at time step t, s t is the state at time step t, is the advantage function, representing the relative advantage of taking action a t in state s t . is typically estimated by:

[0157]

[0158] where Q(s t , a t ) is the Q-value of taking action a t in state s t , which is typically estimated by:

[0159] Q(s t , a t ) = r t + γV(s t+1 )

[0160] where r t is the reward obtained at time step t, γ is the discount factor, and V(s t+1 ) is the value of the next state s t+1 estimated by the Critic network. To limit the step size of policy updates, a clipping function is introduced, which is given by:

[0161]

[0162] where μ is the new policy, is the old policy, and ε is a hyperparameter typically set to 0.2. The clipping function ensures that the probability ratio is within the range [1-ε, 1+ε]. The loss function of the Critic network is used to update the state value V(s), which is given by:

[0163]

[0164] where V (s ) is the state value estimate of the new policy, is the state value estimate of the old policy.

[0165] By gradually optimizing the policy, the intelligent controller can take more optimized actions in the TRUM system, thereby achieving precise control of the TRUM speed. After the policy is updated, the intelligent controller adjusts the parameters of the policy network, including weights and biases, to more accurately select actions at each time step, making the TRUM speed closer to the target speed.

[0166] In the experimental verification step, both simulation experiment verification and actual environment experiment verification are included.

[0167] The simulation experiment uses the traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning to simulate to show the accuracy of the model, by sweeping up and down the TRUM frequency in the [37-41] kHz range, the nonlinear speed response is obtained. After 5 simulation tests, the average error and maximum error between the experimental data and the model prediction data are calculated.

[0168] After the proximal policy optimization algorithm PPO successfully completes the training, the intelligent controller is applied to the actual experimental environment. When deployed, the intelligent controller weights saved automatically by the program trained in the simulation stage are used for the control process of the ultrasonic motor, instead of randomly initializing network parameters. This pre-trained weight strategy not only speeds up the learning speed, but also ensures the stability of the initial behavior of the intelligent controller, which is crucial for the safe operation of the hardware system, and is more reliable than random initialization. The actual environment experiment compares the average response time, average maximum overshoot, maximum steady-state error and minimum steady-state error results of the speed response during the movement of the traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning with traditional controllers (PID, LQR) and intelligent controllers based on other reinforcement learning algorithms (DDPG, SAC).

[0169] Referring to Figure 1 The present application provides a TRUM speed control system based on deep reinforcement learning, which comprises:

[0170] A data collection module for ultrasonic motor data collection;

[0171] A preprocessing module for ultrasonic motor data preprocessing; in this embodiment, the dynamic speed data is normalized to the range of [-1, 1] to obtain a dynamic speed dataset, in order to speed up the training process and improve the model performance;

[0172] A speed model building module for designing and training a dynamic speed model based on LSTM; the input of the dynamic speed model is the frequency value at consecutive time points, and the output of the dynamic speed model is the predicted speed based on historical information; in this embodiment, the dynamic speed dataset is divided into a training set, a validation set and a test set, the training set is used for dynamic speed model training, the validation set is used for hyperparameter optimization, and the test set is used for simulation experiment to evaluate the performance of the model and ensure accurate prediction of the speed change of the TRUM; the work in this step not only optimizes the prediction accuracy of the model, but also provides reliable model support for the design of the subsequent intelligent controller;

[0173] The intelligent control module is used for designing and training an intelligent controller based on a proximal policy optimization algorithm PPO, and the design of the intelligent controller includes a reward function and input and output, and the intelligent controller adjusts policy parameters according to feedback signals provided by the reward function, and pursues maximum accumulated rewards.

[0174] The experimental verification module.

[0175] In the data collection module, the ultrasonic motor data are collected by using an experimental device based on three designed frequency change modes, namely linear change, step change and random jump, and 30,000 groups of dynamic speed data of the TRUM under different working conditions are collected.

[0176] In the speed model building module, the LSTM model has excellent time series data processing capability, can effectively capture the dynamic change and time-varying behavior of the TRUM, and solves the deficiency of the traditional model in dynamic characteristic modeling, therefore, when the dynamic speed model is constructed, 10 continuous time point frequency values of the ultrasonic motor are used to effectively capture the main dynamic characteristics of the system and maintain the calculation efficiency of the model, and the input formula is:

[0177] i t =[f k ,f k+1 ,f k+2 ...f k+9 ]

[0178] Wherein, i t is an array of 10 continuous feature values starting from time step k, and f k is the frequency value at the kth time point.

[0179] The output of the dynamic speed model is the speed predicted according to the historical information, and the formula is:

[0180] o t =[V t ]

[0181] The input and output constitute an input-output pair, and the formula is:

[0182] i t =[f k ,f k+1 ,f k+2 ...f k+9 ]→o t =[V t ]

[0183] The hyperparameters are important parameters of the network structure and training process, including learning rate, number of layers, number of hidden nodes, batch size, and training period, which have a significant impact on the accuracy of estimation, so it is necessary to find the best combination of hyperparameters for training. As shown in Table 1, seven groups of parameters are set, group A selects hyperparameters according to experience as the control group, groups B-F change one hyperparameter each time compared to group A, and group G modifies multiple hyperparameters considering the situation of groups B-F.

[0184] Table 1. Hyperparameter combinations of the network

[0185] Hyperparameter combinations Learning rate Number of layers Number of hidden nodes Batch size Training period A 0.001 2 64 16 20 B 0.001 2 32 16 20 C 0.0001 2 64 16 20 D 0.001 4 64 16 20 E 0.001 2 64 32 20 F 0.001 2 64 32 40 G 0.001 2 64 32 30

[0186] The networks with the above hyperparameter combinations are trained using the same data set, making them fair. In addition, all tests are conducted in the same environment, with the same test standards, and the test results are shown in Table 2. The error indicators for evaluating the performance of these networks consist of two parts: single operating point performance indicators and full envelope performance indicators. Single operating point performance indicators are used to evaluate the performance of the model at a specific operating point, including the maximum average estimation error e max , the minimum average estimation error e min , and the maximum transient estimation error e Emax . An operating point refers to the state of a device or system under a specific operating condition, such as the operating state of a motor at a specific frequency. Full envelope performance indicators are used to evaluate the comprehensive performance of the model in the entire operating range, including the full envelope average estimation error e env , the full envelope estimation error variance e Var , and the average maximum transient estimation error e amax at all operating points. The full envelope range refers to the set of all possible operating conditions of a device or system, covering all operating points.

[0187] Table 2. Training results of different hyperparameter combinations

[0188]

[0189] From the results of Table 2, it can be seen that each parameter has a certain degree of influence on the performance of the model. The C group uses a lower learning rate, and due to insufficient training of the network, the variance of the error increases, and the performance of each working point is inconsistent, resulting in a significant increase in error. The D group uses more layers, making it more difficult for the network to converge under the same training period, resulting in a larger error variance. The E group increases the batch size, which helps to regularize and enhance the generalization ability of the model, reducing the error. The G group considers the performance of the network in groups B-F, and selects a higher learning rate, appropriate number of layers, number of hidden layer nodes, batch size, and training period to balance estimation accuracy and training cost, significantly reducing the error. In this embodiment, the G group of hyperparameters is finally selected.

[0190] In the intelligent control module, the design of the intelligent controller includes a reward function and input and output. The intelligent controller adjusts the policy parameters according to the feedback signal provided by the reward function, and pursues the maximum cumulative reward. The reward function includes a tracking reward, a stability reward, and a penalty term. The tracking reward is used to reduce the speed error, the stability reward ensures the stability of the system under complex working conditions, and the penalty term limits the amount of action change to avoid system instability. Through the design of the three parts, the controller based on the proximal policy optimization algorithm PPO can achieve fast response and high-precision speed tracking, significantly improving the control performance of the TRUM. The reward function formula is:

[0191] R t = R δ + R s - R p

[0192] Wherein, R δ is the tracking reward, R s is the stability reward, and R p is the penalty term.

[0193] The most important task in the speed control problem of ultrasonic motor is to track the speed reference with a small tracking error δ. However, the contribution of the error to the total reward should be different at different stages of control. For example, when the error is large, the speed is far from the target speed, and the intelligent controller should not be given any reward; when the error becomes small, the ultrasonic motor is closer to the target speed, and more reward should be given to encourage the intelligent controller. Therefore, the design of the tracking reward is first considered, and the formula is:

[0194] R δ = max(1- δ, 0)

[0195] Wherein, δ = |V t -V p | represents the absolute value of the difference between the current traveling wave rotary ultrasonic motor speed and the target speed. The smaller the absolute value of the speed difference, the greater the tracking reward.

[0196] The second important task in the speed control problem of ultrasonic motor is to keep the system working continuously. In the initial stage of training, sometimes the intelligent controller cannot find the appropriate convergence direction, and incorrect controller parameters will cause the system to malfunction. In order to avoid the collapse of the system and ensure the stability of the speed and action, a stability reward needs to be designed, which is:

[0197] R s = 0.1

[0198] The penalty term R p gives the intelligent controller a certain punishment according to the change amount of the action at the current time and the previous time, avoids large action changes leading to system instability, and encourages to take smaller actions to maintain stability, which is:

[0199] R p = 0.3log(1+Δa t )

[0200] Where Δa t = |a t -a t-1 | represents the change amount of the action at the current time and the previous time, and this penalty term makes small changes receive less punishment and large changes receive more punishment.

[0201] The intelligent controller is composed of an actor and a critic, and the input includes the input of the actor and the input of the critic, which are the same, that is, the current t time speed v t and the expected speed v r obtained by the ultrasonic motor from the environment, which is:

[0202] S t = [v r , v t ]

[0203] The output includes the output of the actor and the output of the critic; during training, the actor decides the action a t , and the actor output is the average value of the action a t,μ and the variance of the action a t,σ , which is:

[0204] a t,μ = [μ f ], a t,σ = [σ f ]

[0205] The input and output of the actor form an input-output pair, which is:

[0206] S t = [v r , v t ]→at = [a t,μ , a t,σ ]

[0207] The output of the critic is the state value V(S t ) at the current time t, and the input and output of the critic form an input-output pair, which is given by:

[0208] S t = [v r , v t ] → V(S t )

[0209] The intelligent controller based on the proximal policy optimization algorithm interacts with the environment, adjusts the policy according to the feedback obtained, maximizes the expected cumulative reward, and obtains the optimal policy π * , which is given by:

[0210]

[0211] where the argmax function is used to find the policy π that maximizes the expected cumulative reward, E represents the expectation, i.e., the average over all possible trajectories τ, τ represents a trajectory from the initial state to the terminal state, γ is a discount factor that balances the importance of immediate rewards and future rewards, r t is the reward obtained at time t, and T is the length of the trajectory.

[0212] The proximal policy optimization algorithm uses a policy gradient method to guide policy updates, which is given by:

[0213]

[0214] where π θ is the current policy with parameters θ, a t is the action taken at time step t, s t is the state at time step t, is the advantage function, which represents the relative advantage of taking action a t in state s t . The advantage function is usually estimated by:

[0215]

[0216] where Q(s t , a t ) is the Q-value of taking action a t in state s t , which is usually estimated by:

[0217] Q(s ta t )=r t +γV(s t+1 )

[0218] where r t is the reward obtained at time step t, γ is the discount factor, V(s t+1 ) is the value of the next state s t+1 estimated by the Critic network, and the step size of policy update is limited by introducing a clipping function clip, which is defined as:

[0219]

[0220] where μ is the new policy, is the old policy, and ε is a hyperparameter usually set to 0.2; the clipping function clip ensures that the probability ratio is within the range [1-ε, 1+ε], and the loss function of the Critic network is used to update the state value V(s), which is defined as:

[0221]

[0222] where V is the state value estimate of the new policy, is the state value estimate of the old policy.

[0223] By gradually optimizing the policy, the intelligent controller can take more optimized actions in the TRUM system, thereby achieving precise control of the TRUM speed. After the policy is updated, the intelligent controller adjusts the parameters of the policy network, including weights and biases, to select actions more accurately at each time step, making the TRUM speed closer to the target speed.

[0224] In the experimental verification module, simulation experiment verification and actual environment experiment verification are included.

[0225] The simulation experiment uses the TRUM speed control method based on deep reinforcement learning to simulate to show the accuracy of the model. By sweeping up and down the TRUM frequency in the range of [37-41] kHz, the nonlinear speed response is obtained. After 5 simulation tests, the average error and maximum error between the experimental data and the model prediction data are calculated.

[0226] After the proximal policy optimization algorithm PPO successfully completes the training, the intelligent controller is applied to the actual experimental environment. When deployed, the pre-trained weight of the intelligent controller saved automatically during the simulation stage is used for the control process of the ultrasonic motor, rather than randomly initializing the network parameters. This pre-trained weight strategy not only speeds up the learning process, but also ensures the stability of the initial behavior of the intelligent controller, which is crucial for the safe operation of the hardware system and is more reliable than random initialization. The actual environment experiment compares the average response time, average maximum overshoot, maximum steady-state error, and minimum steady-state error of the speed response of the traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning with those of traditional controllers (PID, LQR) and intelligent controllers based on other reinforcement learning algorithms (DDPG, SAC) during the movement process.

[0227] Referring to Figure 2 , a simulation embodiment of the traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning is shown, and the detailed implementation steps are described below.

[0228] The simulation experiment uses the dynamic speed model to perform up-sweep and down-sweep on the TRUM frequency in the [37-41] kHz range, and obtains the nonlinear speed response, as shown in Figure 2 (a). To ensure the stability and reliability of the results, the experiment is performed 5 times independently, each time under the same initial conditions to reduce the influence of random errors, and the average error and the maximum error e max between the experimental data and the model prediction data are calculated, as shown in Figure 2 (b).

[0229] As can be seen from Figure 2 (a), the predicted speed curve coincides with the actual speed curve, indicating that the model well masters the dynamic law of the TRUM.

[0230] As can be seen from Figure 2 (b), during the entire simulation test process, the average error of the frequency down-sweep and the frequency up-sweep is less than 0.5%, and the maximum error is less than 1.6%, indicating that the dynamic speed model has good prediction performance under different frequency change modes and can effectively capture the dynamic characteristics of the TRUM.

[0231] Referring to Figure 3 , a steady-state control test comparison embodiment of the traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning and various controllers in the actual environment is shown, and the detailed implementation steps are described below.

[0232] After the PPO successfully completes the training, the intelligent controller is applied to the actual experimental environment. When deploying, the intelligent controller weights trained in the simulation stage are used instead of randomly initializing network parameters. This pre-training weight strategy not only speeds up the learning speed, but also ensures the stability of the initial behavior of the intelligent controller, which is crucial for the safe operation of the hardware system and is more reliable than random initialization.

[0233] The average response time, average maximum overshoot, maximum steady-state error, and minimum steady-state error of the speed response during the movement of the deep reinforcement learning PPO-based traveling wave rotary ultrasonic motor speed control method are compared with those of the traditional controllers (PID, LQR) and intelligent controllers based on other reinforcement learning algorithms (DDPG, SAC). The performance comparison results of the five control methods are shown in Table 3 and Figure 3 .

[0234] Table 3. Performance comparison of five control methods

[0235] PPO SAC DDPG PID LQR Response time / s 0.034 0.045 0.051 0.053 0.065 Maximum overshoot / % 0.232 0.473 1.214 2.652 1.852 Maximum steady state error / rpm 0.21 0.47 0.56 0.52 0.63 Minimum steady state error / rpm 0 0.07 0.07 0.05 0.09

[0236] From Table 3 and Figure 3 , it can be seen that the PPO controller has better dynamic response capability and robustness in response time, maximum overshoot, and steady-state error compared to the other four controllers. It can achieve a fast response of 0.034 seconds, a maximum steady-state error of 0.21 rpm, and a minimum error of zero error control, with a maximum error reduction of 0.42 rpm compared to the traditional controller.

[0237] Referring to Figure 4 , a disturbance test embodiment of the deep reinforcement learning-based traveling wave rotary ultrasonic motor speed control system is shown. To test whether the deep reinforcement learning-based traveling wave rotary ultrasonic motor speed control system can restore the ultrasonic motor to its original operating state in time when suddenly disturbed, the detailed implementation steps are described below.

[0238] Table 4. Disturbance test results

[0239] Low speed Medium speed High speed Maximum rise time / s 0.06 0.08 0.07 Maximum steady state error / % 0.93 0.15 0.56

[0240] In the initial stage of the test, the deep reinforcement learning-based traveling wave rotary ultrasonic motor speed control system runs at low speed (10 rpm), medium speed (70 rpm), and high speed (140 rpm) initial speed commands, respectively. Then, at 0.5 seconds, 1.5 seconds, and 2.5 seconds, the system is subjected to different amplitude disturbances in turn, and the disturbance amplitude gradually increases. The control system's ability to suppress sudden disturbances and the system's dynamic recovery performance under different operating conditions are comprehensively investigated. The maximum rise time and maximum steady-state error under low-speed, medium-speed, and high-speed conditions are obtained, as shown in Table 4 and Figure 4 .

[0241] From Table 4 and Figure 4 It can be seen that the maximum rising time of the state adjustment process of the system is 0.08 seconds after each disturbance, and the steady-state error is always kept below 0.93%, which shows that the traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning also has good dynamic response ability and robustness when facing different speed and disturbance conditions.

Claims

1. A traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning, characterized by: include: Ultrasonic motor data collection; Design and train a dynamic speed model based on the LSTM model; The input of the dynamic speed model is the frequency value of continuous time points, and the output of the dynamic speed model is the speed predicted based on historical information; Design and train an intelligent controller based on a proximal policy optimization algorithm. The design of the intelligent controller includes a reward function and input and output. The intelligent controller adjusts the policy parameters based on the feedback signal provided by the reward function to maximize the accumulated reward. Experimental verification.

2. The traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning according to claim 1 is characterized in that: The ultrasonic motor data is collected based on three frequency change modes: linear change, step change and random jump.

3. The traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning according to claim 1, characterized in that: The dynamic speed model is designed with input and output according to the characteristics of the ultrasonic motor and the modeling and control requirements; the input is the frequency value of 10 consecutive time points, and the formula is: i t =[f k ,f k+1 ,f k+2 ...f k+9 ] Among them, i t is an array of 10 consecutive eigenvalues ​​starting from time step k, f k is the frequency value at the kth time point; The output of the dynamic speed model is the speed predicted based on historical information, and the formula is: about t =[V t ] The input and output form an input-output pair, and the formula is: i t =[f k ,f k+1 ,f k+2 ...f k+9 ]→o t =[V t ] The dynamic velocity model training is used to find the optimal hyperparameter combination, including learning rate, number of layers, number of hidden layer nodes, batch size, and training cycle.

4. The traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning according to claim 1, characterized in that: The reward function formula is: R t =R δ +R s -R p Among them, R δ To track rewards, R s is the stability reward, R p for punishment; Tracking Rewards R δ The intelligent controller is given a certain reward based on the difference between the current traveling wave rotating ultrasonic motor speed and the target speed. The formula is: R δ =max[(1-δ),0] Where, δ=|V t -V p | represents the absolute value of the difference between the current speed and the target speed. The smaller the absolute value of the speed difference, the greater the tracking reward. Stability Reward R s The formula is: R s =0.1 Penalty term R p The intelligent controller is given a certain penalty based on the change in action between the current moment and the previous moment. The formula is: R p =0.3log(1+Δa t ) Where Δa t =|a t -a t-1 |Indicates the change in the action between the current moment and the previous moment; The intelligent controller consists of an actor and a commentator. The input includes the actor's input and the commentator's input, which are the same and are the state S of the ultrasonic motor at the current time t obtained from the environment. t , including the speed v t and the desired velocity v r , the formula is: S t =[v r ,v t ] The output includes the output of the actor and the output of the critic; during training, the actor decides the action a t , the actor output is the average value a of the action t,μ and the variance of the action a t,σ , the formula is: a t,μ =[μ f ],a t,σ =[s f ] The input and output of the actor form an input-output pair, and the formula is: With t =[in r ,in t ]→a t =[a t,μ ,and t,σ ] The output of the reviewer is the state value V(S t ), the commentator’s input and output form an input-output pair, and the formula is: S t =[v r ,v t ]→V(S t ) The intelligent controller based on the proximal policy optimization algorithm interacts with the environment and continuously adjusts the strategy according to the feedback obtained to maximize the expected cumulative reward and obtain the optimal strategy π * , the formula is: Among them, the argmax function is used to find the strategy π that maximizes the expected cumulative reward, E represents the expectation, that is, the average of all possible trajectories τ, τ represents the trajectory from the initial state to the terminal state, γ is the discount factor used to balance the importance of immediate rewards and future rewards, r t is the reward obtained at time t, T is the length of the trajectory; The proximal policy optimization algorithm uses the policy gradient method ▽ θ J(θ) guides the policy update, the formula is: Among them, π θ is the current policy, with parameters θ, a t is the action taken at time step t, s t is the state at time step t, is the advantage function, which means that in state s t Next take action a t The relative advantage, advantage function It is usually estimated by: Among them, Q(s t ,a t ) is in state s t Next take action a t The Q value is usually estimated as follows: Q(s t ,a t )=r t +γV(s t+1 ) Among them, r t is the reward obtained at time step t, γ is the discount factor, V(s t+1 ) is the next state s estimated by the Critic network t+1 The value of is introduced by the clip function clip to limit the step size of the strategy update. The formula is: in, It's a new strategy. is the old strategy, ε is a hyperparameter, usually set to 0.2; the clip function clip ensures that the probability ratio is Within the range, the loss function of the Critic network is used to update the state value V(s), and the formula is: in, is the state value estimate of the new policy, is the state value estimate of the old policy.

5. The traveling wave rotary ultrasonic motor speed control method based on deep reinforcement learning according to claim 1, characterized in that: The experimental verification includes simulation experiment verification and actual environment experiment verification.

6. A traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning, characterized by: include: Data collection module, used for ultrasonic motor data collection; Speed ​​model building module, used to design and train LSTM-based dynamic speed models; The input of the dynamic speed model is the frequency value of continuous time points, and the output of the dynamic speed model is the speed predicted based on historical information; The intelligent control module is used to design and train an intelligent controller based on a proximal policy optimization algorithm. The design of the intelligent controller includes a reward function and input and output. The intelligent controller adjusts the policy parameters based on the feedback signal provided by the reward function to maximize the accumulated reward. Experimental verification module.

7. The traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning according to claim 1, characterized in that: In the data collection module, the ultrasonic motor data is collected based on three frequency change modes: linear change, step change and random jump.

8. The traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning according to claim 1, characterized in that: In the velocity model building module, the dynamic velocity model is designed for input and output according to the characteristics of the ultrasonic motor and the modeling and control requirements; the input is the frequency value of 10 consecutive time points, and the formula is: i t =[f k ,f k+1 ,f k+2 ...f k+9 ] Among them, i t is an array of 10 consecutive eigenvalues ​​starting from time step k, f k is the frequency value at the kth time point; The output of the dynamic speed model is the speed predicted based on historical information, and the formula is: about t =[V t ] The input and output form an input-output pair, and the formula is: i t =[f k ,f k+1 ,f k+2 ...f k+9 ]→o t =[V t ] The dynamic velocity model training is used to find the optimal hyperparameter combination, including learning rate, number of layers, number of hidden layer nodes, batch size, and training cycle.

9. The traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning according to claim 1, characterized in that: In the intelligent control module, the reward function formula is: R t =R δ +R s -R p Among them, R δ To track rewards, R s is the stability reward, R p for punishment; Tracking Rewards R δ The intelligent controller is given a certain reward based on the difference between the current traveling wave rotating ultrasonic motor speed and the target speed. The formula is: R δ =max[(1-δ),0] Where, δ=|V t -V p | represents the absolute value of the difference between the current speed and the target speed. The smaller the absolute value of the speed difference, the greater the tracking reward. Stability Reward R s The formula is: R s =0.1 Penalty term R p The intelligent controller is given a certain penalty based on the change in action between the current moment and the previous moment. The formula is: R p =0.3log(1+Δa t ) Where Δa t =|a t -a t-1 |Indicates the change in the action between the current moment and the previous moment; The intelligent controller consists of an actor and a commentator. The input includes the actor's input and the commentator's input, which are the same and are the state S of the ultrasonic motor at the current time t obtained from the environment. t , including the speed v t and the desired velocity v r , the formula is: S t =[v r ,v t ] The output includes the output of the actor and the output of the critic; during training, the actor decides the action a t , the actor output is the average value a of the action t,μ and the variance of the action a t,σ , the formula is: a t,μ =[μ f ],a t,σ =[s f ] The input and output of the actor form an input-output pair, and the formula is: With t =[in r ,in t ]→a t =[a t,μ ,and t,σ ] The output of the reviewer is the state value V(S t ), the commentator’s input and output form an input-output pair, and the formula is: S t =[v r ,v t ]→V(S t ) The intelligent controller based on the proximal policy optimization algorithm interacts with the environment and continuously adjusts the strategy according to the feedback obtained to maximize the expected cumulative reward and obtain the optimal strategy π * , the formula is: Among them, the argmax function is used to find the strategy π that maximizes the expected cumulative reward, E represents the expectation, that is, the average of all possible trajectories τ, τ represents the trajectory from the initial state to the terminal state, γ is the discount factor used to balance the importance of immediate rewards and future rewards, r t is the reward obtained at time t, T is the length of the trajectory; The proximal policy optimization algorithm uses the policy gradient method ▽ θ J(θ) guides the policy update, the formula is: Among them, π θ is the current policy, with parameters θ, a t is the action taken at time step t, s t is the state at time step t, is the advantage function, which means that in state s t Next take action a t The relative advantage, advantage function It is usually estimated by: Among them, Q(s t ,a t ) is in state s t Next take action a t The Q value is usually estimated as follows: Q(s t ,a t )=r t +γV(s t+1 ) Among them, r t is the reward obtained at time step t, γ is the discount factor, V(s t+1 ) is the next state s estimated by the Critic network t+1 The value of is introduced by the clip function clip to limit the step size of the strategy update. The formula is: in, It's a new strategy. is the old strategy, ε is a hyperparameter, usually set to 0.2; the clip function clip ensures that the probability ratio is in the range [1-ε, 1+ε]. The loss function of the Critic network is used to update the state value V(s), and the formula is: in, is the state value estimate of the new policy, is the state value estimate of the old policy.

10. The traveling wave rotary ultrasonic motor speed control system based on deep reinforcement learning according to claim 1, characterized in that: The experimental verification module includes simulation experiment verification and actual environment experiment verification.

Citation Information

Cited By

  • Verifiable alignment optimization method for motor parameter reasoning identification and related device

    CN121960810A