Actuator control method and controller based on online neural prediction
Through the actuator control method based on online neural prediction and the controller designed by Zynq system-on-chip, the problems of low control accuracy and insufficient adaptability of traditional control methods and controllers in the face of nonlinearity and parameter uncertainty are solved, and high-precision, multi-channel and scalable control effects are achieved.
Patent Information
- Application Number
- CN202310057037.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-01-18
AI Technical Summary
When traditional actuator control methods and controllers face nonlinearity and parameter uncertainty, their control accuracy is low and their adaptability is insufficient, making it difficult to meet the needs of high precision, multi-channel and scalable.
A method of actuator control based on online neural prediction is proposed. Through hyperparameter adaptation, online system identification and neural prediction control, the neural network is used to simplify nonlinear system modeling, and hardware acceleration and implementation is carried out through the controller designed by Zynq system-on-chip.
It improves the control accuracy and scope of application of the actuator, can better track reference values in various industrial applications, and enhances the scalability and adaptability of the controller.
Smart Images

Figure CN116339136B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an adaptive control system, and in particular to an actuator control method and a controller based on online neural prediction. Background Art
[0002] Actuators are a common output device in motion control systems and are widely used in industrial scenarios such as vibration tables, automatic suspension systems, robots, and offshore gangways. A typical actuator system consists of a controller, a sensor, a load, and the actuator itself. The controller is the carrier of the control method, responsible for receiving sensor signals and outputting drive signals to excite the actuator. The control method and controller performance largely determine the control effect of the actuator.
[0003] There are a lot of nonlinearities and parameter uncertainties in actuators. How to control them with high precision has long been a challenging problem. Most nonlinear or adaptive control methods require detailed theoretical modeling of the nonlinear terms or systems to be compensated in advance, and often rely on full-state feedback, which is difficult to achieve in cost-sensitive or structurally constrained systems. Although the optimal control methods related to model predictive control have lower requirements for the accuracy of theoretical modeling due to the characteristics of rolling optimization, they are generally based on the system model obtained by offline identification for predictive control, and the model is not adjusted as the system dynamics change during the control process. These traditional control methods have, to a certain extent, limited the improvement of control performance and their applicability in various industrial applications.
[0004] The development of controllers in motion control systems has a history of many years. Domestic products are relatively simple and have the following shortcomings: 1) The core architecture is backward, most of which are built on DSP. The serial calculation method limits the increase in control frequency, resulting in low control accuracy; 2) The number of channels is small. In tasks such as high-speed railway subgrade testing, the number of input and output channels required is very large, and traditional controllers cannot meet the needs; 3) The expansion method is not flexible. Generally, the installation and uninstallation of expansions are completed by plugging and unplugging functional modules on limited bus interfaces. The reliability is not high and the number of expansions is limited. Therefore, in the controller design process, high precision, multi-channel and scalability are urgent requirements that need to be considered. Summary of the invention
[0005] Aiming at the shortcomings of traditional control methods and controllers, such as poor control performance and insufficient adaptability of actuators, the present invention proposes an actuator control method and controller based on online neural prediction. The proposed control method and controller are used to control the valve core displacement of the actuator so that it can accurately track the reference value. The controller, as a device carrier of the control method, can perform hardware acceleration and implementation of the proposed control method, and has strong applicability and scalability.
[0006] The objective of the present invention is achieved through the following technical solutions:
[0007] According to a first aspect of the present invention, there is provided an actuator control method based on online neural prediction, the method comprising hyperparameter adaptation, online system identification and neural predictive control, and specifically comprising the following steps:
[0008] Step 1: Hyperparameter adaptation is achieved by maximizing the cumulative reward, which includes the following sub-steps:
[0009] Step 1-1. Optional hyperparameter values for each hyperparameter Calculation control starts at the current time step k and takes optional hyperparameter values The sampling mean of the reward obtained when
[0010] Step 1-2: Use the Hoefding inequality to estimate the upper bound of the reward distribution under the confidence interval of 1 in Select hyperparameters for the start to the current time step k The number of times;
[0011] Step 1-3: Select the upper bound of the confidence interval with the largest value The optimal hyperparameters for As the hyperparameter value of the current time step k;
[0012] Step 2: Online system identification, using a neural network to learn the time-varying dynamics of the actuator system during the control process, includes the following sub-steps:
[0013] Step 2-1: Use a fully connected neural network with parameter θ(k-1) to predict the actuator valve core displacement value x at the current time step k p (k), where the parameter θ(k-1) includes the weights and biases of each layer of the network;
[0014] Step 2-2, using x p (k) and the actual displacement value x measured by the sensor d (k) Obtain the objective function value J θ (k)=(x d (k)-x p (k)) 2 / 2;
[0015] Step 2-3: Find J by back propagation θ (k) Gradient of the neural network parameter θ(k-1)
[0016] Step 2-4: Gradient descent updates neural network parameters Where η is the learning rate;
[0017] Step 3: Neural predictive control obtains the optimal control input based on the neural network system model obtained by online system identification, including the following sub-steps:
[0018] Step 3-1, set the initial iteration t = 0, the initial control amount u c (t|k)=0;
[0019] Step 3-2: Use a neural network with parameter θ(k) to predict the displacement value x at the next time step k+1 p (t|k+1);
[0020] Step 3-3, using x p (t|k+1) and the set reference value r d (k+1) Get the objective function value J u (t|k)=(r d (k+1)-x p (t|k+1)) 2 / 2;
[0021] Step 3-4: Find J by back propagation u (t|k) for u in the neural network input layer c The gradient of (t|k)
[0022] Step 3-5: Gradient descent to update the control amount Where μ is the optimization rate;
[0023] Step 3-6, execute t=t+1, if t<N g , go to step 3-2, otherwise go to step 3-7, where N g is the required number of iterations;
[0024] Step 3-7, iteratively obtained u c (t|k) is the optimal control quantity u to be solved c (k).
[0025] Furthermore, the hyperparameter in step 1 is specifically a learning rate η or an optimization rate μ.
[0026] Furthermore, the hyperparameters in step 1 require manual definition of the form of reward.
[0027] Furthermore, in step 1-1, it is necessary to define several optional hyperparameter values for each hyperparameter. Finite set of Optional set of hyperparameters.
[0028] Furthermore, in step 2-1, it is necessary to define the number of hidden layers N of the neural network.h and the number of hidden layer neurons N n , the input of the neural network is [x d (k-1),…, d (km), c (k-1),…, c (kn)] T , the output of the neural network is x p (k), where m and n are the orders of the displacement state value and the control quantity considered, respectively.
[0029] Furthermore, in step 3-2, the input of the neural network is [x d (k),…, d (k-m+1), c (t|k), c (k-1),…, c (k-n+1)] T , the output of the neural network is x p (t|k+1).
[0030] According to a second aspect of the present specification, there is provided a controller for hardware acceleration and implementation of a control method, the controller comprising a main control module, a power module, a storage module, an input module, an output module, a communication module and a standby module;
[0031] The main control module includes the Zynq system-on-chip and its peripheral circuits. The Zynq system-on-chip is divided into a processing system and programmable logic, which is responsible for the coordination and logical operations between the various modules of the controller. It completes the interaction with the host computer and the hardware acceleration and implementation of the control method through software and hardware programming.
[0032] The power module includes a DC power interface, a filter circuit and a DC / DC regulator, and is used to power the controller;
[0033] The storage module includes DDR, a storage medium whose data disappears when power is off, and is used to cache intermediate variables during the operation of the algorithm; and also includes SDIO, SPI Flash, EEPROM and eMMC, a storage medium whose data does not disappear when power is off, and is used to store data that needs to be used multiple times or for a long time;
[0034] The input module includes a zero adjustment circuit, a programmable gain amplifier, a high-precision ADC, an RS-485 transceiver, a photocoupler, and a Schmitt trigger, and is used to receive analog signals, SSI signals, and switch signals outside the controller;
[0035] The output module includes a high-precision DAC, a bus transceiver and a line driver, which are used to send analog signals and control quantities of switch signal types; the analog signal is used as a drive signal and a standby signal, the digital signal inside the controller is converted into an analog signal output through the DAC, and the switch signal is output with the help of a bus transceiver and a line driver;
[0036] The communication module includes JTAG, USB to UART chip, Ethernet physical layer transceiver and USB transceiver, which provide the signal path for the controller to interact with the outside world; JTAG is used to debug the controller; USB to UART chip is used to output serial port information; Ethernet physical layer transceiver and USB transceiver are used to provide general communication functions, and high-speed Ethernet communication provides conditions for distributed expansion between controllers;
[0037] The backup module includes an interface that directly connects to the programmable logic, which can directly transmit information to the Zynq on-chip system and provide a path for controller function expansion.
[0038] Furthermore, the Zynq system-on-chip in the main control module also combines the software ecosystem of ARM and the parallel processing calculation of FPGA. The processing system is consistent with the functions of a typical processor. The programmable logic can flexibly configure hardware resources and rely on real-time parallel pipeline computing capabilities to perform hardware acceleration on the control method, thereby reducing the control interval and improving the control accuracy.
[0039] Furthermore, the DC power interface in the power module supports large current and is used to connect to an external DC power supply; the filter circuit and the DC / DC regulator stabilize and reduce the voltage of the power input to the DC power interface, providing high-performance and stable power supply for each module, thereby ensuring the reliable operation of the controller.
[0040] Furthermore, the analog signal, SSI signal and switch signal received in the input module include: the external input analog signal passes through the zero adjustment circuit and the programmable gain amplifier and then enters the ADC for sampling to obtain a digital signal; the external input SSI signal is collected through the RS-485 transceiver; the external input switch signal is received through a combination of an optocoupler and a Schmitt trigger.
[0041] The beneficial effect of the present invention is that it provides an actuator control method and controller based on online neural prediction: the control method is based on hyperparameter adaptation, online system identification and neural predictive control, relies on neural networks to simplify the theoretical modeling of nonlinear systems and reduce the required state types, and adapts to time-varying system dynamics based on online learning and predictive control; the controller is designed based on the Zynq system-on-chip as the core architecture, and the control method can be hardware accelerated and implemented while ensuring ease of use, and the number of supported channels is increased through flexible input modules and output modules. In addition, it can be easily expanded by relying on diverse and high-speed communication methods in communication modules and spare modules; the combination of the provided control method and controller can improve the control performance and application scope of actuators in various industrial applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a physical diagram of an actuator system according to an embodiment of the present invention;
[0043] Figure 2 is a schematic diagram of an actuator according to an embodiment of the present invention;
[0044] Figure 3 is a flow chart of an actuator control method based on online neural prediction according to an embodiment of the present invention;
[0045] Figure 4 is a reference displacement waveform of an embodiment of the present invention;
[0046] Figure 5 It is a learning rate adaptation process of hyperparameter adaptation in an embodiment of the present invention;
[0047] Figure 6 It is the optimization rate adaptation process of the hyperparameter adaptation in the embodiment of the present invention;
[0048] Figure 7 is the neural network training error of the online system identification of the embodiment of the present invention;
[0049] Figure 8 is the optimal control quantity of the neural predictive control in the embodiment of the present invention;
[0050] Fig. 9 is the displacement tracking error of the online neural predictive control of an embodiment of the present invention;
[0051] Fig.10 is a structural diagram of a controller according to an embodiment of the present invention;
[0052] Fig.11 This is a physical diagram of a controller according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0054] Figure 1 is a physical diagram of the actuator system of this embodiment, which is mainly composed of an electro-hydraulic actuator, a displacement sensor and a load, wherein the electro-hydraulic actuator is a common type of actuator; Figure 2 is the principle diagram of the actuator of this embodiment, the driving signal acts on the servo valve, the servo valve controls the oil inlet, and finally changes the pressure difference of the hydraulic oil on the hydraulic cylinder valve core, thereby changing the displacement of the valve core; Figure 3 is a flow chart of the actuator control method based on online neural prediction of this embodiment, summarizing the steps indicated in the specification of the present invention; Figure 4 is the reference displacement waveform of this embodiment, the sampling rate is 1000 Hz, and the number of sampling points is 31096. An actuator control method based on online neural prediction is obtained through online hyperparameter adaptation, online system identification and neural predictive control. The specific steps are as follows:
[0055] Step 1: Hyperparameter adaptation is achieved by maximizing the cumulative reward, which includes the following sub-steps:
[0056] Step 1-1. Optional hyperparameter values for each hyperparameter Calculation control starts at the current time step k and takes optional hyperparameter values The sampling mean of the reward obtained when The hyperparameter is specifically a learning rate η or an optimization rate μ; it is necessary to define several optional hyperparameter values for each hyperparameter. Finite set of As an optional set of hyperparameters, in this embodiment, the learning rate η and the optimization rate μ are adaptively adjusted, and the optional set of η is Contains 3×10 -5 , 10 -4 , 3×10 -4 , 10 -3 , the optional set of μ Contains 10 -2 , 3×10 -2 , 10 -1 , 3×10 -1 At the same time, it is also necessary to manually define the form of reward. In this embodiment, the reward form of η is defined as the absolute error of neural network training |x d (k)-x p (k)| is uniformly distributed on [1,0], and the reward form of μ is defined as the absolute error of displacement tracking|r d (k+1)-x p (k+1)| is uniformly distributed on [1,0]. Represents η or μ, and both calculate the optimal value.
[0057] Step 1-2: Use the Hoefding inequality to estimate the upper bound of the reward distribution under the confidence interval of 1 in Select hyperparameters for the start to the current time step k In this embodiment, p=0.01;
[0058] Step 1-3: Select the upper bound of the confidence interval with the largest value The optimal hyperparameters for As the next hyperparameter value; in the hyperparameter adaptation of this embodiment, the adaptive processes of the learning rate η and the optimization rate μ are as follows: Figure 5 and Figure 6 As shown, the learning rate η continuously tries the optional set Finally, a fixed optimal learning rate is obtained at around 700ms, and the optimization rate μ continuously tries the optional set The value in , finally obtains a fixed optimal optimization rate around 4000ms.
[0059] Step 2: Online system identification uses a neural network to learn the time-varying electro-hydraulic actuator system dynamics during the control process, including the following sub-steps:
[0060] Step 2-1: Use a neural network with parameter θ(k-1) to predict the actuator valve core displacement value x at the current time step k. p (k), where the parameters include the weights and biases of each layer of the network; it is necessary to define the number of hidden layers N of the neural network h and the number of hidden layer neurons N n , the input of the neural network is [x d (k-1),…, d (km), c (k-1),…, c (kn)[ T , the output of the neural network is x p (k), where m and n are the orders of the displacement state value and the control quantity considered respectively; in this embodiment, let N h =1,N n =6, m=3, n=3;
[0061] Step 2-2, using x p (k) and the actual displacement value x measured by the sensor d (k) Obtain the objective function value J θ (k)=(x d (k)-x p (k)) 2 / 2;
[0062] Step 2-3: Find J by back propagation θ (k) Gradient of the neural network parameter θ(k-1)
[0063] Step 2-4: Gradient descent updates neural network parameters Where η is the learning rate. In this embodiment, the logarithmic form of the neural network training error identified by the online system is as follows: Figure 7 As shown, it is lower than -3dB, indicating that the neural network used can indeed learn the time-varying dynamics of the electro-hydraulic actuator system online.
[0064] Step 3: Neural predictive control obtains the optimal control input based on the neural network system model obtained through online identification, including the following sub-steps:
[0065] Step 3-1, set the initial iteration t = 0, the initial control amount u c (t|k)=0;
[0066] Step 3-2: Use a neural network with parameter θ(k) to predict the displacement value x at the next time step k+1 p (t|k+1); at this time the input of the neural network is [x d (k),…,x d (k-m+1),u c (t|k),…,u c (k-n+1)] T , the output of the neural network is x p (t|k+1).
[0067] Step 3-3, using x p (t|k+1) and the set reference value r d (k+1) Get the objective function value J u (t|k)=(r d (k+1)-x p (t|k+1)) 2 / 2;
[0068] Step 3-4: Find J by back propagation u (t|k) for u in the neural network input layer c The gradient of (t|k)
[0069] Step 3-5: Gradient descent to update the control amount Where μ is the optimization rate;
[0070] Step 3-6, execute t=t+1, if t<N g , go to step 3-2, otherwise go to step 3-7, where N gis the required number of iterations; in this embodiment, let N g =200;
[0071] Step 3-7, iteratively obtained u c (t|k) is the optimal control quantity u to be solved c (k); the optimal control quantity u c (k) is input to the electro-hydraulic actuator; in this embodiment, the optimal control quantity of the neural predictive control is as follows: Figure 8 As shown in the figure, the actuator displacement tracking error under online neural predictive control is Fig. 9 As shown, it can be seen that the overall error is kept at a low level, which confirms the effectiveness of the method.
[0072] Corresponding to the aforementioned embodiment of an actuator control method based on online neural prediction, the present invention also provides an embodiment of a controller for hardware acceleration and implementation of the control method.
[0073] like Fig.10 As shown, the system includes a main control module, a power module, a storage module, an input module, an output module, a communication module and a backup module;
[0074] The main control module includes the Zynq system-on-chip and its peripheral circuits. The Zynq system-on-chip is divided into a processing system and programmable logic, which is responsible for the coordination and logical operations between the various modules of the controller, and completes the interaction with the host computer and the hardware acceleration and implementation of the control method through software and hardware programming. Among them, the Zynq system-on-chip also combines the software ecology of ARM and the parallel processing and calculation of FPGA. The processing system is consistent with the functions of a typical processor, and the programmable logic can flexibly configure hardware resources, relying on real-time parallel pipeline computing capabilities to perform hardware acceleration on the control method, reducing the control interval and thus improving the control accuracy.
[0075] The Zynq system-on-chip in the main control module provides PS I / O and PLI / O for the processing system and programmable logic, respectively. The PS I / O integrates the function implementation internally, and is mainly connected to the PS-side DDR, SPI FLASH, SDIO of the storage module and the JTAG, USB to UART chip, Ethernet physical layer transceiver and USB transceiver of the communication module; PL I / O is divided into HP and HR: HP is a high-performance I / O supporting 1.2V-1.8V voltage, connected to the backup module, as well as the PL-side DDR and eMMC of the storage module; HR is a high-range I / O supporting 1.2V-3.3V voltage, connected to the control and data ports of the remaining modules. In addition, external oscillators are configured for the PS and PL ends respectively.
[0076] The power module includes a DC power interface, a filter circuit, and a DC / DC regulator, which are used to power the controller. The DC power interface supports large currents and is used to connect to an external DC power source. The filter circuit and DC / DC regulator stabilize and reduce the voltage of the power input from the DC power interface, providing high-performance and stable power for each module, ensuring the reliable operation of the controller.
[0077] The power module needs to generate +12V, -12V, +1.0V, +1.8V, +1.5V, +3.3V and +5.0V stable power with a maximum of 6A. The +12V is connected to the external power supply through a switch by the DC power interface and obtained through a filter circuit. Then the filtered stable +12V power is provided as an energy source to each DC / DC regulator to generate -12V, +1.0V, +1.8V, +1.5V, +3.3V and +5.0V power.
[0078] The storage module includes DDR, a storage medium whose data disappears when power is off, and is used to cache intermediate variables during operation; and also includes SDIO, SPI Flash, EEPROM and eMMC, a storage medium whose data does not disappear when power is off, and is used to store data that needs to be used multiple times or for a long time.
[0079] The storage module provides 1GB DDR for the processing system and programmable logic respectively, provides SDIO, 32MB SPI Flash and 128KB EEPROM with variable capacity for the processing system, and provides 8GB eMMC for the programmable logic.
[0080] The input module includes a zero adjustment circuit, a programmable gain amplifier, a high-precision ADC, an RS-485 transceiver, an optocoupler and a Schmitt trigger, and is used to receive analog signals, SSI signals and switch signals outside the controller. The received analog signals, SSI signals and switch signals include: the external input analog signal passes through the zero adjustment circuit and the programmable gain amplifier and then enters the ADC for sampling to obtain a digital signal; the external input SSI signal is collected through the RS-485 transceiver; the external input switch signal is received through a combination of an optocoupler and a Schmitt trigger.
[0081] The input module has 12 differential analog signal input channels and 4 single-ended analog signal input channels. The input analog signal can be calibrated by the zero adjustment circuit on the controller. The principle of the zero adjustment circuit is to connect a bias voltage to the input analog signal to correct the deviation of the bridge sensor, where the bias voltage comes from the output of the DAC and the resistor voltage divider. There are 6 SSI input channels. The RS-485 transceiver converts the single-ended clock output by the main control module into a differential form and sends it to the SSI sensor. The differential sensor signal is read and then converted into a single-ended form and sent back to the main control module for analysis and processing. There are 16 single-ended switch signal input channels and 4 differential switch signal input channels. The optocoupler performs optoelectronic isolation on the input switch signal to improve safety and anti-interference, and then uses a Schmitt trigger for waveform shaping.
[0082] The output module includes a high-precision DAC, a bus transceiver and a line driver, which are used to send analog signals and switch signal type control quantities; the analog signal is used as a drive signal and a backup signal, the digital signal inside the controller is converted into an analog signal output through the DAC, and the switch signal is output with the help of a bus transceiver and a line driver.
[0083] There are 24 single-ended analog signal output channels in the output module, 12 of which are used for zeroing circuits, and the remaining 12 can be output outside the controller for driving or standby. There are 16 single-ended switch signal output channels, and the bus transceiver can set the signal direction and input and output levels, and output the switch signal after level conversion. There are 4 differential switch signal output channels, and the line driver converts the single-ended switch signal output by the main control module into a differential form output.
[0084] The communication module includes JTAG, USB to UART chip, Ethernet physical layer transceiver and USB transceiver, which provide a signal path for the controller to interact with the outside world; JTAG is used to debug the controller; the USB to UART chip is used to output serial port information; the Ethernet physical layer transceiver and USB transceiver are used to provide general communication functions, and high-speed Ethernet communication provides conditions for distributed expansion between controllers.
[0085] The Ethernet physical layer transceiver of the communication module can communicate with the outside world at a high rate of 1000Mbps, upload data to or download data from the host computer. In addition, under the command of the main control module, it can also use the clock synchronization algorithm to complete the distributed high-precision synchronization expansion between controllers, further increasing the number of control channels; the USB transceiver can be configured as Host or Slave mode by software, and can be flexibly adjusted according to different application scenarios.
[0086] The backup module includes an interface that directly connects to the programmable logic, which can directly transmit information to the Zynq on-chip system and provide a path for controller function expansion.
[0087] The input module, output module, communication module, storage module and backup module are all equipped with physical interfaces. Fig.11 shown.
[0088] The working process in this embodiment is as follows: the host computer issues control instructions to the controller through the communication module, such as parameter setting, etc. The controller receives the sensor signal through the input module, reads the reference signal from the storage module, and after calculation by the main control module, uses the output module to output the drive signal to excite the actuator to control its movement. At the same time, the controller uploads the sensor signal received by the input module to the host computer through the communication module, and displays the control effect in real time on the UI interface of the host computer, which is convenient for users to monitor the control process. It should be pointed out that the host computer is not necessary in the control process, and is only used when parameter modification or process monitoring is required. When the number of channels of a single controller is not enough to meet the needs of motion control tasks, the main control module and the communication module are relied on to complete the interconnection between the controllers, and the clock synchronization algorithm is used to perform distributed high-synchronization expansion to further increase the number of input and output channels. The actuator control method based on online neural prediction is hardware accelerated and implemented on the programmable logic of the Zynq system-on-chip in the controller main control module through a hardware description language. Compared with the implementation on a traditional CPU, the control interval after hardware acceleration is reduced by more than 6 times, and a complete control cycle can be completed in a sufficiently short time to track the displacement reference waveform with a sampling rate of 1000Hz in this embodiment; communication with the host computer is implemented on the processing system of the Zynq system-on-chip in the controller main control module through C language and TCP / IP protocol.
[0089] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. An actuator control method based on online neural prediction, characterized in that: The method includes hyperparameter adaptation, online system identification and neural predictive control, and specifically includes the following steps: Step 1: Hyperparameter adaptation is achieved by maximizing the cumulative reward, which includes the following sub-steps: Step 1-1. Optional hyperparameter values for each hyperparameter Calculation control starts at the current time step k and takes optional hyperparameter values The sampling mean of the reward obtained when Step 1-2: Use the Hoefding inequality to estimate the upper bound of the reward distribution under the confidence interval of 1 in Select hyperparameters for controlling the start to the current time step k The number of times; Step 1-3: Select the upper bound of the confidence interval with the largest value The optimal hyperparameters for As the hyperparameter value of the current time step k; Step 2: Online system identification, using a neural network to learn the time-varying dynamics of the electro-hydraulic actuator system during the control process, includes the following sub-steps: Step 2-1: Use a fully connected neural network with parameter θ(k-1) to predict the actuator valve core displacement value x at the current time step k p (k), where the parameter θ(k-1) includes the weights and biases of each layer of the network; Step 2-2, using x p (k) and the actual displacement value x measured by the sensor d (k) Obtain the objective function value J θ (k)=(x d (k)-x p (k)) 2 / 2; Step 2-3: Find J by back propagation θ (k) Gradient of the neural network parameter θ(k-1) Step 2-4: Gradient descent updates neural network parameters Where η is the learning rate; Step 3: Neural predictive control, which obtains the optimal control input based on the neural network model obtained by online system identification, includes the following sub-steps: Step 3-1, set the initial iteration t = 0, the initial control amount u c (t|k)=0; Step 3-2: Use the neural network with parameter θ(k) to predict the actuator valve core displacement value x at the next time step k+1 p (t|k+1); Step 3-3, using x p (t|+1) and the set reference value r d (k+1) Get the objective function value J u (t|k)=(r d (k+1)-x p (t|k+1)) 2 / 2; Step 3-4: Find J by back propagation u (t|k) for u in the neural network input layer c The gradient of (t|) Step 3-5: Gradient descent to update the control amount Where μ is the optimization rate; Step 3-6, execute t=t+1, if t<N g , go to step 3-2, otherwise go to step 3-7, where N g is the required number of iterations; Step 3-7, iteratively obtained u c (t|k) is the optimal control quantity u to be solved c (k).
2. The actuator control method based on online neural prediction according to claim 1, characterized in that: The hyperparameter in step 1 is specifically the learning rate η or the optimization rate μ.
3. The actuator control method based on online neural prediction according to claim 1, characterized in that: The hyperparameters in step 1 require manual definition of the reward form.
4. The actuator control method based on online neural prediction according to claim 1, characterized in that: In step 1-1, it is necessary to define several optional hyperparameter values for each hyperparameter. Finite set of Optional set of hyperparameters.
5. The actuator control method based on online neural prediction according to claim 1, characterized in that: In step 2-1, it is necessary to define the number of hidden layers N of the neural network h and the number of hidden layer neurons N n , the input of the neural network is [x d (k-1),…,x d (km),u c (k-1),…,u c (kn)] T , the output of the neural network is x p (k), where m and n are the orders of the displacement state value and the control quantity considered, respectively.
6. The actuator control method based on online neural prediction according to claim 1, characterized in that: In step 3-2, the input of the neural network is [x d (k),…,x d (k-m+1),u c (t|k),u c (k-1),…,u c (k-n+1)] T , the output of the neural network is x p (t|k+1).
7. A controller for implementing the method according to any one of claims 1 to 6, characterized in that: The controller includes a main control module, a power module, a storage module, an input module, an output module, a communication module and a standby module; The main control module includes a Zynq system-on-chip and its peripheral circuits. The Zynq system-on-chip is divided into a processing system and a programmable logic, which is responsible for the coordination and logical operations between the various modules of the controller, and completes the interaction with the host computer and the hardware acceleration and implementation of the control method through software and hardware programming; The power module includes a DC power interface, a filter circuit and a DC / DC regulator, and is used to supply power to the controller; The storage module includes a storage medium DDR whose data disappears when power is off, which is used to cache intermediate variables during the operation of the algorithm; and also includes a storage medium SDIO, SPI Flash, EEPROM and eMMC whose data does not disappear when power is off, which is used to store data that needs to be used multiple times or for a long time; The input module includes a zero adjustment circuit, a programmable gain amplifier, a high-precision ADC, an RS-485 transceiver, a photocoupler and a Schmitt trigger, and is used to receive analog signals, SSI signals and switch signals outside the controller; The output module includes a high-precision DAC, a bus transceiver and a line driver, which are used to send analog signals and control quantities of switch signal types; the analog signal is used as a drive signal and a standby signal, the digital signal inside the controller is converted into an analog signal output through the DAC, and the switch signal is output with the help of the bus transceiver and the line driver; The communication module includes JTAG, USB to UART chip, Ethernet physical layer transceiver and USB transceiver, providing a signal path for the controller to interact with the outside world; JTAG is used to debug the controller; USB to UART chip is used to output serial port information; Ethernet physical layer transceiver and USB transceiver are used to provide general communication functions, and high-speed Ethernet communication provides conditions for distributed expansion between controllers; The backup module includes an interface directly connected to the programmable logic, which can directly transmit information with the Zynq system on chip, providing a path for controller function expansion.
8. The controller according to claim 7, characterized in that: The Zynq system-on-chip in the main control module also combines the software ecosystem of ARM and the parallel processing calculation of FPGA. The processing system is consistent with the functions of a typical processor. The programmable logic can flexibly configure hardware resources and rely on real-time parallel pipeline computing capabilities to perform hardware acceleration on the control method, reduce the control interval and thus improve the control accuracy.
9. The controller according to claim 7, characterized in that: The DC power interface in the power module supports large current and is used to connect to an external DC power supply; the filter circuit and the DC / DC regulator stabilize and reduce the voltage of the power input to the DC power interface, providing high-performance and stable power supply for each module, thereby ensuring the reliable operation of the controller.
10. The controller according to claim 7, characterized in that: The analog signal, SSI signal and switch signal received in the input module include: the external input analog signal enters the ADC for sampling to obtain a digital signal after passing through a zero adjustment circuit and a programmable gain amplifier; the external input SSI signal is collected through an RS-485 transceiver; and the external input switch signal is received through a combination of a photocoupler and a Schmitt trigger.