Simulation Turntable Control Method and Device Based on the Fusion of Data-Driven and Active Disturbance Rejection

Through the control method of data-driven and self-immunity fusion, the position controller, self-immunity controller and trained control network are used to solve the problem of the five-axis flight simulation turntable control relying on accurate mathematical model, and the stability and robustness are improved.

CN120233698BActive Publication Date: 2025-07-29TIANJIN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510703160.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-29
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the prior art, the control of the five-axis flight simulation turntable relies on accurate mathematical models, which makes it difficult to control and make it difficult to implement effective control methods.

Method used

Using a control method based on data-driven and self-immunity fusion, the position controller, self-immunity controller and pre-trained control network are used, and the compensation current is output to control the motion of the simulation turntable through the position controller, self-immunity controller and pre-trained control network, combining the behavior network and the evaluation network, thereby avoiding the dependence on the precise mathematical model.

Benefits of technology

It reduces the difficulty of implementing simulation turntable control, improves the stability and robustness of control, and realizes efficient control without the need for precise mathematical models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233698B_ABST
    Figure CN120233698B_ABST
Patent Text Reader

Abstract

The present invention provides a simulation turntable control method and device based on the fusion of data-driven and active disturbance rejection, which can be applied to the field of turntable control technology. The method includes: inputting the desired position at the (t + 1)-th moment and the actual position at the t-th moment of the simulation turntable into a position controller to output the desired speed at the (t + 1)-th moment; inputting the desired speed at the (t + 1)-th moment and the actual speed at the t-th moment of the simulation turntable into an active disturbance rejection controller to output the desired current at the (t + 1)-th moment; inputting the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network to output the compensation current at the (t + 1)-th moment; and controlling the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable. The present invention can reduce the control difficulty of the simulation turntable without relying on an accurate mathematical model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of turntable control, and more particularly, to a simulation turntable control method and device based on the fusion of data-driven and auto-disturbance rejection. Background Art

[0002] The five-axis flight simulation turntable is a simulation device for semi-physical simulation research and testing. The five-axis flight simulation turntable usually consists of a three-axis flight turntable, a two-axis target turntable, a simulation control platform, an inertial navigation system, a target simulator, product installation auxiliary equipment, etc. The three-axis flight turntable is used to simulate the attitude movement of the pod; the two-axis target turntable is equipped with visible light and infrared target sources to reproduce the movement trajectory of the target.

[0003] In the process of implementing the inventive concept, the inventors found that there are at least the following problems in the related art: the control of the five-axis flight simulation turntable in the related art requires relying on an accurate mathematical model, and it is difficult to create the mathematical model, which increases the control difficulty of the turntable and leads to a greater implementation difficulty of the control method for the five-axis flight simulation turntable. Summary of the Invention

[0004] In view of this, the present invention provides a simulation turntable control method and device based on the fusion of data-driven and auto-disturbance rejection.

[0005] One aspect of the present invention provides a simulation turntable control method based on the fusion of data-driven and auto-disturbance rejection, including:

[0006] Input the expected position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the above simulation turntable, and output the expected rotational speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0;

[0007] Input the expected rotational speed of the simulation turntable at the (t + 1)-th moment and the actual rotational speed at the t-th moment into the auto-disturbance rejection controller of the above simulation turntable, and output the expected current of the simulation turntable at the (t + 1)-th moment;

[0008] Input the actual rotational speed at the t-th moment, the expected rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the expected rotational speed at the (t - 1)-th moment of the above simulation turntable into a pre-trained control network, and output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network;

[0009] Control the movement of the above simulation turntable according to the expected current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the above simulation turntable.

[0010] According to an embodiment of the present invention, the above method further includes:

[0011] Based on the motion result of the above simulation turntable at the (t + 1)-th moment, obtain the actual position, actual rotation speed, and actual current of the above simulation turntable at the (t + 1)-th moment;

[0012] Input the desired position of the above simulation turntable at the (t + 2)-th moment and the actual position at the (t + 1)-th moment into the above position controller, and output the desired rotation speed of the above simulation turntable at the (t + 2)-th moment;

[0013] Input the desired rotation speed of the above simulation turntable at the (t + 2)-th moment and the actual rotation speed at the (t + 1)-th moment into the above active disturbance rejection controller, and output the desired current of the above simulation turntable at the (t + 2)-th moment;

[0014] Input the actual rotation speed at the (t + 1)-th moment, the desired rotation speed at the (t + 1)-th moment, the actual rotation speed at the t-th moment, and the desired rotation speed at the t-th moment of the above simulation turntable into the above pre-trained control network, and output the compensation current at the (t + 2)-th moment;

[0015] Control the motion of the above simulation turntable according to the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the above simulation turntable.

[0016] According to an embodiment of the present invention, the step of inputting the actual rotation speed at the t-th moment, the desired rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the desired rotation speed at the (t - 1)-th moment of the above simulation turntable into the pre-trained control network and outputting the compensation current at the (t + 1)-th moment of the above simulation turntable includes:

[0017] Input the actual rotation speed at the t-th moment, the desired rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the desired rotation speed at the (t - 1)-th moment of the above simulation turntable into the above behavior network, and output the initial compensation current at the (t + 1)-th moment;

[0018] Input the initial compensation current at the (t + 1)-th moment, the actual rotation speed at the t-th moment, the desired rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the desired rotation speed at the (t - 1)-th moment into the above evaluation network, and output the evaluation value at the (t + 1)-th moment;

[0019] Adjust the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.

[0020] According to an embodiment of the present invention, the above method further includes:

[0021] Initialize the initial behavior network and the initial evaluation network;

[0022] Input the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the above initial behavior network to output the compensation current at the (t + 1)-th historical moment;

[0023] Input the above compensation current at the (t + 1)-th historical moment, the above actual speed at the t-th historical moment, the above desired speed at the t-th historical moment, the above actual speed at the (t - 1)-th historical moment, and the above desired speed at the (t - 1)-th historical moment into the above initial evaluation network to output the historical evaluation value at the (t + 1)-th historical moment;

[0024] Update the parameters of the above initial evaluation network according to the above historical evaluation value at the (t + 1)-th historical moment, the above actual speed at the t-th historical moment, the above desired speed at the t-th historical moment, the above actual speed at the (t - 1)-th historical moment, and the above desired speed at the (t - 1)-th historical moment to obtain the above evaluation network;

[0025] Update the parameters of the above initial behavior network according to the above historical evaluation value at the (t + 1)-th historical moment to obtain the above behavior network.

[0026] According to an embodiment of the present invention, the above updating the parameters of the above initial evaluation network according to the above historical evaluation value at the (t + 1)-th historical moment, the above actual speed at the t-th historical moment, the above desired speed at the t-th historical moment, the above actual speed at the (t - 1)-th historical moment, and the above desired speed at the (t - 1)-th historical moment to obtain the above evaluation network includes:

[0027] Determine the loss function of the above initial evaluation network according to the historical evaluation value at the t-th historical moment, the reinforcement signal symmetric positive definite matrix, the above historical evaluation value at the (t + 1)-th historical moment, the above actual speed at the t-th historical moment, the above desired speed at the t-th historical moment, the above actual speed at the (t - 1)-th historical moment, and the above desired speed at the (t - 1)-th historical moment;

[0028] Update the weights of the above initial evaluation network according to the loss function of the above initial evaluation network to obtain the above evaluation network.

[0029] According to an embodiment of the present invention, the above updating the parameters of the above initial behavior network according to the above historical evaluation value at the (t + 1)-th historical moment to obtain the above behavior network includes:

[0030] Determine the loss function of the above initial behavior network according to the above historical evaluation value at the (t + 1)-th historical moment;

[0031] Update the weights of the above initial behavior network according to the loss function of the above initial behavior network to obtain the above behavior network.

[0032] According to an embodiment of the present invention, inputting the expected rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable into the auto-disturbance rejection controller of the simulation turntable, and outputting the expected current at the (t + 1)-th moment of the simulation turntable, includes:

[0033] According to the expected rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable, use the extended state observer of the auto-disturbance rejection controller to output the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;

[0034] According to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment, obtain the control voltage at the (t + 1)-th moment of the simulation turntable;

[0035] According to the control voltage at the (t + 1)-th moment, output the expected current at the (t + 1)-th moment of the simulation turntable.

[0036] According to an embodiment of the present invention, the obtaining of the control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment includes:

[0037] Output an intermediate control voltage at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;

[0038] Perform filtering processing on the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.

[0039] According to an embodiment of the present invention, the obtaining of the actual position at the (t + 1)-th moment, the actual rotational speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment includes:

[0040] Obtain the actual position at the (t + 1)-th moment and the actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment;

[0041] Differentiate the actual position at the (t + 1)-th moment to obtain the actual rotational speed at the (t + 1)-th moment.

[0042] Another aspect of the present invention provides a control device for a simulation turntable based on the fusion of data-driven and auto-disturbance rejection, including:

[0043] The first output module is configured to input the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable, and output the desired rotational speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0;

[0044] The second output module is configured to input the desired rotational speed of the simulation turntable at the (t + 1)-th moment and the actual rotational speed at the t-th moment into the active disturbance rejection controller of the simulation turntable, and output the desired current of the simulation turntable at the (t + 1)-th moment;

[0045] The third output module is configured to input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network;

[0046] The first control module is configured to control the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.

[0047] According to an embodiment of the present invention, through a pre-trained control network, the compensation current of the simulation turntable at the (t + 1)-th moment can be output based on data such as the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment, avoiding the establishment of a specific mathematical model according to the disturbance of the simulation turntable. Based on the output data at each moment, the control of the simulation turntable can be realized without the need to rely on an accurate mathematical model, reducing the implementation difficulty of the simulation turntable control method. Moreover, the active disturbance rejection controller is used to output the desired current at the (t + 1)-th moment, and the active disturbance rejection controller can further improve the stability and robustness of the control method. Description of the Drawings

[0048] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0049] Figure 1 Shows a flowchart of a control method for a simulation turntable based on the fusion of data-driven and active disturbance rejection according to an embodiment of the present invention;

[0050] Figure 2 Shows a flowchart of a training method for a control network according to an embodiment of the present invention;

[0051] Figure 3 Shows a schematic diagram of the network architecture of a control network according to an embodiment of the present invention;

[0052] Figure 4 shows the control schematic diagram of the active disturbance rejection controller according to an embodiment of the present invention;

[0053] Figure 5 shows the schematic diagram of the control architecture of the simulation turntable according to an embodiment of the present invention;

[0054] Figure 6A shows the schematic diagram of the update of the behavior network weights in the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention;

[0055] Figure 6B shows the schematic diagram of the update of the evaluation network weights in the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention;

[0056] Figure 7 shows the schematic diagram of the comparison of the simulation effects between the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention and the control method in the related art;

[0057] Figure 8 shows the block diagram of the simulation turntable control device based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention. Detailed implementation manners

[0058] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, obviously, one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0059] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0060] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0061] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0062] In the embodiments of the present invention, in aspects such as the collection, update, analysis, processing, use, transmission, provision, disclosure, and storage of the involved data (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs.

[0063] In the related art, the main means to evaluate the performance of the guidance system and navigation equipment components is hardware-in-the-loop simulation. The method of hardware-in-the-loop simulation is much lower in cost and risk than actual physical tests because it can be tested without building a complete hardware system. At the same time, compared with real-time simulation that completely relies on computer models, hardware-in-the-loop simulation can also provide results closer to the actual system behavior.

[0064] During the casting or welding process of the structural components in the five-axis flight simulation turntable, there will be a certain degree of difference from the design drawings, resulting in multiple resonance points. Moreover, according to the different positions between the axes, there will also be resonance points with different frequencies. During the movement of the five-axis flight simulation turntable, it will also be affected by multi-source disturbances such as multi-axis coupling, cogging torque fluctuation, friction torque fluctuation, and electromagnetic interference, making it difficult to establish an accurate mathematical model and having a greater control difficulty.

[0065] In view of this, the embodiments of the present invention provide a simulation turntable control method based on the fusion of data-driven and active disturbance rejection, including: inputting the expected position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable to output the expected speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0; inputting the expected speed of the simulation turntable at the (t + 1)-th moment and the actual speed at the t-th moment into the active disturbance rejection controller of the simulation turntable to output the expected current of the simulation turntable at the (t + 1)-th moment; inputting the actual speed of the simulation turntable at the t-th moment, the expected speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the expected speed at the (t - 1)-th moment into a pre-trained control network to output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network; controlling the movement of the simulation turntable according to the expected current of the simulation turntable at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment.

[0066] Figure 1The flowchart of the simulation turntable control method based on the fusion of data-driven and active disturbance rejection according to an embodiment of the present invention is shown.

[0067] As Figure 1 shown, the method includes operations S110 to S140.

[0068] In operation S110, the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment are input into the position controller of the simulation turntable, and the desired speed of the simulation turntable at the (t + 1)-th moment is output, where t is an integer greater than 0.

[0069] In operation S120, the desired speed of the simulation turntable at the (t + 1)-th moment and the actual speed at the t-th moment are input into the active disturbance rejection controller of the simulation turntable, and the desired current of the simulation turntable at the (t + 1)-th moment is output.

[0070] In operation S130, the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable are input into a pre-trained control network, and the compensation current of the simulation turntable at the (t + 1)-th moment is output. The control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network.

[0071] In operation S140, the simulation turntable is controlled to move according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.

[0072] According to an embodiment of the present invention, the simulation turntable proposed by the present invention may be a two-axis target turntable in a five-axis flight simulation turntable, and the moment proposed by the present invention may be a servo control period of the simulation turntable.

[0073] According to an embodiment of the present invention, the desired position of the simulation turntable at the (t + 1)-th moment may be the final position of the simulation turntable input from the outside, or may be an intermediate position decomposed according to the final position of the simulation turntable input from the outside. The desired position at the (t + 1)-th moment may be the position where it is desired for the simulation turntable to be located at the end of the (t + 1)-th moment.

[0074] According to an embodiment of the present invention, the actual position of the simulation turntable at the t-th moment may be the position of the simulation turntable at the end of the t-th moment, that is, the position of the simulation turntable at the beginning of the (t + 1)-th moment.

[0075] According to an embodiment of the present invention, the position controller may calculate the desired position at the (t + 1)-th moment and the actual position at the t-th moment to obtain the desired speed of the simulation turntable at the (t + 1)-th moment.

[0076] According to an embodiment of the present invention, the actual rotational speed of the simulation turntable at the t-th moment can be the rotational speed of the simulation turntable at the end of the t-th moment, that is, the rotational speed of the simulation turntable at the start of the (t + 1)-th moment.

[0077] According to an embodiment of the present invention, the desired rotational speed at the (t + 1)-th moment can be the desired speed of the simulation turntable at the end of the (t + 1)-th moment. According to an embodiment of the present invention, based on the desired rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment, the desired current of the simulation turntable at the (t + 1)-th moment can be output through an active disturbance rejection controller, such that the simulation turntable rotates according to the desired current at the (t + 1)-th moment, and the speed at the end of the (t + 1)-th moment can be the desired rotational speed at the (t + 1)-th moment.

[0078] According to an embodiment of the present invention, the active disturbance rejection controller can use the active disturbance rejection control (ADRC) method for output. The ADRC method does not rely on an accurate mathematical model, treats all modeling errors as disturbances, and only requires a rough process model to design a control loop. Its linear form is equivalent to a special case of classical state - space control based on the internal - model principle, with disturbance estimation and compensation, and has high robustness and anti - disturbance performance.

[0079] According to an embodiment of the present invention, the control network can be a network pre - trained for compensating the desired current. The control network includes a behavior network and an evaluation network. The behavior network can output a specific compensation current, and the evaluation network can evaluate the compensation current output by the behavior network to determine whether the output result of the behavior network is optimal.

[0080] According to an embodiment of the present invention, based on the actual rotational speed and the desired rotational speed at the t-th moment, the rotational speed error at the t-th moment can be determined. Based on the actual rotational speed and the desired rotational speed at the (t - 1)-th moment, the rotational speed error at the (t - 1)-th moment can be determined. Inputting the rotational speed error at the t-th moment and the rotational speed error at the (t - 1)-th moment into the pre - trained control network can output the compensation current at the (t + 1)-th moment. That is, the control network can output the compensation current for the next moment based on the rotational speed errors of the two adjacent previous moments.

[0081] According to an embodiment of the present invention, based on the desired current and the compensation current at the (t + 1)-th moment of the simulation turntable, the control current at the (t + 1)-th moment can be obtained. The control current at the (t + 1)-th moment can be the current used to control the simulation turntable at the (t + 1)-th moment. Based on the control current at the (t + 1)-th moment and the actual current at the t-th moment, the movement of the simulation turntable is controlled.

[0082] According to an embodiment of the present invention, through a pre-trained control network, the compensation current at the (t + 1)-th moment of the simulation turntable can be output based on data such as the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment. This avoids establishing a specific mathematical model based on the disturbance of the simulation turntable, reducing the implementation difficulty of the simulation turntable control method. Moreover, the auto-disturbance rejection controller is used to output the desired current at the (t + 1)-th moment, and the auto-disturbance rejection controller can further improve the stability and robustness of the control method.

[0083] According to an embodiment of the present invention, the method may further include: obtaining the actual position, actual speed, and actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment; inputting the desired position at the (t + 2)-th moment and the actual position at the (t + 1)-th moment of the simulation turntable into a position controller to output the desired speed at the (t + 2)-th moment of the simulation turntable; inputting the desired speed at the (t + 2)-th moment and the actual speed at the (t + 1)-th moment of the simulation turntable into an auto-disturbance rejection controller to output the desired current at the (t + 2)-th moment of the simulation turntable; inputting the actual speed at the (t + 1)-th moment, the desired speed at the (t + 1)-th moment, the actual speed at the t-th moment, and the desired speed at the t-th moment of the simulation turntable into a pre-trained control network to output the compensation current at the (t + 2)-th moment; and controlling the motion of the simulation turntable according to the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment.

[0084] According to an embodiment of the present invention, after the simulation turntable moves at the (t + 1)-th moment, the position, speed, and current of the simulation turntable can be detected to obtain the actual position, actual speed, and actual current at the (t + 1)-th moment.

[0085] According to an embodiment of the present invention, the motion control of the simulation turntable at each moment will refer to the actual position, actual speed, and actual current at the previous moment, and adjust the output of the simulation turntable according to the difference between the actual position and the desired position, the difference between the actual speed and the desired speed, and the difference between the actual current and the desired current.

[0086] According to an embodiment of the present invention, the motion of the simulation turntable at the (t + 2)-th moment can refer to the description at the (t + 1)-th moment, which will not be elaborated here.

[0087] According to an embodiment of the present invention, inputting the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network to output the compensation current at the (t + 1)-th moment of the simulation turntable, including: inputting the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a behavior network to output the initial compensation current at the (t + 1)-th moment; inputting the initial compensation current at the (t + 1)-th moment, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment into an evaluation network to output the evaluation value at the (t + 1)-th moment; and adjusting the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.

[0088] According to an embodiment of the present invention, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment can be input into a behavior network, and the behavior network can output the initial compensation current at the (t + 1)-th moment. Since it is uncertain whether the initial compensation current at the (t + 1)-th moment meets the requirements, the initial compensation current at the (t + 1)-th moment, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment can be input into an evaluation network to output the evaluation value at the (t + 1)-th moment for the initial compensation current at the (t + 1)-th moment. When the evaluation value at the (t + 1)-th moment indicates that the initial compensation current at the (t + 1)-th moment can make the position of the simulation turntable at the end of the (t + 1)-th moment the same as the desired position at the (t + 1)-th moment, the initial compensation current at the (t + 1)-th moment can be output as the compensation current at the (t + 1)-th moment. When the evaluation value at the (t + 1)-th moment cannot make the position of the simulation turntable at the end of the (t + 1)-th moment the same as the desired position at the (t + 1)-th moment, the initial compensation current at the (t + 1)-th moment can be adjusted to obtain the compensation current at the (t + 1)-th moment. During the process of adjusting the initial compensation current at the (t + 1)-th moment, multiple intermediate compensation currents at the (t + 1)-th moment can be obtained, and the intermediate compensation currents at the (t + 1)-th moment can be input into the evaluation network to obtain the corresponding evaluation values, and the compensation current at the (t + 1)-th moment can be obtained according to the evaluation values.

[0089] Figure 2 The flowchart of the training method of the control network according to an embodiment of the present invention is shown.

[0090] As Figure 2 shown, the method includes operations S210 to S250.

[0091] In operation S210, initialize the initial behavior network and the initial evaluation network.

[0092] In operation S220, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment are input into the initial behavior network, and the compensation current at the (t + 1)-th historical moment is output.

[0093] In operation S230, the compensation current at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment are input into the initial evaluation network, and the historical evaluation value at the (t + 1)-th historical moment is output.

[0094] In operation S240, based on the historical evaluation value at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment, the parameters of the initial evaluation network are updated to obtain the evaluation network.

[0095] In operation S250, based on the historical evaluation value at the (t + 1)-th historical moment, the parameters of the initial behavior network are updated to obtain the behavior network.

[0096] According to an embodiment of the present invention, the control network can be trained based on the historical action data of the simulation turntable.

[0097] Figure 3 FIG. shows a schematic diagram of the network architecture of the control network according to an embodiment of the present invention.

[0098] Figure 3 FIG. a schematically shows the network architecture of the initial behavior network. The initial behavior network may include an input layer, a hidden layer, and an output layer. The input layer of the initial behavior network is used to input the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment. The hidden layer of the initial behavior network obtains the features based on the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment. The output layer of the initial behavior network is used to output the compensation current at the (t + 1)-th historical moment according to the features in the hidden layer. Among them, represents the rotational speed error at the historical moment input into the initial behavior network, represents the weight from the input layer to the hidden layer of the initial behavior network, represents the weight from the hidden layer to the output layer of the initial behavior network, represents the compensation current at the historical moment of the output of the initial behavior network.

[0099] Figure 3Figure b shows the network architecture of the initial evaluation network. The initial evaluation network may include an input layer, a hidden layer, and an output layer. The input layer of the initial evaluation network is used to input the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment. The hidden layer of the initial evaluation network obtains features based on the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment. The output layer of the initial evaluation network is used to output the historical evaluation value at the (t + 1)-th historical moment according to the features in the hidden layer. Among them, represents the weight from the input layer to the hidden layer of the initial evaluation network, represents the weight from the hidden layer to the output layer of the initial evaluation network, represents the historical evaluation value at the historical moment output by the initial evaluation network.

[0100] According to an embodiment of the present invention, according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment, the parameters of the initial evaluation network can be updated to obtain an evaluation network. According to the historical evaluation value at the (t + 1)-th historical moment, the parameters of the initial behavior network can be updated to obtain a behavior network.

[0101] The design of the behavior network is as follows:

[0102] (1)

[0103] (2)

[0104] (3)

[0105] (4)

[0106] Among them, , , represent the intermediate variables output by the initial behavior network at the t-th historical moment, represents the number of nodes in the input layer of the initial behavior network, i represents the i-th node in the input layer of the initial behavior network, represents the number of nodes in the hidden layer of the initial behavior network, j represents the j-th node in the hidden layer of the initial behavior network, represents the base of the natural logarithm, represents the output weight from the i-th node in the input layer to the j-th node in the hidden layer of the initial behavior network at the t-th historical moment, denotes the output weight of the j-th node in the hidden layer of the initial behavior network at the t-th historical moment, , which represents the combination of the rotational speed error at the t-th historical moment and the rotational speed error at the (t - 1)-th historical moment, where, represents the error between the actual rotational speed and the desired rotational speed at the t-th historical moment, represents the error between the actual rotational speed and the desired rotational speed at the (t - 1)-th historical moment, represents the compensated current at the (t + 1)-th historical moment output by the initial behavior network.

[0107] The input of the evaluation network is , which represents the combination of the rotational speed error at the t-th historical moment, the rotational speed error at the (t - 1)-th historical moment, and the compensated current at the (t + 1)-th historical moment. The design of the evaluation network is as follows:

[0108] (5)

[0109] (6)

[0110] (7)

[0111] Wherein, , represents the intermediate variable output by the initial evaluation network at the t-th historical moment, represents the number of nodes in the input layer of the initial evaluation network, represents the number of nodes in the hidden layer of the initial evaluation network, represents the output weight from the i-th node in the input layer to the j-th node in the hidden layer of the initial evaluation network at the t-th historical moment, represents the output weight of the j-th node in the hidden layer of the initial evaluation network at the t-th historical moment, represents the historical evaluation value at the (t + 1)-th historical moment output by the initial evaluation network.

[0112] According to the embodiments of the present invention, based on the historical evaluation value at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment, the parameters of the initial evaluation network are updated to obtain an evaluation network, including: determining the loss function of the initial evaluation network according to the historical evaluation value at the t-th historical moment, the positive definite matrix of the reinforcement signal symmetry, the historical evaluation value at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment; updating the weights of the initial evaluation network according to the loss function of the initial evaluation network to obtain an evaluation network.

[0113] According to an embodiment of the present invention, the historical evaluation value at the t-th historical moment can be obtained by referring to the historical evaluation value at the (t + 1)-th historical moment, which will not be elaborated here.

[0114] According to an embodiment of the present invention, the reinforcement signal symmetric positive definite matrix can be a matrix that is both a symmetric matrix and a positive definite matrix to ensure the stability of the update of the initial evaluation network.

[0115] According to an embodiment of the present invention, the loss function of the initial evaluation network is as follows:

[0116] (8)

[0117] where represents the reinforcement signal at the t-th historical moment of the initial evaluation network, is the reinforcement signal symmetric positive definite matrix, represents the transpose of, is the loss factor, represents the historical evaluation value at the (t + 1)-th historical moment output by the initial evaluation network, represents the historical evaluation value at the t-th historical moment output by the initial evaluation network.

[0118] According to the gradient descent method, the weight update of the initial evaluation network is as follows:

[0119] (9)

[0120] (10)

[0121] where represents the training learning rate of the initial evaluation network, .

[0122] According to an embodiment of the present invention, according to the historical evaluation value at the (t + 1)-th historical moment, the parameters of the initial behavior network are updated to obtain the behavior network, including: determining the loss function of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment; updating the weights of the initial behavior network according to the loss function of the initial behavior network to obtain the behavior network.

[0123] The loss function of the initial behavior network is:

[0124] (11)

[0125] where represents the square of the historical evaluation value at the (t + 1)-th historical moment output by the initial evaluation network .

[0126] The weights of the initial behavior network are updated as follows:

[0127] (12)

[0128] (13)

[0129] wherein, represents the training learning rate of the initial behavior network, .

[0130] According to the embodiments of the present invention, the weights of the initial evaluation network and the initial behavior network can be iteratively updated, and the iteration is stopped when the number of iterations reaches a threshold or the initial evaluation network and the initial behavior network meet the requirements, so as to obtain a behavior network and an evaluation network, that is, a control network. Based on data such as the rotation speed and position of the simulation turntable, the control network can output a compensation current, without establishing a corresponding disturbance model for the simulation turntable, reducing the control difficulty of the simulation turntable.

[0131] Figure 4 shows the control schematic diagram of the active disturbance rejection controller according to the embodiments of the present invention.

[0132] As Figure 4 shown, the active disturbance rejection controller combines the classical proportional-integral-derivative (PID) control method and has ease of use. The input of the active disturbance rejection controller is the desired rotation speed at the current moment and the actual rotation speed at the previous moment, and the output control voltage is . After low-pass filtering and notch filtering, the desired current at the current moment can be obtained, and then the desired current at the current moment is sent to the driver through analog quantity to control the action of the simulation turntable. The active disturbance rejection controller includes an extended state observer and a PID controller. The outputs of the extended state observer are , and . represents the proportional coefficient of the PID controller, represents the differential coefficient of the PID controller, represents the critical gain in the active disturbance rejection controller.

[0133] According to the embodiments of the present invention, represents the observed speed of the simulation turntable, represents the observed acceleration of the simulation turntable, represents the observed total disturbance of the simulation turntable.

[0134] Let , the output of the extended state observer can be obtained as follows:

[0135] (14)

[0136] (15)

[0137] where, represents the estimation of the extended state observer at the th moment obtained by using the estimation at the th moment, represents the control voltage output by the active disturbance rejection controller at the th moment, represents the output of the extended state observer at the th moment obtained by using the estimation at the th moment, represents the output at the th moment obtained by using the estimation at the th moment, represents the actual rotational speed at the th moment.

[0138] where, , , , are the gains of the extended state observer, represents the servo cycle of the simulation turntable, represents the critical gain. , and represent the three-dimensional components of the gain of the extended state observer, , and are expressed as follows:

[0139] (16)

[0140] (17)

[0141] (18)

[0142] where, is the bandwidth of the extended state observer.

[0143] The output of the active disturbance rejection controller is:

[0144] (19)

[0145] (20)

[0146] (21)

[0147] Wherein, represents the control voltage output by the active disturbance rejection controller at the th moment, represents the desired rotational speed at the th moment, represents the observed speed of the simulation turntable at the th moment, represents the observed acceleration of the simulation turntable at the th moment, represents the observed total disturbance of the simulation turntable at the th moment, is the bandwidth of the active disturbance rejection controller. The control voltage output by the active disturbance rejection controller at the th moment can output the desired current at the th moment after low-pass filtering and notch filtering.

[0148] Specifically, in the embodiment of the present invention, according to the desired rotational speed at the (t + 1)th moment and the actual rotational speed at the tth moment of the simulation turntable, the extended state observer of the active disturbance rejection controller can output the observed speed at the (t + 1)th moment, the observed acceleration at the (t + 1)th moment, and the observed total disturbance at the (t + 1)th moment; according to the observed speed at the (t + 1)th moment, the observed acceleration at the (t + 1)th moment, and the observed total disturbance at the (t + 1)th moment, the control voltage at the (t + 1)th moment of the simulation turntable can be calculated; according to the control voltage at the (t + 1)th moment, after low-pass filtering and notch filtering, the desired current at the (t + 1)th moment of the simulation turntable is output.

[0149] According to the embodiment of the present invention, obtaining the control voltage at the (t + 1)th moment of the simulation turntable according to the observed speed at the (t + 1)th moment, the observed acceleration at the (t + 1)th moment, and the observed total disturbance at the (t + 1)th moment may include: outputting the intermediate control voltage at the (t + 1)th moment according to the observed speed at the (t + 1)th moment, the observed acceleration at the (t + 1)th moment, and the observed total disturbance at the (t + 1)th moment; performing filtering processing on the intermediate control voltage at the (t + 1)th moment to obtain the control voltage at the (t + 1)th moment.

[0150] According to the embodiment of the present invention, the filtering processing may include low-pass filtering and notch filtering. By performing filtering processing on the intermediate control voltage at the (t + 1)th moment, the interference signal in the intermediate control voltage at the (t + 1)th moment can be filtered out to obtain a more stable and accurate control voltage at the (t + 1)th moment.

[0151] According to an embodiment of the present invention, the ADRC method combines the ease of use of the classical PID control method, does not rely on an accurate mathematical model, treats all modeling errors as disturbances, and only requires a very rough process model to design a control loop. Its linear form is equivalent to a special case of classical state - space control based on the internal - model principle, has disturbance estimation and compensation, and has extremely strong robustness and anti - disturbance performance.

[0152] According to an embodiment of the present invention, based on the motion result of the simulation turntable at the (t + 1) - th moment, the actual position, the actual rotation speed, and the actual current of the simulation turntable at the (t + 1) - th moment are obtained, including: based on the motion result of the simulation turntable at the (t + 1) - th moment, the actual position at the (t + 1) - th moment and the actual current of the simulation turntable at the (t + 1) - th moment are obtained; differentiating the actual position at the (t + 1) - th moment to obtain the actual rotation speed at the (t + 1) - th moment.

[0153] According to an embodiment of the present invention, the simulation turntable can be driven by a DC brushless torque motor and is equipped with a high - precision absolute encoder. The absolute encoder can obtain the actual position at the (t + 1) - th moment through information fusion technology, and differentiating the actual position at the (t + 1) - th moment to obtain the actual rotation speed at the (t + 1) - th moment.

[0154] Figure 5 The schematic diagram of the control architecture of the simulation turntable according to an embodiment of the present invention is shown.

[0155] As Figure 5 shown, the position controller can receive the actual position of the simulation turntable in real - time. According to the input desired position and the actual position, it outputs the desired rotation speed to the active disturbance rejection controller (ADRC). The active disturbance rejection controller can receive the actual rotation speed of the simulation turntable in real - time, and outputs the desired current through the active disturbance rejection controller and low - pass and notch algorithms. The actual rotation speed of the simulation turntable is also input into the control network. As known from the foregoing, the control network is a pre - trained network, and the utility function is a parameter related to the convergence condition of the control network set during the training of the control network. The convergence condition of the control network is that the error between the actual rotation speed and the desired rotation speed at the current moment is equal to zero. Therefore, during the training of the control network in the present invention, the utility function can be set to 0. The control network can save the rotation speed errors of the previous two moments, and input the rotation speed errors of the previous two moments and the compensation current at the (t + 1) - th moment output by the behavior network into the evaluation network. The evaluation network outputs the evaluation value at the (t + 1) - th moment, and feeds it back to through the loss factor and the reinforcement signal . The behavior network outputs a compensation current according to the final , and the final input current is obtained from the desired current and the compensation current . The input current is the direct-axis current, represents the quadrature-axis current. According to and the actual current , the motion of the simulation turntable is controlled. Among them, the driver of the simulation turntable can be equivalent to the proportional-integral controller (PI), inverse Park transformation, space vector modulation (SVPWM), and three-phase inverter in the figure. The inverse Park transformation is used to convert the voltages and in the dq coordinate system into the voltages and in the αβ coordinate system. The SVPWM is used to utilize the DC bus voltage to improve the output capacity of the inverter. The three-phase inverter is used to convert direct current into sinusoidal alternating current. Finally, the controller outputs the voltage to control the motion of the brushless DC motor (BLDC). At the same time, the BLDC outputs the actual current to the Clark transformation. The Clark transformation converts into the currents and in the αβ coordinate system. The Park transformation converts the currents and in the αβ coordinate system into the currents and in the dq coordinate system and feeds them back to the driver. During the motion of the BLDC, the interference term can be regarded as multi-source interference and is reflected in the frame axis. The absolute encoder obtains the real-time position through the frame axis, and after information fusion, it outputs the actual position to the inverse Park transformation, Park transformation, and position controller, and outputs the actual speed to the ADRC through differentiation. Figure 5 . The corresponding moments of the desired position, actual position, actual speed, etc. in

[0156] can be understood in combination with other embodiments of the present invention, and will not be described in detail here. Figure 5 As can be seen from , in the control architecture of the simulation turntable, a closed-loop control of the current loop, speed loop, and position loop is formed through current feedback, speed feedback, and position feedback. Before the simulation turntable operates, it is necessary to adjust the parameters in each loop. The parameters required for the current loop can be adjusted first, and then the parameters in the speed loop and position loop can be adjusted in turn. The parameters in the speed loop can be adjusted by using the frequency-domain sweep method to debug the parameters of the low-pass filter and notch filter, and then adjust the critical gain and the bandwidth of the extended state observer The parameters of the position loop can also be adjusted using the amplitude-frequency and phase-frequency characteristic curves to obtain a control system with stable operation, and then the control network is added.

[0157] The control method of the embodiment of the present invention is subjected to simulation analysis, and the resistance of the DC brushless motor is designed , the inductance is 5.25 mH, the number of pole pairs is 4, the magnetic flux is 0.187 Wb, and the moment of inertia is , the parameters of the speed loop: , , . The parameters of the control network: , , , , , , , , different rotational speeds of the simulation turntable are selected, and the weight updates of the behavior network and the evaluation network are obtained as shown in FIG. 6.

[0158] Figure 6A FIG. shows a schematic diagram of the weight update of the behavior network in the simulation analysis of the simulation turntable control method based on data-driven and auto-disturbance rejection fusion according to an embodiment of the present invention.

[0159] Figure 6B FIG. shows a schematic diagram of the weight update of the evaluation network in the simulation analysis of the simulation turntable control method based on data-driven and auto-disturbance rejection fusion according to an embodiment of the present invention.

[0160] As Figure 6A shown, the weights of the behavior network in different nodes of the control network can be obtained. From top to bottom, they respectively represent the weights of the behavior network at the 1st node of the input layer to the 1st, 2nd, 3rd, 4th, 5th, and 6th nodes of the hidden layer. Among them, the abscissa represents the moment of weight update, and the ordinate represents the numerical change of the weight. As Figure 6B shown, the weights of the evaluation network in different nodes of the control network can be obtained. From top to bottom, they respectively represent the weights of the evaluation network at the 1st node of the input layer to the 1st, 2nd, 3rd, 4th, 5th, and 6th nodes of the hidden layer. Among them, the abscissa represents the moment of weight update, and the ordinate represents the numerical change of the weight. It should be noted that Figure 6A and Figure 6B only show partial node weight update processes.

[0161] Figure 7 FIG. shows a schematic diagram of the comparison of the simulation effects of the simulation analysis of the simulation turntable control method based on data-driven and auto-disturbance rejection fusion according to an embodiment of the present invention and the control method in the related art.

[0162] AsFigure 7 As shown, the abscissa represents the running time of the simulation turntable, and the ordinate represents the speed of the simulation turntable. The input of the simulation turntable is the rotational speed with multi-source disturbances, and a step disturbance signal is added at 1.5 s. Among them, the straight line (ADRC) is the simulation effect without adding the control network, and the dashed line (ADRC-Aux) represents the simulation effect after adding the control network. It can be seen that the speed of the simulation turntable after adding the control network is more stable, and there are significant improvements in both the anti-disturbance ability and the rate smoothness.

[0163] According to an embodiment of the present invention, the control method proposed by the present invention is applied to a five-axis flight simulation turntable, and the rate smoothness is detected to obtain the rate smoothness of the simulation turntable at 15 ms intervals, as shown in Table 1 and Table 2 below. Among them, Table 1 is the inspection record form of the pitch axis rate smoothness of the simulation turntable, and Table 2 is the inspection record form of the azimuth axis rate smoothness of the simulation turntable. The horizontal header represents the input desired rate command, and the vertical header represents the actual angle turned by the pitch axis or azimuth axis of the simulation turntable when running at the desired rate every 15 ms interval.

[0164] Table 1

[0165]

[0166] Table 2

[0167]

[0168] From the actual test results in the above table, it can be concluded that after the simulation turntable is added with the control method of the simulation turntable based on data-driven and auto-disturbance rejection fusion, its rate smoothness can reach (15 ms fixed time interval), and there is an obvious improvement in the performance of the rate smoothness compared with the conventional turntable.

[0169] Figure 8 The block diagram of the simulation turntable control device based on data-driven and auto-disturbance rejection fusion according to an embodiment of the present invention is shown.

[0170] As Figure 8 shown, the device 800 includes a first output module 810, a second output module 820, a third output module 830, and a first control module 840.

[0171] The first output module 810 is configured to input the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable, and output the desired rotational speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0;

[0172] The second output module 820 is configured to input the desired rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable into the auto-disturbance rejection controller of the simulation turntable, and output the desired current at the (t + 1)-th moment of the simulation turntable.

[0173] The third output module 830 is configured to input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 1)-th moment of the simulation turntable, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network.

[0174] The first control module 840 is configured to control the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.

[0175] According to an embodiment of the present invention, the apparatus 800 further includes:

[0176] The first obtaining module is configured to obtain the actual position at the (t + 1)-th moment, the actual rotational speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the movement result of the simulation turntable at the (t + 1)-th moment.

[0177] The fourth output module is configured to input the desired position at the (t + 2)-th moment and the actual position at the (t + 1)-th moment of the simulation turntable into a position controller, and output the desired rotational speed at the (t + 2)-th moment of the simulation turntable.

[0178] The fifth output module is configured to input the desired rotational speed at the (t + 2)-th moment and the actual rotational speed at the (t + 1)-th moment of the simulation turntable into the auto-disturbance rejection controller, and output the desired current at the (t + 2)-th moment of the simulation turntable.

[0179] The sixth output module is configured to input the actual rotational speed at the (t + 1)-th moment, the desired rotational speed at the (t + 1)-th moment, the actual rotational speed at the t-th moment, and the desired rotational speed at the t-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 2)-th moment.

[0180] The second control module is configured to control the movement of the simulation turntable according to the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable.

[0181] According to an embodiment of the present invention, the third output module 830 configured to input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 1)-th moment of the simulation turntable includes:

[0182] The first output unit is configured to input the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into the behavior network, and output the initial compensation current at the (t + 1)-th moment;

[0183] The second output unit is configured to input the initial compensation current at the (t + 1)-th moment, the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment into the evaluation network, and output the evaluation value at the (t + 1)-th moment;

[0184] The third output unit is configured to adjust the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.

[0185] According to an embodiment of the present invention, the device 800 further includes:

[0186] An initialization module, configured to initialize the initial behavior network and the initial evaluation network;

[0187] The seventh output module is configured to input the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial behavior network, and output the compensation current at the (t + 1)-th historical moment;

[0188] The eighth output module is configured to input the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial evaluation network, and output the historical evaluation value at the (t + 1)-th historical moment;

[0189] The second obtaining module is configured to update the parameters of the initial evaluation network according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network;

[0190] The third obtaining module is configured to update the parameters of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment to obtain the behavior network.

[0191] According to an embodiment of the present invention, the second obtaining module for updating the parameters of the initial evaluation network according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network includes:

[0192] A first obtaining unit, configured to determine a loss function of an initial evaluation network according to a historical evaluation value at the t-th historical moment, a symmetric positive definite matrix of a reinforcement signal, a historical evaluation value at the (t + 1)-th historical moment, an actual rotation speed at the t-th historical moment, an expected rotation speed at the t-th historical moment, an actual rotation speed at the (t - 1)-th historical moment, and an expected rotation speed at the (t - 1)-th historical moment;

[0193] A second obtaining unit, configured to update weights of the initial evaluation network according to the loss function of the initial evaluation network to obtain an evaluation network.

[0194] According to an embodiment of the present invention, a third obtaining module for updating parameters of an initial behavior network according to a historical evaluation value at the (t + 1)-th historical moment to obtain a behavior network includes:

[0195] A third obtaining unit, configured to determine a loss function of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment;

[0196] A fourth obtaining unit, configured to update weights of the initial behavior network according to the loss function of the initial behavior network to obtain a behavior network.

[0197] According to an embodiment of the present invention, a second output module 820 for inputting an expected rotation speed at the (t + 1)-th moment and an actual rotation speed at the t-th moment of a simulation turntable into an active disturbance rejection controller of the simulation turntable and outputting an expected current at the (t + 1)-th moment of the simulation turntable includes:

[0198] A fourth output unit, configured to output an observed speed at the (t + 1)-th moment, an observed acceleration at the (t + 1)-th moment, and an observed total disturbance at the (t + 1)-th moment by using an extended state observer of the active disturbance rejection controller according to the expected rotation speed at the (t + 1)-th moment and the actual rotation speed at the t-th moment of the simulation turntable;

[0199] A fifth output unit, configured to obtain a control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;

[0200] A sixth output unit, configured to output an expected current at the (t + 1)-th moment of the simulation turntable according to the control voltage at the (t + 1)-th moment.

[0201] According to an embodiment of the present invention, the fifth output unit for obtaining a control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment includes:

[0202] A first output subunit, configured to output an intermediate control voltage at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;

[0203] A second output subunit, configured to filter the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.

[0204] According to an embodiment of the present invention, the first obtaining module for obtaining the actual position, the actual rotation speed, and the actual current of the simulation turntable at the (t + 1)-th moment according to the motion result of the simulation turntable at the (t + 1)-th moment includes:

[0205] A fifth obtaining unit, configured to obtain the actual position at the (t + 1)-th moment and the actual current of the simulation turntable at the (t + 1)-th moment according to the motion result of the simulation turntable at the (t + 1)-th moment;

[0206] A sixth obtaining unit, configured to differentiate the actual position at the (t + 1)-th moment to obtain the actual rotation speed at the (t + 1)-th moment.

[0207] According to an embodiment of the present invention, any plurality of the modules, sub-modules, units, and sub-units, or at least part of the functions of any of them can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging the circuit in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0208] For example, any combination of the first output module 810, the second output module 820, the third output module 830, and the first control module 840 can be implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present invention, at least one of the first output module 810, the second output module 820, the third output module 830, and the first control module 840 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, at least one of the first output module 810, the second output module 820, the third output module 830, and the first control module 840 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0210] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A control method for a simulation turntable based on the fusion of data-driven and active disturbance rejection, characterized in that, Including: Input the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable, and output the desired rotational speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0; Input the desired rotational speed of the simulation turntable at the (t + 1)-th moment and the actual rotational speed at the t-th moment into the active disturbance rejection controller of the simulation turntable, and output the desired current of the simulation turntable at the (t + 1)-th moment; Input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network; Control the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.

2. The method according to claim 1, characterized in that The method further includes: Obtain the actual position at the (t + 1)-th moment, the actual rotational speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the movement result of the simulation turntable at the (t + 1)-th moment; Input the desired position of the simulation turntable at the (t + 2)-th moment and the actual position at the (t + 1)-th moment into the position controller, and output the desired rotational speed of the simulation turntable at the (t + 2)-th moment; Input the desired rotational speed of the simulation turntable at the (t + 2)-th moment and the actual rotational speed at the (t + 1)-th moment into the active disturbance rejection controller, and output the desired current of the simulation turntable at the (t + 2)-th moment; Input the actual rotational speed at the (t + 1)-th moment, the desired rotational speed at the (t + 1)-th moment, the actual rotational speed at the t-th moment, and the desired rotational speed at the t-th moment of the simulation turntable into the pre-trained control network, and output the compensation current at the (t + 2)-th moment; Control the movement of the simulation turntable according to the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable.

3. The method according to claim 1, wherein The step of inputting the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and outputting the compensation current of the simulation turntable at the (t + 1)-th moment includes: Input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into the behavior network, and output the initial compensation current at the (t + 1)-th moment; Input the initial compensation current at the (t + 1)-th moment, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment into the evaluation network, and output the evaluation value at the (t + 1)-th moment; Adjust the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Initialize the initial behavior network and the initial evaluation network; Input the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial behavior network to output the compensation current at the (t + 1)-th historical moment; Input the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial evaluation network to output the historical evaluation value at the (t + 1)-th historical moment; Update the parameters of the initial evaluation network according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network; Update the parameters of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment to obtain the behavior network.

5. The method according to claim 4, wherein The step of updating the parameters of the initial evaluation network according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network includes: Determine the loss function of the initial evaluation network according to the historical evaluation value at the t-th historical moment, the reinforcement signal symmetric positive definite matrix, the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment; Update the weights of the initial evaluation network according to the loss function of the initial evaluation network to obtain the evaluation network.

6. The method according to claim 4, wherein The step of updating the parameters of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment to obtain the behavior network includes: Determine the loss function of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment; Update the weights of the initial behavior network according to the loss function of the initial behavior network to obtain the behavior network.

7. The method according to any one of claims 1 to 3, characterized in that The step of inputting the desired speed at the (t + 1)-th moment and the actual speed at the t-th moment of the simulation turntable into the active disturbance rejection controller of the simulation turntable to output the desired current at the (t + 1)-th moment of the simulation turntable includes: Output the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment by using the extended state observer of the active disturbance rejection controller according to the desired speed at the (t + 1)-th moment and the actual speed at the t-th moment of the simulation turntable; Obtain the control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment; Output the desired current at the (t + 1)-th moment of the simulation turntable according to the control voltage at the (t + 1)-th moment.

8. The method according to claim 7, wherein Obtaining the control voltage of the simulation turntable at the (t + 1)-th moment according to the observed speed, the observed acceleration, and the observed total disturbance at the (t + 1)-th moment includes: Outputting an intermediate control voltage at the (t + 1)-th moment according to the observed speed, the observed acceleration, and the observed total disturbance at the (t + 1)-th moment; Performing a filtering process on the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.

9. The method according to claim 2, wherein Obtaining the actual position, the actual rotation speed, and the actual current of the simulation turntable at the (t + 1)-th moment according to the motion result of the simulation turntable at the (t + 1)-th moment includes: Obtaining the actual position at the (t + 1)-th moment and the actual current of the simulation turntable at the (t + 1)-th moment according to the motion result of the simulation turntable at the (t + 1)-th moment; Differentiating the actual position at the (t + 1)-th moment to obtain the actual rotation speed at the (t + 1)-th moment.

10. A simulation turntable control device based on the fusion of data-driven and active disturbance rejection, characterized in that, Including: A first output module, configured to input the desired position at the (t + 1)-th moment and the actual position at the t-th moment of the simulation turntable into the position controller of the simulation turntable, and output the desired rotation speed at the (t + 1)-th moment of the simulation turntable, where t is an integer greater than 0; A second output module, configured to input the desired rotation speed at the (t + 1)-th moment and the actual rotation speed at the t-th moment of the simulation turntable into the active disturbance rejection controller of the simulation turntable, and output the desired current at the (t + 1)-th moment of the simulation turntable; A third output module, configured to input the actual rotation speed at the t-th moment, the desired rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the desired rotation speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 1)-th moment of the simulation turntable, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network; A first control module, configured to control the motion of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.

Citation Information

Patent Citations

  • Method and apparatus for controlling tunnel ventilation

    JP2007100503A

  • Servo control device

    WO2020217282A1