Simulation turntable control method and device based on data driving and auto-disturbance rejection fusion
Through the control method of data-driven and self-immunity fusion, the position controller, self-immunity controller and control network are used to solve the problem of the five-axis flight simulation turntable control relying on accurate mathematical model, and the stability and robustness are improved.
Patent Information
- Application Number
- CN202510703160.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing control method of five-axis flight simulation turntable relies on accurate mathematical models, which makes it difficult to achieve effective control.
The control method based on data-driven and self-immunity fusion is adopted to output compensation current through position controllers, self-immunity controllers and pre-trained control networks (including behavioral networks and evaluation networks), reducing dependence on precise mathematical models, and improving control stability and robustness.
The effective control of the five-axis flight simulation turntable can be achieved without establishing an accurate mathematical model, which reduces the control difficulty and improves the stability and disturbance resistance of the control method.
Smart Images

Figure CN120233698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of turntable control, and more specifically, to a simulation turntable control method and device based on the fusion of data-driven and auto-disturbance rejection. Background Art
[0002] The five-axis flight simulation turntable is a simulation device for semi-physical simulation research and testing. The five-axis flight simulation turntable usually consists of a three-axis flight turntable, a two-axis target turntable, a simulation control platform, an inertial navigation system, a target simulator, product installation auxiliary equipment, etc. The three-axis flight turntable is used to simulate the attitude movement of the pod; the two-axis target turntable carries visible light and infrared target sources to reproduce the movement trajectory of the target.
[0003] In the process of implementing the inventive concept, the inventors found that there are at least the following problems in the related art: the control of the five-axis flight simulation turntable in the related art requires relying on an accurate mathematical model, and it is difficult to create the mathematical model, which increases the control difficulty of the turntable, resulting in a greater difficulty in implementing the control method of the five-axis flight simulation turntable. Summary of the Invention
[0004] In view of this, the present invention provides a simulation turntable control method and device based on the fusion of data-driven and auto-disturbance rejection.
[0005] One aspect of the present invention provides a simulation turntable control method based on the fusion of data-driven and auto-disturbance rejection, including:
[0006] Input the expected position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the above simulation turntable, and output the expected rotational speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0;
[0007] Input the expected rotational speed of the simulation turntable at the (t + 1)-th moment and the actual rotational speed at the t-th moment into the auto-disturbance rejection controller of the above simulation turntable, and output the expected current of the simulation turntable at the (t + 1)-th moment;
[0008] Input the actual rotational speed at the t-th moment, the expected rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the expected rotational speed at the (t - 1)-th moment of the above simulation turntable into a pre-trained control network, and output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network;
[0009] Control the movement of the above simulation turntable according to the expected current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the above simulation turntable.
[0010] According to an embodiment of the present invention, the above method further includes:
[0011] Based on the motion result of the above simulation turntable at the (t + 1)-th moment, obtain the actual position, actual rotation speed, and actual current of the above simulation turntable at the (t + 1)-th moment;
[0012] Input the expected position of the above simulation turntable at the (t + 2)-th moment and the actual position at the (t + 1)-th moment into the above position controller, and output the expected rotation speed of the above simulation turntable at the (t + 2)-th moment;
[0013] Input the expected rotation speed of the above simulation turntable at the (t + 2)-th moment and the actual rotation speed at the (t + 1)-th moment into the above active disturbance rejection controller, and output the expected current of the above simulation turntable at the (t + 2)-th moment;
[0014] Input the actual rotation speed at the (t + 1)-th moment, the expected rotation speed at the (t + 1)-th moment, the actual rotation speed at the t-th moment, and the expected rotation speed at the t-th moment of the above simulation turntable into the above pre-trained control network, and output the compensation current at the (t + 2)-th moment;
[0015] Control the movement of the above simulation turntable according to the expected current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the above simulation turntable.
[0016] According to an embodiment of the present invention, the step of inputting the actual rotation speed at the t-th moment, the expected rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the expected rotation speed at the (t - 1)-th moment of the above simulation turntable into the pre-trained control network to output the compensation current at the (t + 1)-th moment of the above simulation turntable includes:
[0017] Input the actual rotation speed at the t-th moment, the expected rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the expected rotation speed at the (t - 1)-th moment of the above simulation turntable into the above behavior network, and output the initial compensation current at the (t + 1)-th moment;
[0018] Input the initial compensation current at the (t + 1)-th moment, the actual rotation speed at the t-th moment, the expected rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the expected rotation speed at the (t - 1)-th moment into the above evaluation network, and output the evaluation value at the (t + 1)-th moment;
[0019] Adjust the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.
[0020] According to an embodiment of the present invention, the above method further includes:
[0021] Initialize the initial behavior network and the initial evaluation network;
[0022] Input the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment into the above initial behavior network, and output the compensation current at the (t + 1)-th historical moment;
[0023] Input the above compensation current at the (t + 1)-th historical moment, the above actual rotational speed at the t-th historical moment, the above desired rotational speed at the t-th historical moment, the above actual rotational speed at the (t - 1)-th historical moment, and the above desired rotational speed at the (t - 1)-th historical moment into the above initial evaluation network, and output the historical evaluation value at the (t + 1)-th historical moment;
[0024] Update the parameters of the above initial evaluation network according to the above historical evaluation value at the (t + 1)-th historical moment, the above actual rotational speed at the t-th historical moment, the above desired rotational speed at the t-th historical moment, the above actual rotational speed at the (t - 1)-th historical moment, and the above desired rotational speed at the (t - 1)-th historical moment, to obtain the above evaluation network;
[0025] Update the parameters of the above initial behavior network according to the above historical evaluation value at the (t + 1)-th historical moment, to obtain the above behavior network.
[0026] According to an embodiment of the present invention, the above updating the parameters of the above initial evaluation network according to the above historical evaluation value at the (t + 1)-th historical moment, the above actual rotational speed at the t-th historical moment, the above desired rotational speed at the t-th historical moment, the above actual rotational speed at the (t - 1)-th historical moment, and the above desired rotational speed at the (t - 1)-th historical moment, to obtain the above evaluation network, includes:
[0027] Determine the loss function of the above initial evaluation network according to the historical evaluation value at the t-th historical moment, the reinforcement signal symmetric positive definite matrix, the above historical evaluation value at the (t + 1)-th historical moment, the above actual rotational speed at the t-th historical moment, the above desired rotational speed at the t-th historical moment, the above actual rotational speed at the (t - 1)-th historical moment, and the above desired rotational speed at the (t - 1)-th historical moment;
[0028] Update the weights of the above initial evaluation network according to the loss function of the above initial evaluation network, to obtain the above evaluation network.
[0029] According to an embodiment of the present invention, the above updating the parameters of the above initial behavior network according to the above historical evaluation value at the (t + 1)-th historical moment, to obtain the above behavior network, includes:
[0030] Determine the loss function of the above initial behavior network according to the above historical evaluation value at the (t + 1)-th historical moment;
[0031] Update the weights of the above initial behavior network according to the loss function of the above initial behavior network, to obtain the above behavior network.
[0032] According to an embodiment of the present invention, inputting the expected rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable into the auto-disturbance rejection controller of the simulation turntable, and outputting the expected current at the (t + 1)-th moment of the simulation turntable, includes:
[0033] According to the expected rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable, use the extended state observer of the auto-disturbance rejection controller to output the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;
[0034] According to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment, obtain the control voltage at the (t + 1)-th moment of the simulation turntable;
[0035] According to the control voltage at the (t + 1)-th moment, output the expected current at the (t + 1)-th moment of the simulation turntable.
[0036] According to an embodiment of the present invention, the obtaining the control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment includes:
[0037] Output an intermediate control voltage at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;
[0038] Perform filtering processing on the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.
[0039] According to an embodiment of the present invention, the obtaining the actual position at the (t + 1)-th moment, the actual rotational speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment includes:
[0040] Obtain the actual position at the (t + 1)-th moment and the actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment;
[0041] Differentiate the actual position at the (t + 1)-th moment to obtain the actual rotational speed at the (t + 1)-th moment.
[0042] Another aspect of the present invention provides a control device for a simulation turntable based on the fusion of data-driven and auto-disturbance rejection, including:
[0043] The first output module is configured to input the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the above-mentioned simulation turntable, and output the desired rotational speed of the above-mentioned simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0;
[0044] The second output module is configured to input the desired rotational speed of the above-mentioned simulation turntable at the (t + 1)-th moment and the actual rotational speed at the t-th moment into the auto-disturbance rejection controller of the above-mentioned simulation turntable, and output the desired current of the above-mentioned simulation turntable at the (t + 1)-th moment;
[0045] The third output module is configured to input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the above-mentioned simulation turntable into a pre-trained control network, and output the compensation current of the above-mentioned simulation turntable at the (t + 1)-th moment, where the above-mentioned control network includes a behavior network and an evaluation network, and the above-mentioned behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the above-mentioned evaluation network;
[0046] The first control module is configured to control the movement of the above-mentioned simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the above-mentioned simulation turntable.
[0047] According to an embodiment of the present invention, through a pre-trained control network, the compensation current of the simulation turntable at the (t + 1)-th moment can be output through data such as the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment, avoiding the establishment of a specific mathematical model based on the disturbance of the simulation turntable. Based on the output data at each moment, the control of the simulation turntable can be realized, without the need to rely on an accurate mathematical model, reducing the implementation difficulty of the control method for the simulation turntable. Moreover, the auto-disturbance rejection controller is used to output the desired current at the (t + 1)-th moment, and the auto-disturbance rejection controller can further improve the stability and robustness of the control method. Description of the Drawings
[0048] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0049] Figure 1 Shows a flowchart of a control method for a simulation turntable based on the fusion of data-driven and auto-disturbance rejection according to an embodiment of the present invention;
[0050] Figure 2 Shows a flowchart of a training method for a control network according to an embodiment of the present invention;
[0051] Figure 3 Shows a schematic diagram of the network architecture of a control network according to an embodiment of the present invention;
[0052] Figure 4 shows the control schematic diagram of the active disturbance rejection controller according to an embodiment of the present invention;
[0053] Figure 5 shows the schematic diagram of the control architecture of the simulation turntable according to an embodiment of the present invention;
[0054] Figure 6A shows the schematic diagram of the update of the behavior network weights in the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention;
[0055] Figure 6B shows the schematic diagram of the update of the evaluation network weights in the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention;
[0056] Figure 7 shows the schematic diagram of the comparison of the simulation effects between the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention and the control method in the related art;
[0057] Figure 8 shows the block diagram of the simulation turntable control device based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention. Detailed implementation manners
[0058] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0059] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0060] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0061] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0062] In the embodiments of the present invention, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the involved data (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs.
[0063] In the related art, the main means to evaluate the performance of the guidance system and navigation equipment components is hardware-in-the-loop simulation. The method of hardware-in-the-loop simulation is much lower in cost and risk than actual physical tests because it can be tested without building a complete hardware system. At the same time, compared with real-time simulation that completely relies on computer models, hardware-in-the-loop simulation can also provide results closer to the actual system behavior.
[0064] During the casting or welding process of the structural parts in the five-axis flight simulation turntable, there will be a certain degree of difference from the design drawings, resulting in multiple resonance points. Moreover, according to the different positions between the axes, there will also be resonance points with different frequencies. During the movement of the five-axis flight simulation turntable, it will also be affected by multi-source disturbances such as multi-axis coupling, cogging torque fluctuation, friction torque fluctuation, and electromagnetic interference, making it difficult to establish an accurate mathematical model and having a greater control difficulty.
[0065] In view of this, the embodiments of the present invention provide a simulation turntable control method based on the fusion of data-driven and active disturbance rejection, including: inputting the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable to output the desired speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0; inputting the desired speed of the simulation turntable at the (t + 1)-th moment and the actual speed at the t-th moment into the active disturbance rejection controller of the simulation turntable to output the desired current of the simulation turntable at the (t + 1)-th moment; inputting the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network to output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network; controlling the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.
[0066] Figure 1The flowchart of the simulation turntable control method based on the fusion of data-driven and auto-disturbance rejection according to an embodiment of the present invention is shown.
[0067] As Figure 1 shown, the method includes operations S110 to S140.
[0068] In operation S110, the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment are input into the position controller of the simulation turntable, and the desired rotational speed of the simulation turntable at the (t + 1)-th moment is output, where t is an integer greater than 0.
[0069] In operation S120, the desired rotational speed of the simulation turntable at the (t + 1)-th moment and the actual rotational speed at the t-th moment are input into the auto-disturbance rejection controller of the simulation turntable, and the desired current of the simulation turntable at the (t + 1)-th moment is output.
[0070] In operation S130, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable are input into a pre-trained control network, and the compensation current of the simulation turntable at the (t + 1)-th moment is output, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network.
[0071] In operation S140, the simulation turntable is controlled to move according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.
[0072] According to an embodiment of the present invention, the simulation turntable proposed by the present invention may be a two-axis target turntable in a five-axis flight simulation turntable, and the moment proposed by the present invention may be a servo control cycle of the simulation turntable.
[0073] According to an embodiment of the present invention, the desired position of the simulation turntable at the (t + 1)-th moment may be the final position of the simulation turntable input from the outside, or may be an intermediate position decomposed according to the final position of the simulation turntable input from the outside. The desired position at the (t + 1)-th moment may be the position where it is desired for the simulation turntable to be located at the end of the (t + 1)-th moment.
[0074] According to an embodiment of the present invention, the actual position of the simulation turntable at the t-th moment may be the position of the simulation turntable at the end of the t-th moment, that is, the position of the simulation turntable at the beginning of the (t + 1)-th moment.
[0075] According to an embodiment of the present invention, the position controller may calculate the desired position at the (t + 1)-th moment and the actual position at the t-th moment to obtain the desired rotational speed of the simulation turntable at the (t + 1)-th moment.
[0076] According to an embodiment of the present invention, the actual rotational speed of the simulation turntable at the t-th moment can be the rotational speed of the simulation turntable at the end of the t-th moment, that is, the rotational speed of the simulation turntable at the start of the (t + 1)-th moment.
[0077] According to an embodiment of the present invention, the desired rotational speed at the (t + 1)-th moment can be the desired speed of the simulation turntable at the end of the (t + 1)-th moment. According to an embodiment of the present invention, based on the desired rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment, the desired current of the simulation turntable at the (t + 1)-th moment can be output through an active disturbance rejection controller, such that the simulation turntable rotates according to the desired current at the (t + 1)-th moment, and the speed at the end of the (t + 1)-th moment can be the desired rotational speed at the (t + 1)-th moment.
[0078] According to an embodiment of the present invention, the active disturbance rejection controller can use the active disturbance rejection control (ADRC) method for output. The ADRC method does not rely on an accurate mathematical model, treats all modeling errors as disturbances, and only requires a rough process model to design a control loop. Its linear form is equivalent to a special case of classical state space control based on the internal model principle, with disturbance estimation and compensation, and has high robustness and anti-disturbance performance.
[0079] According to an embodiment of the present invention, the control network can be a network pre-trained for compensating the desired current. The control network includes a behavior network and an evaluation network. The behavior network can output a specific compensation current, and the evaluation network can evaluate the compensation current output by the behavior network to determine whether the output result of the behavior network is optimal.
[0080] According to an embodiment of the present invention, based on the actual rotational speed and the desired rotational speed at the t-th moment, the rotational speed error at the t-th moment can be determined. Based on the actual rotational speed and the desired rotational speed at the (t - 1)-th moment, the rotational speed error at the (t - 1)-th moment can be determined. Inputting the rotational speed error at the t-th moment and the rotational speed error at the (t - 1)-th moment into the pre-trained control network can output the compensation current at the (t + 1)-th moment. That is, the control network can output the compensation current for the next moment based on the rotational speed errors of the two adjacent previous moments.
[0081] According to an embodiment of the present invention, based on the desired current and the compensation current at the (t + 1)-th moment of the simulation turntable, the control current at the (t + 1)-th moment can be obtained. The control current at the (t + 1)-th moment can be the current used to control the simulation turntable at the (t + 1)-th moment. Based on the control current at the (t + 1)-th moment and the actual current at the t-th moment, the movement of the simulation turntable is controlled.
[0082] According to an embodiment of the present invention, through a pre-trained control network, the compensation current of the simulation turntable at the (t + 1)-th moment can be output based on data such as the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment. This avoids establishing a specific mathematical model based on the disturbance of the simulation turntable, reducing the implementation difficulty of the control method for the simulation turntable. Moreover, the auto-disturbance rejection controller is used to output the desired current at the (t + 1)-th moment, and the auto-disturbance rejection controller can further improve the stability and robustness of the control method.
[0083] According to an embodiment of the present invention, the method may further include: obtaining the actual position, actual speed, and actual current of the simulation turntable at the (t + 1)-th moment based on the motion result of the simulation turntable at the (t + 1)-th moment; inputting the desired position at the (t + 2)-th moment and the actual position at the (t + 1)-th moment of the simulation turntable into a position controller to output the desired speed of the simulation turntable at the (t + 2)-th moment; inputting the desired speed at the (t + 2)-th moment and the actual speed at the (t + 1)-th moment of the simulation turntable into an auto-disturbance rejection controller to output the desired current of the simulation turntable at the (t + 2)-th moment; inputting the actual speed at the (t + 1)-th moment, the desired speed at the (t + 1)-th moment, the actual speed at the t-th moment, and the desired speed at the t-th moment of the simulation turntable into a pre-trained control network to output the compensation current at the (t + 2)-th moment; and controlling the motion of the simulation turntable based on the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable.
[0084] According to an embodiment of the present invention, after the simulation turntable moves at the (t + 1)-th moment, the position, speed, and current of the simulation turntable can be detected to obtain the actual position, actual speed, and actual current at the (t + 1)-th moment.
[0085] According to an embodiment of the present invention, the motion control of the simulation turntable at each moment will refer to the actual position, actual speed, and actual current at the previous moment, and adjust the output of the simulation turntable according to the differences between the actual position and the desired position, the actual speed and the desired speed, and the actual current and the desired current.
[0086] According to an embodiment of the present invention, the motion of the simulation turntable at the (t + 2)-th moment can refer to the description at the (t + 1)-th moment, which will not be elaborated here.
[0087] According to an embodiment of the present invention, inputting the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network to output the compensation current at the (t + 1)-th moment of the simulation turntable, including: inputting the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a behavior network to output the initial compensation current at the (t + 1)-th moment; inputting the initial compensation current at the (t + 1)-th moment, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment into an evaluation network to output the evaluation value at the (t + 1)-th moment; and adjusting the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.
[0088] According to an embodiment of the present invention, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment can be input into a behavior network, and the behavior network can output the initial compensation current at the (t + 1)-th moment. Since it is uncertain whether the initial compensation current at the (t + 1)-th moment meets the requirements, the initial compensation current at the (t + 1)-th moment, the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment can be input into an evaluation network to output the evaluation value at the (t + 1)-th moment for the initial compensation current at the (t + 1)-th moment. When the evaluation value at the (t + 1)-th moment indicates that the initial compensation current at the (t + 1)-th moment can make the position of the simulation turntable at the end of the (t + 1)-th moment the same as the desired position at the (t + 1)-th moment, the initial compensation current at the (t + 1)-th moment can be output as the compensation current at the (t + 1)-th moment. When the evaluation value at the (t + 1)-th moment cannot make the position of the simulation turntable at the end of the (t + 1)-th moment the same as the desired position at the (t + 1)-th moment, the initial compensation current at the (t + 1)-th moment can be adjusted to obtain the compensation current at the (t + 1)-th moment. During the adjustment of the initial compensation current at the (t + 1)-th moment, multiple intermediate compensation currents at the (t + 1)-th moment can be obtained, and the intermediate compensation currents at the (t + 1)-th moment can be input into the evaluation network to obtain the corresponding evaluation values, and the compensation current at the (t + 1)-th moment can be obtained according to the evaluation values.
[0089] Figure 2 The flowchart of the training method of the control network according to an embodiment of the present invention is shown.
[0090] As Figure 2 shown, the method includes operation S210 to operation S250.
[0091] In operation S210, initialize the initial behavior network and the initial evaluation network.
[0092] In operation S220, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment are input into the initial behavior network, and the compensation current at the (t + 1)-th historical moment is output.
[0093] In operation S230, the compensation current at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment are input into the initial evaluation network, and the historical evaluation value at the (t + 1)-th historical moment is output.
[0094] In operation S240, based on the historical evaluation value at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment, the parameters of the initial evaluation network are updated to obtain the evaluation network.
[0095] In operation S250, based on the historical evaluation value at the (t + 1)-th historical moment, the parameters of the initial behavior network are updated to obtain the behavior network.
[0096] According to an embodiment of the present invention, the control network can be trained based on the historical action data of the simulation turntable.
[0097] Figure 3 FIG. shows a schematic diagram of the network architecture of the control network according to an embodiment of the present invention.
[0098] Figure 3 In FIG. a schematically shows the network architecture of the initial behavior network. The initial behavior network may include an input layer, a hidden layer, and an output layer. The input layer of the initial behavior network is used to input the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment. The hidden layer of the initial behavior network obtains features based on the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment. The output layer of the initial behavior network is used to output the compensation current at the (t + 1)-th historical moment according to the features in the hidden layer. Among them, represents the rotational speed error at the historical moment input into the initial behavior network, represents the weight from the input layer to the hidden layer of the initial behavior network, represents the weight from the hidden layer to the output layer of the initial behavior network, represents the compensation current at the historical moment of the output of the initial behavior network.
[0099] Figure 3Figure b shows the network architecture of the initial evaluation network. The initial evaluation network may include an input layer, a hidden layer, and an output layer. The input layer of the initial evaluation network is used to input the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment. The hidden layer of the initial evaluation network obtains features based on the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment. The output layer of the initial evaluation network is used to output the historical evaluation value at the (t + 1)-th historical moment according to the features in the hidden layer. Among them, represents the weight from the input layer to the hidden layer of the initial evaluation network, represents the weight from the hidden layer to the output layer of the initial evaluation network, represents the historical evaluation value at the historical moment output by the initial evaluation network.
[0100] According to an embodiment of the present invention, based on the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment, the parameters of the initial evaluation network can be updated to obtain an evaluation network. According to the historical evaluation value at the (t + 1)-th historical moment, the parameters of the initial behavior network can be updated to obtain a behavior network.
[0101] The design of the behavior network is as follows:
[0102] (1)
[0103] (2)
[0104] (3)
[0105] (4)
[0106] Among them, , , represent intermediate variables output by the initial behavior network at the t-th historical moment, represents the number of nodes in the input layer of the initial behavior network. i represents the i-th node in the input layer of the initial behavior network, represents the number of nodes in the hidden layer of the initial behavior network. j represents the j-th node in the hidden layer of the initial behavior network, represents the base of the natural logarithm, represents the output weight from the i-th node in the input layer to the j-th node in the hidden layer of the initial behavior network at the t-th historical moment, Denote the output weight of the j-th node in the hidden layer of the initial behavior network at the t-th historical moment. , which represents the combination of the rotational speed error at the t-th historical moment and the rotational speed error at the (t - 1)-th historical moment, where Denote the error between the actual rotational speed and the desired rotational speed at the t-th historical moment. Denote the error between the actual rotational speed and the desired rotational speed at the (t - 1)-th historical moment. Denote the compensated current at the (t + 1)-th historical moment output by the initial behavior network.
[0107] The input of the evaluation network is , which represents the combination of the rotational speed error at the t-th historical moment, the rotational speed error at the (t - 1)-th historical moment, and the compensated current at the (t + 1)-th historical moment. The design of the evaluation network is as follows:
[0108] (5)
[0109] (6)
[0110] (7)
[0111] Where , Denote the intermediate variable output by the initial evaluation network at the t-th historical moment. Denote the number of nodes in the input layer of the initial evaluation network. Denote the number of nodes in the hidden layer of the initial evaluation network. Denote the output weight from the i-th node in the input layer to the j-th node in the hidden layer of the initial evaluation network at the t-th historical moment. Denote the output weight of the j-th node in the hidden layer of the initial evaluation network at the t-th historical moment. Denote the historical evaluation value at the (t + 1)-th historical moment output by the initial evaluation network.
[0112] According to the embodiments of the present invention, based on the historical evaluation value at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment, update the parameters of the initial evaluation network to obtain an evaluation network, including: determining the loss function of the initial evaluation network according to the historical evaluation value at the t-th historical moment, the positive definite matrix of the reinforcement signal symmetry, the historical evaluation value at the (t + 1)-th historical moment, the actual rotational speed at the t-th historical moment, the desired rotational speed at the t-th historical moment, the actual rotational speed at the (t - 1)-th historical moment, and the desired rotational speed at the (t - 1)-th historical moment; updating the weights of the initial evaluation network according to the loss function of the initial evaluation network to obtain an evaluation network.
[0113] According to an embodiment of the present invention, the historical evaluation value at the t-th historical moment can be obtained by referring to the process of the historical evaluation value at the (t + 1)-th historical moment, which will not be elaborated here.
[0114] According to an embodiment of the present invention, the reinforcement signal symmetric positive definite matrix can be a matrix that is both a symmetric matrix and a positive definite matrix to ensure the stability of the update of the initial evaluation network.
[0115] According to an embodiment of the present invention, the loss function of the initial evaluation network is as follows:
[0116] (8)
[0117] where represents the reinforcement signal at the t-th historical moment of the initial evaluation network, is the reinforcement signal symmetric positive definite matrix, represents the transpose of , is the loss factor, represents the historical evaluation value at the (t + 1)-th historical moment output by the initial evaluation network, represents the historical evaluation value at the t-th historical moment output by the initial evaluation network.
[0118] According to the gradient descent method, the weight update of the initial evaluation network is as follows:
[0119] (9)
[0120] (10)
[0121] where represents the training learning rate of the initial evaluation network, .
[0122] According to an embodiment of the present invention, according to the historical evaluation value at the (t + 1)-th historical moment, the parameters of the initial behavior network are updated to obtain the behavior network, including: determining the loss function of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment; updating the weights of the initial behavior network according to the loss function of the initial behavior network to obtain the behavior network.
[0123] The loss function of the initial behavior network is:
[0124] (11)
[0125] where represents the square of the historical evaluation value at the (t + 1)-th historical moment output by the initial evaluation network .
[0126] The weights of the initial behavior network are updated as follows:
[0127] (12)
[0128] (13)
[0129] where represents the training learning rate of the initial behavior network, .
[0130] According to the embodiments of the present invention, the weights of the initial evaluation network and the initial behavior network can be iteratively updated, and the iteration is stopped when the number of iterations reaches a threshold or when the initial evaluation network and the initial behavior network meet the requirements, obtaining a behavior network and an evaluation network, that is, a control network. Based on data such as the rotation speed and position of the simulation turntable, the control network can output a compensation current, eliminating the need to establish a corresponding disturbance model for the simulation turntable and reducing the control difficulty of the simulation turntable.
[0131] Figure 4 FIG. shows the control schematic diagram of the active disturbance rejection controller according to the embodiments of the present invention.
[0132] As Figure 4 shown, the active disturbance rejection controller combines the classical proportional-integral-derivative (PID) control method with ease of use. The input of the active disturbance rejection controller is the desired rotation speed at the current moment and the actual rotation speed at the previous moment. The output control voltage is . After low-pass filtering and notch filtering, the desired current at the current moment can be obtained, and then the desired current at the current moment is sent to the driver through analog quantity to control the action of the simulation turntable. The active disturbance rejection controller includes an extended state observer and a PID controller. The outputs of the extended state observer are , and . represents the proportional coefficient of the PID controller, represents the differential coefficient of the PID controller, represents the critical gain in the active disturbance rejection controller.
[0133] According to the embodiments of the present invention, represents the observed speed of the simulation turntable, represents the observed acceleration of the simulation turntable, represents the observed total disturbance of the simulation turntable.
[0134] Let , the output of the extended state observer can be obtained as follows:
[0135] (14)
[0136] (15)
[0137] where, denotes the estimation of the extended state observer at the th moment obtained by using the estimation at the th moment, denotes the control voltage output by the active disturbance rejection controller at the th moment, denotes the output of the extended state observer at the th moment obtained by using the estimation at the th moment, denotes the output at the th moment obtained by using the estimation at the th moment, denotes the actual rotational speed at the th moment.
[0138] where, , , , are the gains of the extended state observer, denotes the servo cycle of the simulation turntable, denotes the critical gain. , and denote the three-dimensional components of the gain of the extended state observer, , and are expressed as follows:
[0139] (16)
[0140] (17)
[0141] (18)
[0142] where, is the bandwidth of the extended state observer.
[0143] The output of the active disturbance rejection controller is:
[0144] (19)
[0145] (20)
[0146] (21)
[0147] Among them, represents the control voltage output by the active disturbance rejection controller at the th moment, represents the desired rotational speed at the th moment, represents the observed speed of the simulation turntable at the th moment, represents the observed acceleration of the simulation turntable at the th moment, represents the observed total disturbance of the simulation turntable at the th moment, is the bandwidth of the active disturbance rejection controller. The control voltage output by the active disturbance rejection controller at the th moment can output the desired current at the th moment after low-pass filtering and notch filtering.
[0148] Specifically, in the embodiment of the present invention, according to the desired rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable, the extended state observer of the active disturbance rejection controller can output the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment; according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment, the control voltage at the (t + 1)-th moment of the simulation turntable can be calculated; according to the control voltage at the (t + 1)-th moment, after low-pass filtering and notch filtering, the desired current at the (t + 1)-th moment of the simulation turntable is output.
[0149] According to the embodiment of the present invention, obtaining the control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment may include: outputting the intermediate control voltage at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment; performing filtering processing on the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.
[0150] According to the embodiment of the present invention, the filtering processing may include low-pass filtering and notch filtering. By performing filtering processing on the intermediate control voltage at the (t + 1)-th moment, the interference signal in the intermediate control voltage at the (t + 1)-th moment can be filtered out, and a more stable and accurate control voltage at the (t + 1)-th moment can be obtained.
[0151] According to an embodiment of the present invention, the ADRC method combines the ease of use of the classical PID control method, does not rely on an accurate mathematical model, treats all modeling errors as disturbances, and only requires a very rough process model to design a control loop. Its linear form is equivalent to a special case of classical state - space control based on the internal - model principle, has disturbance estimation and compensation, and has extremely strong robustness and anti - disturbance performance.
[0152] According to an embodiment of the present invention, based on the motion result of the simulation turntable at the (t + 1)-th moment, the actual position, actual rotation speed, and actual current of the simulation turntable at the (t + 1)-th moment are obtained, including: based on the motion result of the simulation turntable at the (t + 1)-th moment, the actual position at the (t + 1)-th moment and the actual current of the simulation turntable at the (t + 1)-th moment are obtained; differentiating the actual position at the (t + 1)-th moment to obtain the actual rotation speed at the (t + 1)-th moment.
[0153] According to an embodiment of the present invention, the simulation turntable can be driven by a DC brushless torque motor and is equipped with a high - precision absolute encoder. The absolute encoder can obtain the actual position at the (t + 1)-th moment through information fusion technology, and differentiating the actual position at the (t + 1)-th moment to obtain the actual rotation speed at the (t + 1)-th moment.
[0154] Figure 5 The schematic diagram of the control architecture of the simulation turntable according to an embodiment of the present invention is shown.
[0155] As Figure 5 shown, the position controller can receive the actual position of the simulation turntable in real - time. According to the input desired position and actual position, it outputs the desired rotation speed to the active disturbance rejection controller (ADRC). The active disturbance rejection controller can receive the actual rotation speed of the simulation turntable in real - time and outputs the desired current through the active disturbance rejection controller and low - pass and notch algorithms. The actual rotation speed of the simulation turntable is also input into the control network. As known from the foregoing, the control network is a pre - trained network, and the utility function is a parameter related to the convergence condition of the control network set during the training of the control network. The convergence condition of the control network is that the error between the actual rotation speed and the desired rotation speed at the current moment is equal to zero. Therefore, during the training of the control network in the present invention, the utility function can be set to 0. The control network can store the rotation speed errors of the previous two moments, and input the rotation speed errors of the previous two moments and the compensation current at the (t + 1)-th moment output by the behavior network into the evaluation network. The evaluation network outputs the evaluation value at the (t + 1)-th moment, and feeds it back through the loss factor , the reinforcement signal at the t - th moment, and the evaluation value at the t - th moment to . The behavior network outputs a compensation current according to the final , and the final input current is obtained from the desired current and the compensation current . The input current is the direct-axis current, represents the quadrature-axis current. According to and the actual current , the motion of the simulation turntable is controlled. Among them, the driver of the simulation turntable can be equivalent to the proportional-integral controller (PI), inverse Park transformation, space vector modulation (SVPWM), and three-phase inverter in the figure. The inverse Park transformation is used to convert the voltages and in the dq coordinate system into voltages and in the αβ coordinate system. The SVPWM is used to utilize the DC bus voltage to improve the output ability of the inverter. The three-phase inverter is used to convert direct current into sinusoidal alternating current. Finally, the controller outputs a voltage to control the motion of the brushless DC motor (BLDC). At the same time, the BLDC outputs the actual current to the Clark transformation. The Clark transformation converts into currents and in the αβ coordinate system. The Park transformation converts the currents and in the αβ coordinate system into currents and in the dq coordinate system and feeds them back to the driver. During the motion of the BLDC, the interference term can be regarded as multi-source interference and is reflected in the frame axis. The absolute encoder obtains the real-time position through the frame axis, and outputs the actual position to the inverse Park transformation, Park transformation, and position controller through information fusion, and outputs the actual speed to the ADRC through differentiation. Figure 5 The corresponding moments of the desired position, actual position, actual speed, etc. in
[0156] can be understood in combination with other embodiments of the present invention and will not be elaborated here. Figure 5 It can be seen from that in the control architecture of the simulation turntable, a closed-loop control of the current loop, speed loop, and position loop is formed through current feedback, speed feedback, and position feedback. Before the simulation turntable operates, the parameters in each loop need to be adjusted. The parameters required in the current loop can be adjusted first, and then the parameters in the speed loop and position loop are adjusted in turn. The parameter adjustment in the speed loop can use the frequency-domain sweep method to debug the parameters of the low-pass filter and notch filter, and then adjust the critical gain and the bandwidth of the extended state observer The parameters of the position loop can also be adjusted using the amplitude-frequency and phase-frequency characteristic curves to obtain a stable operating control system, and then a control network is added.
[0157] Perform simulation analysis on the control method of the embodiment of the present invention, and design the resistance of the DC brushless motor , inductance 5.25 mH, number of pole pairs 4, magnetic flux 0.187 Wb, moment of inertia , parameters of the speed loop: , , . Parameters of the control network: , , , , , , , , select different rotational speeds of the simulation turntable, and obtain the weight updates of the behavior network and the evaluation network as shown in Fig. 6.
[0158] Figure 6A Shows a schematic diagram of the weight update of the behavior network in the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention.
[0159] Figure 6B Shows a schematic diagram of the weight update of the evaluation network in the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention.
[0160] As Figure 6A shown, the weights of the behavior network in different nodes of the control network can be obtained. From top to bottom, they represent the weights of the behavior network from the 1st node in the input layer to the 1st node, 2nd node, 3rd node, 4th node, 5th node, and 6th node in the hidden layer. Among them, the abscissa represents the moment of weight update, and the ordinate represents the numerical change of the weight. As Figure 6B shown, the weights of the evaluation network in different nodes of the control network can be obtained. From top to bottom, they represent the weights of the evaluation network from the 1st node in the input layer to the 1st node, 2nd node, 3rd node, 4th node, 5th node, and 6th node in the hidden layer. Among them, the abscissa represents the moment of weight update, and the ordinate represents the numerical change of the weight. It should be noted that Figure 6A and Figure 6B only show partial node weight update processes.
[0161] Figure 7 Shows a schematic diagram of the comparison of the simulation effects between the simulation analysis of the simulation turntable control method based on data-driven and active disturbance rejection fusion according to an embodiment of the present invention and the control method in the related art.
[0162] AsFigure 7 As shown, the abscissa represents the running time of the simulation turntable, and the ordinate represents the speed of the simulation turntable. The input of the simulation turntable is the rotational speed with multi-source disturbances, and a step disturbance signal is added at 1.5 s. Among them, the straight line (ADRC) is the simulation effect without adding the control network, and the dashed line (ADRC-Aux) represents the simulation effect after adding the control network. It can be seen that the speed of the simulation turntable after adding the control network is more stable, and there are significant improvements in both the anti-disturbance ability and the rate smoothness.
[0163] According to an embodiment of the present invention, the control method proposed by the present invention is applied to a five-axis flight simulation turntable, and the rate smoothness is detected to obtain the rate smoothness of the simulation turntable at 15 ms intervals, as shown in Table 1 and Table 2 below. Among them, Table 1 is the inspection record form of the pitch axis rate smoothness of the simulation turntable, and Table 2 is the inspection record form of the azimuth axis rate smoothness of the simulation turntable. The horizontal header represents the input desired rate command, and the vertical header represents the actual angle turned by the pitch axis or azimuth axis of the simulation turntable when running at the desired rate every 15 ms interval.
[0164] Table 1
[0165]
[0166] Table 2
[0167]
[0168] From the actual test results in the above table, it can be concluded that after the simulation turntable adds the control method of the simulation turntable based on data-driven and auto-disturbance rejection fusion, its rate smoothness can reach (15 ms fixed time interval), and the performance in terms of rate smoothness is significantly improved compared with the conventional turntable.
[0169] Figure 8 The block diagram of the simulation turntable control device based on data-driven and auto-disturbance rejection fusion according to an embodiment of the present invention is shown.
[0170] As Figure 8 shown, the device 800 includes a first output module 810, a second output module 820, a third output module 830, and a first control module 840.
[0171] The first output module 810 is configured to input the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable, and output the desired rotational speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0;
[0172] The second output module 820 is configured to input the desired rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable into the auto-disturbance rejection controller of the simulation turntable, and output the desired current at the (t + 1)-th moment of the simulation turntable;
[0173] The third output module 830 is configured to input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 1)-th moment of the simulation turntable, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network;
[0174] The first control module 840 is configured to control the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.
[0175] According to an embodiment of the present invention, the apparatus 800 further includes:
[0176] The first obtaining module is configured to obtain the actual position at the (t + 1)-th moment, the actual rotational speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the movement result of the simulation turntable at the (t + 1)-th moment;
[0177] The fourth output module is configured to input the desired position at the (t + 2)-th moment and the actual position at the (t + 1)-th moment of the simulation turntable into a position controller, and output the desired rotational speed at the (t + 2)-th moment of the simulation turntable;
[0178] The fifth output module is configured to input the desired rotational speed at the (t + 2)-th moment and the actual rotational speed at the (t + 1)-th moment of the simulation turntable into the auto-disturbance rejection controller, and output the desired current at the (t + 2)-th moment of the simulation turntable;
[0179] The sixth output module is configured to input the actual rotational speed at the (t + 1)-th moment, the desired rotational speed at the (t + 1)-th moment, the actual rotational speed at the t-th moment, and the desired rotational speed at the t-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 2)-th moment;
[0180] The second control module is configured to control the movement of the simulation turntable according to the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable.
[0181] According to an embodiment of the present invention, the third output module 830 configured to input the actual rotational speed at the t-th moment, the desired rotational speed at the t-th moment, the actual rotational speed at the (t - 1)-th moment, and the desired rotational speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 1)-th moment of the simulation turntable includes:
[0182] The first output unit is configured to input the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into the behavior network, and output the initial compensation current at the (t + 1)-th moment;
[0183] The second output unit is configured to input the initial compensation current at the (t + 1)-th moment, the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment into the evaluation network, and output the evaluation value at the (t + 1)-th moment;
[0184] The third output unit is configured to adjust the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.
[0185] According to an embodiment of the present invention, the apparatus 800 further includes:
[0186] An initialization module, configured to initialize the initial behavior network and the initial evaluation network;
[0187] The seventh output module is configured to input the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial behavior network, and output the compensation current at the (t + 1)-th historical moment;
[0188] The eighth output module is configured to input the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial evaluation network, and output the historical evaluation value at the (t + 1)-th historical moment;
[0189] The second obtaining module is configured to update the parameters of the initial evaluation network according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network;
[0190] The third obtaining module is configured to update the parameters of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment to obtain the behavior network.
[0191] According to an embodiment of the present invention, the second obtaining module, which is configured to update the parameters of the initial evaluation network according to the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network, includes:
[0192] A first obtaining unit, configured to determine a loss function of an initial evaluation network according to a historical evaluation value at the t-th historical moment, a symmetric positive definite matrix of a reinforcement signal, a historical evaluation value at the (t + 1)-th historical moment, an actual rotational speed at the t-th historical moment, a desired rotational speed at the t-th historical moment, an actual rotational speed at the (t - 1)-th historical moment, and a desired rotational speed at the (t - 1)-th historical moment;
[0193] A second obtaining unit, configured to update weights of the initial evaluation network according to the loss function of the initial evaluation network to obtain an evaluation network.
[0194] According to an embodiment of the present invention, a third obtaining module for updating parameters of an initial behavior network according to a historical evaluation value at the (t + 1)-th historical moment to obtain a behavior network includes:
[0195] A third obtaining unit, configured to determine a loss function of the initial behavior network according to the historical evaluation value at the (t + 1)-th historical moment;
[0196] A fourth obtaining unit, configured to update weights of the initial behavior network according to the loss function of the initial behavior network to obtain a behavior network.
[0197] According to an embodiment of the present invention, a second output module 820 for inputting a desired rotational speed at the (t + 1)-th moment and an actual rotational speed at the t-th moment of a simulation turntable into an active disturbance rejection controller of the simulation turntable and outputting a desired current at the (t + 1)-th moment of the simulation turntable includes:
[0198] A fourth output unit, configured to output an observed speed at the (t + 1)-th moment, an observed acceleration at the (t + 1)-th moment, and an observed total disturbance at the (t + 1)-th moment by using an extended state observer of the active disturbance rejection controller according to the desired rotational speed at the (t + 1)-th moment and the actual rotational speed at the t-th moment of the simulation turntable;
[0199] A fifth output unit, configured to obtain a control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;
[0200] A sixth output unit, configured to output a desired current at the (t + 1)-th moment of the simulation turntable according to the control voltage at the (t + 1)-th moment.
[0201] According to an embodiment of the present invention, the fifth output unit for obtaining a control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment includes:
[0202] A first output subunit, configured to output an intermediate control voltage at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment;
[0203] A second output subunit, configured to filter the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.
[0204] According to an embodiment of the present invention, a first obtaining module for obtaining the actual position, the actual rotation speed, and the actual current of the simulation turntable at the (t + 1)-th moment according to the motion result of the simulation turntable at the (t + 1)-th moment includes:
[0205] A fifth obtaining unit, configured to obtain the actual position at the (t + 1)-th moment and the actual current of the simulation turntable at the (t + 1)-th moment according to the motion result of the simulation turntable at the (t + 1)-th moment;
[0206] A sixth obtaining unit, configured to differentiate the actual position at the (t + 1)-th moment to obtain the actual rotation speed at the (t + 1)-th moment.
[0207] According to an embodiment of the present invention, any multiple of the modules, sub-modules, units, and sub-units, or at least part of the functions of any multiple of them, can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, in the form of hardware or firmware, or in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0208] For example, any combination of the first output module 810, the second output module 820, the third output module 830, and the first control module 840 may be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units may be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units may be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present invention, at least one of the first output module 810, the second output module 820, the third output module 830, and the first control module 840 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first output module 810, the second output module 820, the third output module 830, and the first control module 840 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0210] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present invention.
Claims
1. A control method for a simulation turntable based on the fusion of data-driven and active disturbance rejection, characterized in that, Including: Input the desired position of the simulation turntable at the (t + 1)-th moment and the actual position at the t-th moment into the position controller of the simulation turntable, and output the desired speed of the simulation turntable at the (t + 1)-th moment, where t is an integer greater than 0; Input the desired speed of the simulation turntable at the (t + 1)-th moment and the actual speed at the t-th moment into the active disturbance rejection controller of the simulation turntable, and output the desired current of the simulation turntable at the (t + 1)-th moment; Input the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current of the simulation turntable at the (t + 1)-th moment, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network; Control the movement of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.
2. The method according to claim 1, wherein The method further includes: Obtain the actual position at the (t + 1)-th moment, the actual speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the movement result of the simulation turntable at the (t + 1)-th moment; Input the desired position of the simulation turntable at the (t + 2)-th moment and the actual position at the (t + 1)-th moment into the position controller, and output the desired speed of the simulation turntable at the (t + 2)-th moment; Input the desired speed of the simulation turntable at the (t + 2)-th moment and the actual speed at the (t + 1)-th moment into the active disturbance rejection controller, and output the desired current of the simulation turntable at the (t + 2)-th moment; Input the actual speed at the (t + 1)-th moment, the desired speed at the (t + 1)-th moment, the actual speed at the t-th moment, and the desired speed at the t-th moment of the simulation turntable into the pre-trained control network, and output the compensation current at the (t + 2)-th moment; Control the movement of the simulation turntable according to the desired current at the (t + 2)-th moment, the compensation current at the (t + 2)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable.
3. The method according to claim 1, wherein The step of inputting the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network and outputting the compensation current at the (t + 1)-th moment of the simulation turntable includes: Input the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment of the simulation turntable into the behavior network, and output the initial compensation current at the (t + 1)-th moment; Input the initial compensation current at the (t + 1)-th moment, the actual speed at the t-th moment, the desired speed at the t-th moment, the actual speed at the (t - 1)-th moment, and the desired speed at the (t - 1)-th moment into the evaluation network, and output the evaluation value at the (t + 1)-th moment; Adjust the initial compensation current at the (t + 1)-th moment according to the evaluation value at the (t + 1)-th moment to obtain the compensation current at the (t + 1)-th moment.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Initialize the initial behavior network and the initial evaluation network; Input the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial behavior network to output the compensation current at the (t + 1)-th historical moment; Input the compensation current at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment into the initial evaluation network to output the historical evaluation value at the (t + 1)-th historical moment; Update the parameters of the initial evaluation network based on the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network; Update the parameters of the initial behavior network based on the historical evaluation value at the (t + 1)-th historical moment to obtain the behavior network.
5. The method according to claim 4, wherein The step of updating the parameters of the initial evaluation network based on the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment to obtain the evaluation network includes: Determine the loss function of the initial evaluation network based on the historical evaluation value at the t-th historical moment, the reinforcement signal symmetric positive definite matrix, the historical evaluation value at the (t + 1)-th historical moment, the actual speed at the t-th historical moment, the desired speed at the t-th historical moment, the actual speed at the (t - 1)-th historical moment, and the desired speed at the (t - 1)-th historical moment; Update the weights of the initial evaluation network according to the loss function of the initial evaluation network to obtain the evaluation network.
6. The method according to claim 4, wherein The step of updating the parameters of the initial behavior network based on the historical evaluation value at the (t + 1)-th historical moment to obtain the behavior network includes: Determine the loss function of the initial behavior network based on the historical evaluation value at the (t + 1)-th historical moment; Update the weights of the initial behavior network according to the loss function of the initial behavior network to obtain the behavior network.
7. The method according to any one of claims 1 to 3, characterized in that The step of inputting the desired speed at the (t + 1)-th moment and the actual speed at the t-th moment of the simulation turntable into the active disturbance rejection controller of the simulation turntable to output the desired current at the (t + 1)-th moment of the simulation turntable includes: Output the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment by using the extended state observer of the active disturbance rejection controller according to the desired speed at the (t + 1)-th moment and the actual speed at the t-th moment of the simulation turntable; Obtain the control voltage at the (t + 1)-th moment of the simulation turntable according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment; Output the desired current at the (t + 1)-th moment of the simulation turntable according to the control voltage at the (t + 1)-th moment.
8. The method according to claim 7, characterized in that, Obtaining the control voltage of the simulation turntable at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment includes: Outputting an intermediate control voltage at the (t + 1)-th moment according to the observed speed at the (t + 1)-th moment, the observed acceleration at the (t + 1)-th moment, and the observed total disturbance at the (t + 1)-th moment; Performing a filtering process on the intermediate control voltage at the (t + 1)-th moment to obtain the control voltage at the (t + 1)-th moment.
9. The method according to claim 2, characterized in that Obtaining the actual position at the (t + 1)-th moment, the actual rotation speed at the (t + 1)-th moment, and the actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment includes: Obtaining the actual position at the (t + 1)-th moment and the actual current at the (t + 1)-th moment of the simulation turntable according to the motion result of the simulation turntable at the (t + 1)-th moment; Differentiating the actual position at the (t + 1)-th moment to obtain the actual rotation speed at the (t + 1)-th moment.
10. A simulation turntable control device based on the fusion of data-driven and active disturbance rejection, characterized in that, Including: A first output module, configured to input the desired position at the (t + 1)-th moment and the actual position at the t-th moment of the simulation turntable into the position controller of the simulation turntable, and output the desired rotation speed at the (t + 1)-th moment of the simulation turntable, where t is an integer greater than 0; A second output module, configured to input the desired rotation speed at the (t + 1)-th moment and the actual rotation speed at the t-th moment of the simulation turntable into the active disturbance rejection controller of the simulation turntable, and output the desired current at the (t + 1)-th moment of the simulation turntable; A third output module, configured to input the actual rotation speed at the t-th moment, the desired rotation speed at the t-th moment, the actual rotation speed at the (t - 1)-th moment, and the desired rotation speed at the (t - 1)-th moment of the simulation turntable into a pre-trained control network, and output the compensation current at the (t + 1)-th moment of the simulation turntable, where the control network includes a behavior network and an evaluation network, and the behavior network outputs the compensation current at the (t + 1)-th moment according to the evaluation value of the evaluation network; A first control module, configured to control the motion of the simulation turntable according to the desired current at the (t + 1)-th moment, the compensation current at the (t + 1)-th moment, and the actual current at the t-th moment of the simulation turntable.
Citation Information
Patent Citations
Rotary table servo system control method based on deep reinforcement learning
CN118655764A
Semi-physical simulation system and method for gliding extended-range guided aircraft
CN119414729A
Method and device for generating aircraft attitude control model, equipment and medium
CN119645096A
Method and apparatus for controlling tunnel ventilation
JP2007100503A
Servo control device
WO2020217282A1