Plasma virtual control surface and control mechanism control distribution method for reentry flight of high-speed aircraft

By using reinforcement learning algorithms to design control distribution models during high-speed aircraft reentry flight, and processing nonlinear characteristics of plasma and aerodynamic characteristics, the control distribution problem during high-speed aircraft reentry flight is solved, and more stable attitude control is achieved.

CN120215549AActive Publication Date: 2025-06-27AIR FORCE UNIV PLA

Patent Information

Application Number
CN202510152209.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-27
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

During the re-entry process of high-speed aircraft, how to reasonably allocate control means on the basis of considering the nonlinear characteristics of plasma excitation power and control output, as well as the nonlinear and uncertainty of aerodynamic characteristics, and effectively deal with the control problems of the transition process between mode switching.

Method used

A plasma virtual rudder surface and control distribution method for high-speed aircraft reentry flight is adopted, and the control distribution model of plasma active flow control system, a pneumatic rudder surface control mechanism and RCS reverse thrust control system is designed.

Benefits of technology

The attitude control of the high-speed aircraft during the re-entry flight stage is realized, the stability of the control system is improved, and the control means can be reasonably allocated according to different flight stages, solving the control problem of the mode switching transition process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215549A_ABST
    Figure CN120215549A_ABST
Patent Text Reader

Abstract

The invention provides a plasma virtual control surface and control mechanism control distribution method for reentry flight of a high-speed aircraft. The method comprises the following steps: firstly, determining that the aircraft enters a reentry flight stage; secondly, the needed torque is calculated according to an attitude controller; then designing control distribution models of a plasma active flow control system, a pneumatic control surface control mechanism and an RCS backstepping control system based on a depth deterministic strategy gradient reinforcement learning algorithm; and the high-speed aircraft completes reentry flight based on the control distribution parameters obtained by the control distribution model. Auxiliary control is achieved through the plasma jet, and a new way is provided for accurate control over the aircraft. According to the method, control means can be reasonably distributed according to the early stage, the middle stage and the later stage of reentry flight, and relevant parameters of all control systems are determined. The reinforcement learning algorithm is used for processing the nonlinear characteristics of the input power and control output of the plasma exciter and the nonlinearity and uncertainty of aerodynamic characteristics in an attitude control model, the stability of a control system can be improved, and the method has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of reentry flight control allocation for high-speed aircraft, and particularly to a control allocation method for a plasma virtual rudder surface and a control mechanism for high-speed aircraft reentry flight. Background Art

[0002] The research on high-speed aircraft reentry flight is an important means to promote the progress of aerospace technology and expand the understanding of nature. However, high-speed aircraft reentry flight also faces many problems, such as extremely high temperatures posing severe challenges to thermal protection technology, complex aerodynamic environments increasing the difficulty of navigation and control, and high R & D costs also restricting its development.

[0003] Plasma jets can assist in controlling high-speed aircraft. By adjusting the jet parameters, the flow field around the aircraft can be changed to achieve attitude adjustment. During flight, it can enhance the stability and maneuverability of the aircraft to cope with complex flight environments, providing a new way for the precise control of high-speed aircraft to improve its performance. However, the introduction of auxiliary control will inevitably bring control allocation problems. In the early, middle, and late stages of reentry flight, the required control means need to be controlled through different working modes, but the control during the transition process between mode switches is difficult, and it is difficult to accurately control the activation timing of various control means. In addition, there are large non-linear characteristics between the excitation power of the plasma and the control output, and the attitude control model is also a simplified uncertain model, which itself has strong non-linearity with respect to aerodynamic characteristics. Summary of the Invention

[0004] Technical Problems to be Solved

[0005] How to reasonably allocate control means on the basis of considering the non-linear characteristics of plasma excitation power and control output, as well as the non-linearity and uncertainty of aerodynamic characteristics, and effectively address the control challenges during the transition process between mode switches is the technical problem to be solved by the present invention.

[0006] In view of the above technical problems, the present invention proposes a control allocation method for a plasma virtual rudder surface and a control mechanism for high-speed aircraft reentry flight. By using a reinforcement learning algorithm to handle the non-linear relationships and uncertainty effects therein, the attitude control of high-speed aircraft during the reentry flight stage is realized, and the stability of the control system is improved.

[0007] The technical solution of the present invention is as follows:

[0008] A control allocation method for a plasma virtual rudder surface and a control mechanism for high-speed aircraft reentry flight, comprising the following steps:

[0009] Step 1: Determine the reentry flight phase entered by the aircraft according to the flight state parameters; the reentry flight phase is divided into three stages: the early stage of reentry flight, the middle stage of reentry flight, and the late stage of reentry flight;

[0010] Step 2: Solve the desired moment M according to the aircraft attitude angular velocity;

[0011] Step 3: Design the control allocation models of the plasma active flow control system, the aerodynamic control surface manipulation mechanism, and the RCS reverse thrust control system based on the deep deterministic policy gradient reinforcement learning algorithm;

[0012] Step 4: The aircraft completes the reentry flight based on the control allocation parameters obtained from the control allocation model.

[0013] Further, in the early stage of reentry flight, the flight control actuator consists only of the RCS reverse thrust system;

[0014] In the middle stage of reentry flight, the flight control actuator consists of a hybrid structure of three different types: aerodynamic control surfaces, the RCS reverse thrust control system, and the plasma active flow control system;

[0015] In the late stage of reentry flight, the flight control actuator consists of aerodynamic control surfaces and the plasma active flow control system.

[0016] Further, in Step 2, according to the formula

[0017]

[0018] Solve the desired moment M, where ω = [p, q, r] T is the representation form of the aircraft attitude angular rate vector, and I is the inertia tensor matrix.

[0019] Further, the process of establishing the control allocation model in Step 3 is as follows:

[0020] Set the agent structure: The agent is a policy model based on the deep deterministic policy gradient;

[0021] Set the environmental state: The environmental state includes the pitch angle, yaw angle, roll angle of the aircraft, as well as the flight speed, angle of attack, and flight altitude;

[0022] Set the agent action: The agent action is the control parameters of the plasma active flow control system, the aerodynamic control surface manipulation mechanism, and the RCS reverse thrust control system;

[0023] Set the state transition process: The state transition process is obtained based on the aircraft six-degree-of-freedom model;

[0024] Set up a reward mechanism: The reward mechanism is calculated according to different flight phases, taking the negative of the absolute value of the difference between the control moment generated by the output action and the desired moment as the first reward value, and taking the negative of the absolute value of the difference between the control parameters of the plasma active flow control system, the aerodynamic rudder surface control mechanism, and the RCS thrust reverse control system and their physical constraints as the remaining reward values;

[0025] Set up an experience pool; and use the temporal difference as the model objective function.

[0026] Furthermore, the agent is composed of a multi-layer fully connected network, and the input and output dimensions are determined by the environmental state and the agent's actions respectively. The intermediate hidden layer is set as a multi-layer fully connected neural network.

[0027] Furthermore, the agent's actions are the control parameters of the plasma active flow control system, the aerodynamic rudder surface control mechanism, and the RCS thrust reverse control system; among them, the control parameters of the plasma active flow control system include the input power of the plasma actuators in each channel w = [wp1,..., wp n T , where wp1 is the input power of the plasma actuator in the first channel of the plasma active flow control system; the control parameters of the aerodynamic rudder surface control mechanism include the deflection angles of each aerodynamic rudder surface δ = [δ1,..., δ m T , where δ1 is the deflection angle of the first aerodynamic rudder surface in the aerodynamic rudder surface control mechanism; the control parameters of the RCS thrust reverse control system include the proportionality coefficients of each nozzle P = [p1,..., p l T , where p1 is the proportionality coefficient of the first nozzle in the thrust reverse control system.

[0028] Furthermore, the calculation formula of the reward mechanism is:

[0029] Reward = Reward1 + Reward2 + Reward3 + Reward4

[0030] Reward1 = -‖M - M action *mask i ‖, i = {1, 2, 3}

[0031] mask1 = {1 1*n , 0 1*m , 0 1*l}

[0032] mask2 = {1 1*n , 1 1*m , 1 1*l}

[0033] mask3 = {0 1*n ​​​,1 1*m ,1 1*l}

[0034] Reward2 = -‖δ - D A ‖

[0035] Reward3 = -‖P - D R ‖

[0036] Reward4 = -‖w - D pl ‖

[0037] Where Reward is the reward value, Reward1 is the first reward value, Reward2 is the reward value corresponding to the pneumatic rudder surface control mechanism, Reward3 is the reward value corresponding to the RCS reverse thrust control system, Reward4 is the reward value corresponding to the plasma active flow control system, mask1 is the control state vector in the early stage of reentry flight, mask2 is the control state vector in the middle stage of reentry flight, mask3 is the control state vector in the late stage of reentry flight, D A , D R and D pl are the physical constraints of the pneumatic rudder surface deflection angle, the RCS reverse thrust control nozzle proportional coefficient, and the plasma actuator input power respectively, M action = B a U T , where B a is the control allocation matrix of the plasma active flow control system, the pneumatic rudder surface control mechanism, and the RCS reverse thrust control system, and U is the agent action.

[0038] In addition, the present invention also proposes an electronic device, including a processor and a memory for storing a program, the program includes instructions, and the instructions, when executed by the processor, cause the processor to execute the above method. And a computer-readable storage medium for storing a program is proposed, the program includes instructions, and the instructions, when executed by the processor of an electronic device, cause the electronic device to execute the above method.

[0039] Advantageous Effects

[0040] A plasma virtual rudder surface and control mechanism control allocation method for high-speed aircraft reentry flight proposed by the present invention uses plasma jets to achieve auxiliary control, providing a new way for the precise control of aircraft. This method can reasonably allocate control means according to the early, middle, and late stages of reentry flight, and clarify the relevant parameters of each control system. Using the reinforcement learning algorithm to deal with the non-linear characteristics of the plasma actuator input power and control output, as well as the non-linearity and uncertainty of the aerodynamic characteristics in the attitude control model, can improve the stability of the control system and has good application prospects.

[0041] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0043] Figure 1 is the design flow chart of the present invention.

[0044] Figure 2 is the training framework based on DDPG reinforcement learning involved in the present invention.

[0045] Figure 3 is the block diagram of the reentry flight phase control system of the hypersonic vehicle carrying the DDPG policy network involved in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Embodiments of the present invention will be described in detail below. The embodiments are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.

[0047] This embodiment mainly addresses the control allocation problem of the plasma active flow control system (i.e., plasma virtual rudder surface), the aerodynamic rudder surface control mechanism, and the RCS thrust reverse control system during the reentry flight process of a hypersonic vehicle, and proposes a control allocation method for the plasma virtual rudder surface and the control mechanism for a hypersonic vehicle during reentry flight. First, determine the reentry flight phase entered by the vehicle, secondly, calculate the required torque according to the attitude controller, then design a control allocation model for the plasma active flow control system, the aerodynamic rudder surface control mechanism, and the RCS thrust reverse control system based on the deep deterministic policy gradient (DDPG) reinforcement learning algorithm, and finally, the hypersonic vehicle completes the reentry flight based on the control allocation parameters obtained from the control allocation model; as Figure 1 shown, the specific steps are as follows:

[0048] Step 1: Determine the reentry flight phase entered by the vehicle according to the flight state parameters; the reentry flight phase is divided into three stages: the early stage of reentry flight, the middle stage of reentry flight, and the late stage of reentry flight.

[0049] In the early stage of reentry flight, the air is thin, resulting in a small dynamic pressure, and the control efficiency of the aerodynamic rudder surface control mechanism is low. In order to stably control the hypersonic vehicle at the initial stage of the reentry phase, the aerodynamic rudder surface is not used, and the flight control actuator is only composed of the RCS thrust reverse system.

[0050] In the middle stage of reentry flight, the increasing dynamic pressure makes the control efficiency of the pneumatic operating mechanism continuously increase. Therefore, in this stage, a hybrid heterogeneous execution structure is adopted to perform composite control on the aircraft. At this time, the flight control actuator consists of three heterogeneous hybrid structures: pneumatic control surfaces, RCS thrust reverse control systems, and plasma active flow control systems.

[0051] In the later stage of reentry flight, the gradually increasing dynamic pressure makes the pneumatic control surface operating mechanism capable of completing the control of the aerospace high-speed aircraft. Therefore, it is not necessary to perform control allocation on the RCS thrust reverse control system. At this time, the flight control actuator consists of pneumatic control surfaces and plasma active flow control systems.

[0052] Step 2: Calculate the required torque according to the dynamic equation of the angular rate loop in the attitude controller of the aerospace high-speed aircraft.

[0053] Using the attitude angular velocity of the high-speed aircraft, according to the formula

[0054]

[0055] Calculate the required torque M, that is, the desired torque, where ω = [p, q, r] T is the representation form of the aircraft attitude angular rate vector, and I is the inertia tensor matrix.

[0056] Step 3: Design the control allocation models for the plasma active flow control system, pneumatic control surface operating mechanism, and RCS thrust reverse control system based on the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm.

[0057] The reinforcement learning training environment can be built based on a well-known simulator. Referring to Figure 2 , the training framework based on DDPG reinforcement learning involved in the present invention includes four parts: environment, policy network, evaluation network, and experience pool. The environment provides the basic conditions for the operation of the high-speed aircraft. The policy network outputs policies to guide the flight of the high-speed aircraft. The evaluation network is used to score the actions of the policies. The policy network stores the current state, action, transferred state, reward, and discount obtained at each step of execution into the experience pool. Every certain period, soft updates are realized between the policy network and the target policy network, and between the action network and the target action network respectively. Then, the weight parameters are copied between the policy network and the action value network using a well-known AC reinforcement learning framework. The design of its related model composition is as follows:

[0058] (1) Set the agent structure

[0059] The agent is the policy model of DDPG, which is composed of a multi-layer fully connected network. The input and output dimensions are determined by the environmental state and the agent's actions respectively. The middle hidden layer is set as two layers of fully connected neural networks, with 128 * 128 nodes in each layer.

[0060] (2) Set the environmental state

[0061] The environmental state s includes the flight attitude data of the aircraft, including pitch angle, yaw angle, roll angle, as well as flight speed, angle of attack, and flight altitude.

[0062] (3) Set the agent action

[0063] The agent action is the control parameters of the plasma active flow control system, the pneumatic rudder surface control mechanism, and the RCS thrust reverse control system. Among them, the control parameters of the plasma active flow control system include the input power w of each channel plasma actuator = [wp1,…wp n T , where wp1 is the input power of the first channel plasma actuator in the plasma active flow control system; the control parameters of the pneumatic rudder surface control mechanism include the deflection angles δ of each pneumatic rudder surface = [δ1,…,δ m T , where δ1 is the deflection angle of the first pneumatic rudder surface in the pneumatic rudder surface control mechanism; the control parameters of the RCS thrust reverse control system include the proportionality coefficients P of each nozzle = [p1,…,p l T , where p1 is the proportionality coefficient of the first nozzle in the thrust reverse control system; comprehensively, the agent action is:

[0064] U = [w; δ; P]

[0065] In this embodiment, taking the aircraft entering the mid-stage of reentry flight as an example, this stage adopts pneumatic rudders (8 in number) + plasma active flow control (2 in number) + RCS (4 in number), then the agent action is:

[0066] U = [wp1,wp2; δ1,…,δ8; p1,…,p4]

[0067] Each action component is directly associated with the input of the corresponding mechanism, which is convenient for online learning.

[0068] (4) Set the state transition process

[0069] The state transition is a continuous transformation process during the movement of the agent in the environment. The state transition of the present invention is obtained through a six-degree-of-freedom model of a high-speed aircraft. This model usually involves the translational dynamics of the aircraft's center of mass and the rotational dynamics around the center of mass. Usually, the lateral force can be ignored compared with the lift and drag. Therefore, the following nonlinear motion equations are cited in the design of the guidance method:

[0070]

[0071] ​​​Wherein, R is the distance from the aircraft to the earth's center, V is the aircraft speed, σ is the bank angle of the aircraft, θ and φ are the longitude and latitude respectively, γ and ψ are the flight path inclination angle and the course angle of the aircraft respectively, g is the acceleration due to gravity, L is the lift force received by the aircraft, D is the drag force received by the aircraft, ω E is the angular velocity of the rotation of the earth coordinate system (since the aircraft flies at a high altitude, the rotation of the earth coordinate system needs to be considered), and m is the mass of the aircraft.

[0072] Calculations are performed using the output strategy of the 6-degree-of-freedom model following agent. Under each strategy, the corresponding state components can be output, thereby providing the environmental state under each iteration process.

[0073] (5) Set up the reward mechanism

[0074] The reward mechanism scores according to the environmental state of the aircraft and the actions output by the agent to evaluate the quality of the strategy. The actions of the present invention involve three parts: the plasma active flow control system, the aerodynamic control surface operating mechanism, and the RCS reverse thrust control system. Therefore, the rewards are jointly weighted by these three parts.

[0075] After the attitude controller calculates the required torque M, it is necessary to determine the actual actuator composition according to the flight mode and then determine the actual flight torque. The entire flight process is divided into three stages: the early stage of reentry flight, the middle stage of reentry flight, and the late stage of reentry flight. The actuator compositions are pure RCS reverse thrust control, aerodynamic control surface + plasma active flow control + RCS reverse thrust control, and aerodynamic control surface + plasma active flow control respectively. Therefore, the present invention designs a mask array to calculate the torque according to the priority. It is divided into three groups, corresponding to the three flight stages respectively. After the strategy model outputs an action, different masks need to be used according to the flight stage to calculate the control torque, and the absolute value of the difference between it and the desired torque is taken as the first reward value, and the absolute value of the difference between the control parameters of the plasma active flow control system, the aerodynamic control surface operating mechanism, and the RCS reverse thrust control system and the physical constraints is taken as the remaining reward values: specifically as follows:

[0076] Reward = Reward1 + Reward2 + Reward3 + Reward4

[0077] Reward1 = -‖M - M action *mask i ‖, i = {1, 2, 3}

[0078] mask1 = {1 1*n , 0 1*m , 0 1*l}

[0079] mask2 = {1 1*n , 11*m , 1 1*l}

[0080] mask3 = {0 1*n , 1 1*m , 1 1*l}

[0081] Reward2 = -‖δ - D A ‖

[0082] Reward3 = -‖P - D R ‖

[0083] Reward4 = -‖w - D pl ‖

[0084] Wherein, D A , D R and D pl are respectively the physical constraints of the deflection angle of the pneumatic rudder surface, the proportional coefficient of the RCS reverse thrust control nozzle, and the input power of the plasma actuator. M action = B a U T , where B a is the control allocation matrix of three different types of actuators, that is, the matrix mapping the actual output and torque, which is determined by the actuator composition and internal principle, and U is the agent action, that is, the actual output of the actuator.

[0085] (6) Set up the experience pool

[0086] In reinforcement learning, the experience pool plays an important role. It can store the data generated by the interaction between the agent and the environment, and then use it through random sampling, which can break the data correlation and improve the data utilization rate. The experience pool storage tuple of the present invention contains <s, a, s ′ , r, τ>, where s is the environmental state, s ′ is the new state obtained through the state transition equation, a is the output value of the policy model, r is the immediate reward of the corresponding policy model, and τ is the discounted return. After batch extracting the data in the experience pool, the temporal difference (TD-error) is used as the model objective function:

[0087] TD = r + ρV(s ′ ) - V(s)

[0088] Wherein, V(.) represents the state value, and ρ is the set coefficient. Then, train the neuron weights in the network according to the gradient descent backpropagation, and store the weights once every H iterations.

[0089] Regarding the extraction of experience pool samples, the present invention adopts the well-known prioritized experience pool algorithm to sort the experience pool samples according to the TD difference, and then selects, for example, 64 or 128 samples from the samples with higher rankings for extraction. In this way, the convergence speed and reward value can be effectively improved, and the learning ability of the intelligent agent can be significantly enhanced.

[0090] The final control allocation model is obtained by online training of the control allocation model constructed in the above process.

[0091] Step 4: The hypersonic vehicle completes the reentry flight based on the control allocation parameters obtained from the control allocation model.

[0092] Reference Figure 3 , first, the attitude controller calculates the required torque, and then the deflection angle of the aerodynamic rudder surface under the corresponding policy control, the input power of the plasma jet exciter, and the switching state of the nozzle of the RCS thrust control system are output through the policy network of the trained DDPG intelligent agent, so as to control the attitude stability of the hypersonic vehicle and realize the whole process control in the early, middle, and late stages of the reentry flight.

[0093] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and purposes of the present invention.

Claims

1. A method for controlling and allocating plasma virtual control surfaces and control mechanisms for high-speed aircraft reentry flight, characterized in that: The following steps are involved: Step 1: Determine the reentry flight phase that the aircraft enters according to the flight state parameters; the reentry flight phase is divided into three phases: early reentry flight phase, mid reentry flight phase and late reentry flight phase; Step 2: Calculate the expected torque M according to the aircraft attitude angular velocity; Step 3: Design the control allocation model of plasma active flow control system, pneumatic control surface control mechanism and RCS thrust reverse control system based on deep deterministic policy gradient reinforcement learning algorithm; Step 4: The aircraft completes reentry flight based on the control allocation parameters obtained by the control allocation model.

2. The method for controlling and allocating plasma virtual control surfaces and operating mechanisms for high-speed aircraft reentry flight according to claim 1, characterized in that: In the early stage of the reentry flight, the flight control actuator is composed only of the RCS reverse thrust system; In the mid-stage of the reentry flight, the flight control actuator is composed of three heterogeneous hybrid structures: aerodynamic control surfaces, RCS reverse thrust control system, and plasma active flow control system; In the later stage of the reentry flight, the flight control actuator is composed of aerodynamic control surfaces and a plasma active flow control system.

3. The method for controlling and allocating plasma virtual control surfaces and maneuvering mechanisms for high-speed aircraft reentry flight according to claim 1, characterized in that: In step 2, according to the formula Solve for the desired torque M, where ω = [p, q, r] T It is the vector representation of the vehicle's attitude angular rate, and I is the inertia tensor matrix.

4. The method for controlling and allocating plasma virtual control surfaces and operating mechanisms for high-speed aircraft reentry flight according to claim 2, characterized in that: The process of establishing the control allocation model in step 3 is: Setting the agent structure: the agent is a policy model based on deep deterministic policy gradient; Set the environmental status: The environmental status includes the pitch angle, yaw angle, roll angle, flight speed, angle of attack and flight altitude of the aircraft: Setting agent actions: the agent actions are control parameters of the plasma active flow control system, the pneumatic control surface control mechanism, and the RCS reverse thrust control system; Setting a state transfer process: the state transfer process is obtained based on a 6-DOF model of the aircraft; Setting a reward mechanism: The reward mechanism is calculated according to different flight phases, and the absolute value of the difference between the control torque generated by the output action and the expected torque is taken as the first reward value, and the absolute value of the difference between the control parameters of the plasma active flow control system, the pneumatic control surface control mechanism and the RCS reverse thrust control system and the physical constraints is taken as the remaining reward values; Set up experience pool; And the timing difference is used as the model objective function.

5. The method for controlling and allocating plasma virtual control surfaces and operating mechanisms for high-speed aircraft reentry flight according to claim 4, characterized in that: The agent is composed of a multi-layer fully connected network, the input and output dimensions are determined by the environment state and the agent action respectively, and the middle hidden layer is set as a multi-layer fully connected neural network.

6. The method for controlling and allocating plasma virtual control surfaces and control mechanisms for high-speed aircraft reentry flight according to claim 4, characterized in that: The intelligent body action is the control parameters of the plasma active flow control system, the pneumatic control surface control mechanism and the RCS reverse thrust control system; wherein the control parameters of the plasma active flow control system include the input power w of each channel plasma actuator = [wp1, ...wp n ] T , wp1 is the input power of the first channel plasma actuator in the plasma active flow control system; the control parameters of the pneumatic control surface control mechanism include the deflection angle of each pneumatic control surface δ = [δ1,…,δ m ] T , δ1 is the deflection angle of the first pneumatic control surface in the pneumatic control surface control mechanism; the control parameters of the RCS reverse thrust control system include the proportional coefficients of each nozzle P = [p1, ..., p l ] T , p1 is the proportional coefficient of the first nozzle in the reverse thrust control system.

7. The method for controlling and allocating plasma virtual control surfaces and maneuvering mechanisms for high-speed aircraft reentry flight according to claim 4, characterized in that: The calculation formula of the reward mechanism is: Reward=Reward1+Reward2+Reward3+Reward4 Reward1=-||M-M action *mask i ||,i={1,2,3} mask1={1 1*n ,0 1*m ,0 1*l } mask2={1 1*n ,1 1*m ,1 1*l } mask3={0 1*n ,1 1*m ,1 1*l } Reward2=-||δ-D A || Reward3=-||P-D R || Reward4=-||w-D pl || Where Reward is the reward value, Reward1 is the first reward value, Reward2 is the reward value corresponding to the aerodynamic control surface control mechanism, Reward3 is the reward value corresponding to the RCS reverse thrust control system, Reward4 is the reward value corresponding to the plasma active flow control system, mask1 is the control state vector in the early stage of reentry flight, mask2 is the control state vector in the middle stage of reentry flight, mask3 is the control state vector in the late stage of reentry flight, D A , D R and D pl are the physical constraints of the aerodynamic control surface deflection angle, RCS reverse thrust control nozzle proportional coefficient and plasma actuator input power, M action =B a U T , where B a is the control allocation matrix of the plasma active flow control system, the pneumatic control surface control mechanism and the RCS thrust reverse control system, and U is the intelligent agent action.

8. An electronic device comprising a processor and a memory storing a program, wherein the program comprises instructions, characterized in that: When the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a program, the program comprising instructions, characterized in that: When the instructions are executed by a processor of an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Plasma immersion ion implantation with highly uniform chamber seasoning process for a toroidal source reactor

    CN101308784A

  • Closed loop simulation system suitable for controlling attitude of reentry vehicle

    CN103488814A

  • Device for relieving influence on high-speed aircraft reentry communication by space plasma

    CN103796407A

  • Surface plasmon DC pulse attitude control and propulsion assisting system for hypersonic aircraft

    CN109850143A

  • Aircraft electromagnetic characteristic modeling method in high-speed flight state

    CN112257350A

Cited By

  • Aircraft control method and device based on deep reinforcement learning, equipment and medium

    CN121187203A