A morphing aircraft maneuver trajectory design method based on a SAC algorithm, an electronic device, and a storage medium

Through the intelligent agent neural network model based on the SAC algorithm, the maneuvering trajectory design problem of hypersonic aircraft in complex environments was solved, more flexible and intelligent aircraft maneuvering trajectory adjustment was achieved, and the aircraft's maneuverability and environmental adaptability were improved.

CN119828724BActive Publication Date: 2025-10-17HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411937915.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-10-17
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively design maneuvering trajectories for hypersonic aircraft in complex and dynamic flight environments, especially for aircraft combined with variable wing sweep angles, resulting in insufficient flexibility and intelligence in the selection and execution of maneuvering schemes.

Method used

An intelligent agent neural network model based on the SAC algorithm is adopted to design the aircraft maneuver trajectory by constructing a strategy network and a value network, combining Markov process modeling and adaptive exploration mechanism, and adjusting the wing sweep angle and expected state in real time to optimize the maneuver effect.

Benefits of technology

It improves the training stability and learning efficiency of the aircraft in the continuous action space, achieves more accurate and real-time flight state adjustment, enhances maneuverability, and adapts to complex and dynamic flight environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119828724B_ABST
    Figure CN119828724B_ABST
Patent Text Reader

Abstract

A morphing aircraft maneuver trajectory design method based on SAC algorithm, electronic equipment and storage medium belong to the technical field of aircraft trajectory design. In order to improve the training stability and learning efficiency of aircraft in continuous action space, the present application includes constructing an agent neural network model for morphing aircraft maneuver trajectory design; The maneuver problem of the morphing aircraft under the interception condition is modeled by Markov process, and the simulation environment is designed; Collect flight data for the agent neural network model for morphing aircraft maneuver trajectory design; The SAC algorithm is used to train the agent neural network model for morphing aircraft maneuver trajectory design, and an adaptive exploration mechanism is added in the training process, the strength of exploration is adjusted according to the current flight state and environmental conditions, and the trained agent neural network model for morphing aircraft maneuver trajectory design is obtained; Real-time trajectory adjustment and morphing decision in flight control loop.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of aircraft trajectory design, and particularly relates to a morphing aircraft maneuver trajectory design method based on a SAC algorithm, an electronic device and a storage medium. BACKGROUND

[0002] With the development of modern warfare technology, hypersonic aircraft has been widely used in the military field. Among them, the morphing aircraft, as a kind of aircraft type with adjustable wing surface sweep angle, can flexibly adjust the flight state in different flight stages, so as to realize higher flight performance. However, the sequential maneuver trajectory design for multiple interceptors is still a challenging task because it involves numerous parameters and needs to be adjusted in real time according to complex flight environment and enemy trajectory.

[0003] Traditional maneuver trajectory design mainly relies on artificially preset flight schemes and fixed decision logic. Although these methods work well in some scenarios, they often seem inadequate when dealing with complex, dynamic and uncertain flight environments. In this case, the selection and execution of maneuver schemes need more flexible and intelligent methods.

[0004] In recent years, reinforcement learning as a self-adaptive and autonomous learning method has been widely studied and applied in many fields. Among them, the SAC algorithm as a variant of reinforcement learning has attracted widespread attention for its stability and efficiency in continuous action space. However, the application of reinforcement learning and SAC algorithm to the maneuver trajectory design of hypersonic aircraft, especially combined with the characteristics of aircraft with variable wing surface sweep angle, is still relatively rare.

[0005] Due to the complexity of hypersonic aircraft flight state, the variability of interceptor information, and the real-time requirements of maneuver trajectory, traditional flight control and trajectory design methods often perform poorly in these aspects. Therefore, it is of great theoretical significance and practical value to study a method that can combine the characteristics of hypersonic aircraft, effectively utilize reinforcement learning technology, especially SAC algorithm, to realize maneuver trajectory design. This not only helps to improve the maneuverability of aircraft, but also helps to further combine and develop aircraft guidance technology and reinforcement learning algorithm. SUMMARY

[0006] The problem to be solved by the present application is to improve the training stability and learning efficiency of aircraft in continuous action space, and a morphing aircraft maneuver trajectory design method based on SAC algorithm, an electronic device and a storage medium are proposed.

[0007] To achieve the above-mentioned purpose, the technical scheme is as follows:

[0008] A morphing aircraft maneuver trajectory design method based on SAC algorithm, comprising the following steps:

[0009] S1. Constructing an agent neural network model for morphing aircraft maneuver trajectory design, including a policy network π for generating actions and a value network Q for estimating the value of each state-action pair;

[0010] S2. Modeling the Markov process of the morphing aircraft maneuver problem under interception conditions, designing the simulation environment, including designing the aircraft state, action, reward function, and aircraft state transition relationship;

[0011] S3. Simulating various flight scenarios in the simulation environment designed in step S2, including different initial flight conditions, states of hostile interceptors, and different target points, collecting flight data for the agent neural network model for morphing aircraft maneuver trajectory design;

[0012] S4. Training the agent neural network model for morphing aircraft maneuver trajectory design using the SAC algorithm, adding an adaptive exploration mechanism during training to adjust the strength of exploration according to the current flight state and environmental conditions, obtaining the trained agent neural network model for morphing aircraft maneuver trajectory design;

[0013] S5. Deploying the trained agent neural network model for morphing aircraft maneuver trajectory design in the onboard computer for real-time trajectory adjustment and morphing decision-making in the flight control loop.

[0014] Further, the specific implementation method of step S1 includes the following steps:

[0015] S1.1. Constructing a policy network π for generating actions; the policy network includes an input layer, a hidden layer, and an output layer;

[0016] The size of the input layer is set to the dimension of the aircraft state and interceptor information, including height, speed, heading angle, 3 parameters of aircraft state, and position, speed, model, 3 parameters of interceptor information, the size of the input layer is set to 6 layers;

[0017] The hidden layer is set to two hidden layers, each hidden layer includes 256 neurons;

[0018] The size of the output layer is set to the dimension of the action, including the wing surface sweep angle and the expected position, expected speed, 3 parameters, the size of the output layer is set to 3 layers;

[0019] The expression of the policy network is:

[0020] a=π(s) (1)

[0021] Wherein, a is action, s is aircraft state, and represents a nonlinear function relationship expressed by a policy network.

[0022] S1.2. Construct a value network Q for estimating the value of each state-action pair. The value network Q includes a value network and a target value network, which have the same structure and are used to estimate the value of each state-action pair, including an input layer, a hidden layer, and an output layer.

[0023] The size of the input layer is set to the dimension of the aircraft state, interceptor information, and action. The aircraft state includes three parameters: height, speed, and heading angle. The interceptor information includes three parameters: position, speed, and model. The action includes three parameters: wing surface sweep angle, desired position, and desired speed. The size of the input layer is set to 9 layers.

[0024] The hidden layer is set to two hidden layers, each including 256 neurons.

[0025] The size of the output layer is set to the value of the state-action pair. The size of the output layer is set to 1 layer.

[0026] The expression of the value network is:

[0027]

[0028] Wherein, Q is the value of the optimal problem, is a nonlinear function relationship expressed by a value network, i.e., an optimality indicator.

[0029] S1.3. Set the calculation rules of the policy network and the value network: calculated by the forward calculation method of the neural network, set the output of the lth layer as a (l) The calculation of each layer is divided into two steps: linear transformation and nonlinear activation function, expressed as:

[0030] z (l) =W (l) a (l-1) +b (l) (3)

[0031] a (l) =σ (l) (z (l) ) (4)

[0032] Wherein, W (l) is the weight matrix of the lth layer, b (l) is the bias vector, and σ (l) is the activation function of the lth layer, set as σ (l) (x) = tanh(x).

[0033] The expression of the output of the policy network is:

[0034] a = a (2) (5)

[0035] The output of the value network is expressed as:

[0036] Q = a (2) (6).

[0037] Further, the specific implementation method of step S2 includes the following steps:

[0038] S2.1. Designing the aircraft state: the aircraft state includes the current height, speed, heading angle parameters, interceptor information, and the expression is:

[0039] s = [h, v, θ, d] (7)

[0040] wherein h, v, θ are the height, speed and heading angle of the aircraft at time t, respectively;

[0041] Setting the range of aircraft state parameters: the height h is 5000 meters-20000 meters; the speed is 1000m / s-3000m / s; the heading angle is 0 degrees-360 degrees;

[0042] d is the interceptor information, and the expression is:

[0043] d = [p I ,v I ,L I ] (8)

[0044] wherein d represents the interceptor information at time t, p I , v I , L I represent the position, speed and model of the interceptor at time t, respectively;

[0045] The range of interceptor information parameters: the position p I is 0 meters-50000 meters; the speed v I is 500m / s-1500m / s; the model L I is a discrete integer;

[0046] S2.2. Designing actions:

[0047] According to the state of the aircraft and the information of the interceptor, the action is decided through the strategy network, and the expression is:

[0048] a = π(s, d) (9)

[0049] wherein a is defined as:

[0050]

[0051] where Λ is the wing sweep angle, h c is the desired height, v c is the desired velocity;

[0052] S2.3. Designing the reward function:

[0053] The reward function R(s t ,a r,t ) is composed of several parts:

[0054] R(s t ,a t ) = w1R target - w2R intercept - w3R energy + w4R stability (11)

[0055] where R target is the target reward, R intercept is the interception penalty, R energy is the energy penalty, and R stability is the stability reward;

[0056]

[0057] where d(p r,t ,p target ) is the Euclidean distance between the aircraft position and the target point, d(p r,t ,p b,t ) is the Euclidean distance between the aircraft position and the interceptor position, d thresh is the distance threshold for the interception penalty, s desired is the desired aircraft state;

[0058] R target represents the reward when the aircraft approaches the target region; R intercept represents the penalty for being hit: given when the aircraft is too close to the interceptor; R energy represents the penalty for consuming more energy when performing a maneuvering action, encouraging energy efficiency, R stability represents that the aircraft should be as close as possible to the original flight state when no maneuver is needed;

[0059] S2.4. State transition relationship:

[0060] Let the state at time t be s t , then the state at time t+1 is s t+1 = P(s t ,a), where the process used for training is a deterministic dynamics integration P, expressed as:

[0061]

[0062] where x e , y n and h are the x-axis displacement, y-axis displacement and z-axis displacement in the local East-North-Up coordinate system ENU, in which the x-axis points to the east of the earth origin, the y-axis points to the north of the earth origin, and the z-axis is determined by the right-hand rule; V, ξ, ψ are the speed, flight path angle and heading angle of the aircraft respectively, and g is the local standard gravity; L is the lift, and D is the drag;

[0063] The expressions of the lift and the drag are as follows:

[0064] L = c L (Λ)qS ref (14)

[0065] D = c D (Λ)qS ref (15)

[0066] where S ref is the reference area of the aircraft, q is the dynamic pressure, q = 1 / 2 ρ(h) V 2 , ρ(h) = ρ0e -βh , ρ(h) is the atmospheric density, ρ0 is the atmospheric density at sea level, β is a constant, c L (Λ) is the lift coefficient of the aircraft, which is expressed as a function of the sweepback angle, and c D (Λ) is the drag coefficient of the aircraft, which is also expressed as a function of the sweepback angle, both of which are obtained by interpolation from the aircraft aerodynamic parameter table.

[0067] Further, the specific implementation method of step S3 includes the following steps:

[0068] S3.1. Initialization: first, the weights of the policy network, the value network and the target value network are initialized, and the initial values are set to random numbers in the range of (0, 1) in the initialization stage;

[0069] S3.2. State and action sampling: in each iteration, the state s t at time t is sampled from the state space according to the current policy network, and the action a t at time t is sampled from the action space, and the expression is as follows:

[0070]

[0071] The sample flight data (s t , a t , r t , s t+1 ) is stored in the experience replay buffer.

[0072] Further, the specific implementation method of step S4 includes the following steps:

[0073] S4.1. Based on the SAC algorithm, set the agent in state s t Take action a t The reward observed after is R t , and the exploration weighting factor is β(s t ), then the expression of the improved exploration strategy is:

[0074]

[0075] Where ∈ is a random variable sampled from a Gaussian distribution, α is a temperature parameter for adjusting the balance between exploration and utilization; the exploration weighting factor β(s t ) is dynamically calculated based on the access frequency or variance of the value function of state s t ;

[0076] S4.2. Set the processing method of delayed reward: when facing delayed reward, use the improved reward credit distribution mechanism to set the adjusted future reward G t at time t, and use the improved time difference error to update the value function, the expression is:

[0077] G t = r t + γr t+1 + γ 2 r t+2 +... + γ n-1 r t+n-1 + γ n Q(s t+n , a t+n )(18)

[0078] δ t = G t - Q(s t , a t )(19)

[0079] Where γ is the discount factor, n is the number of steps, R t represents the reward at step t;

[0080] S4.3. Update the policy network, use δ t to update the policy network weight, the update steps are as follows:

[0081]

[0082] Where δ t is the time difference error;

[0083] S4.4. Update the value network by processing the sample flight data obtained in step S3 and calculating the target value y, which is expressed as:

[0084] y=r+γ(V(s)-αlogπ θ′ (a′|s′))(21)

[0085] in, is the value function;

[0086] Then use the mean square error to update the network weights, the expression is:

[0087]

[0088] in, Indicates the expectation of random variable b under sample a, φ 1,2 Represents the weights of the value network and target value network used by the SAC algorithm;

[0089] Use the soft update method to update the value network weight and target value network weight respectively. The expression is:

[0090] φ i ′←τφ i +(1-τ)φ i '(twenty three)

[0091] When i is 1, φ i ′ is the value network weight, when i is 2, φ i ′ is the target value network weight, τ is the soft update coefficient;

[0092] S4.5. Adjust and optimize algorithm parameters based on training results: Adjust the entropy weight α and soft update coefficient τ to optimize the comprehensive training index; the expression for adjusting α is:

[0093]

[0094] in, is the target entropy, which is usually set to the value of the negative action space dimension to encourage sufficient exploration; the entropy weight α is adjusted. If the agent exhibits over-exploration, that is, the strategy is too random, α is reduced; if the agent does not explore enough, α is increased; then τ is adjusted according to the performance of the agent in the environment to obtain a trained agent neural network model for maneuver trajectory design of deformable aircraft.

[0095] Furthermore, in step S5, network pruning and quantization techniques are applied to set the model of the agent to a size suitable for onboard computing power, with the goal of setting the model size to no more than 50MB.

[0096] An electronic device comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm.

[0097] A computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm.

[0098] The beneficial effects of the present application are:

[0099] The method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm proposes an improved SAC algorithm, which enhances the training stability and learning efficiency of the algorithm in the continuous action space by introducing an adaptive exploration strategy and a multi-task learning mechanism, and is suitable for the maneuvering trajectory design of a morphing aircraft. The adaptive exploration mechanism adjusts the exploration intensity according to the current flight state and environmental conditions. In critical maneuvers or near targets, the exploration intensity is automatically reduced to enhance the trajectory accuracy and operational stability. The exploration strategy uses real-time feedback of flight state and expected maneuver to dynamically adjust, to optimize the responsiveness and efficiency of the aircraft. An improved reward mechanism is added, which is specifically used to handle decisions related to delayed rewards. This mechanism optimizes flight strategies that may initially appear non-optimal but have better long-term effects by considering long-term benefits. Through analysis of historical data and prediction models, the algorithm can evaluate the potential effects of various operations to ensure optimality in a multi-step decision-making environment.

[0100] The method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm first trains an agent in a ground environment, enabling it to learn and make real-time decisions on the optimal wing sweep angle and desired position and speed of the aircraft with variable wing sweep angle. During the training process, the agent dynamically adjusts the wing sweep angle and flight desired state based on the flight state of the aircraft and the position, speed, and model information of the interceptor, to achieve optimal maneuvering effect.

[0101] The method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm combines the adaptive learning ability of reinforcement learning and the deformable characteristics of hypersonic aircraft, achieving more accurate and real-time adjustment of flight state, thereby significantly improving the maneuverability of the aircraft. Compared with traditional aircraft trajectory design methods, the present application can more flexibly cope with complex, dynamic and uncertain flight environments, providing a new solution for the field of morphing aircraft maneuvering trajectory design.

[0102] The method for designing the maneuvering trajectory of the morphing aircraft based on the SAC algorithm deploys the trained model in the on-board computer, and is used for real-time trajectory adjustment and morphing decision in the flight control loop. BRIEF DESCRIPTION OF DRAWINGS

[0103] Figure 1 The flow chart of the method for designing the maneuvering trajectory of the morphing aircraft based on the SAC algorithm;

[0104] Figure 2 The flow chart of the method for designing the maneuvering trajectory of the morphing aircraft based on the SAC algorithm;

[0105] Figure 3 The flow chart of the SAC algorithm training;

[0106] Figure 4 The comparison chart of the improved range of the random strategy (SAC);

[0107] Figure 5 The reward function change curve in the training;

[0108] Figure 6 The overload curve of the attack and defense sides obtained;

[0109] Figure 7 The aircraft-interceptor position curve obtained. DETAILED DESCRIPTION

[0110] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application, that is, the specific embodiments described are only part of the embodiments of the present application, but not all the specific embodiments. The components of the specific embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations, and the present application can also have other embodiments.

[0111] Therefore, the detailed description of the specific embodiments of the present application provided in the drawings below is not intended to limit the scope of the claimed present application, but only represents selected specific embodiments of the present application. Based on the specific embodiments of the present application, all other specific embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0112] In order to further understand the inventive content, characteristics and effects of the present application, the following specific embodiments are exemplified, and the accompanying drawings are Figure 1 -attached Figure 7 The detailed description is as follows:

[0113] Example 1:

[0114] A morphing aircraft maneuver trajectory design method based on SAC algorithm, comprising the following steps:

[0115] S1. Construct an agent neural network model for morphing aircraft maneuver trajectory design, including a policy network π for generating actions and a value network Q for estimating the value of each state-action pair;

[0116] Further, the specific implementation method of step S1 comprises the following steps:

[0117] S1.1. Construct a policy network π for generating actions; the policy network includes an input layer, a hidden layer and an output layer;

[0118] The size of the input layer is set to the dimension of the aircraft state and interceptor information, and the aircraft state includes three parameters of height, speed and heading angle, and the interceptor information includes three parameters of position, speed and model, and the size of the input layer is set to 6 layers;

[0119] The hidden layer is set to two hidden layers, each of which includes 256 neurons;

[0120] The size of the output layer is set to the dimension of the action, which includes three parameters of wing surface sweep angle, desired position and desired speed, and the size of the output layer is set to 3 layers;

[0121] The expression of the policy network is:

[0122] a=π(s) (1)

[0123] Where a is the action, s is the aircraft state, and π represents a nonlinear function relationship expressed by the policy network;

[0124] S1.2. Construct a value network Q for estimating the value of each state-action pair; the value network includes a value network and a target value network, and the structure of the value network and the target value network is exactly the same, for estimating the value of each state-action pair, including an input layer, a hidden layer and an output layer;

[0125] The size of the input layer is set to the dimension of the aircraft state, interceptor information and action, the aircraft state includes three parameters of height, speed and heading angle, the interceptor information includes three parameters of position, speed and model, and the action includes three parameters of wing surface sweep angle, desired position and desired speed, and the size of the input layer is set to 9 layers;

[0126] The hidden layer is set to two hidden layers, each including 256 neurons;

[0127] The output layer size is set to the value of state-action pair, and the size of the output layer is set to 1 layer;

[0128] The expression of the value network is:

[0129]

[0130] Wherein, Q is the value of the optimal problem, is a nonlinear function relationship expressed by the value network, that is, the optimality index;

[0131] S1.3. Set the calculation rules of the policy network and the value network: calculated by the forward calculation method of the neural network, set the output of the l-th layer to a (l) The calculation of each layer is divided into two steps: linear transformation and nonlinear activation function, expressed as:

[0132] z (l) = W (l) a (l-1) + b (l) (3)

[0133] a (l) = σ (l) (z (l) ) (4)

[0134] Wherein, W (l) is the weight matrix of the l-th layer, b (l) is the bias vector, σ (l) is the activation function of the l-th layer, set to σ (l) (x) = tanh(x);

[0135] The expression of the output of the policy network is:

[0136] a = a (2) (5)

[0137] The expression of the output of the value network is:

[0138] Q = a (2) (6)

[0139] S2. Model the maneuvering problem of the morphing aircraft under the interception condition as a Markov process, and design the simulation environment, including designing the aircraft state, action, reward function, and aircraft state transition relationship;

[0140] Further, the specific implementation method of step S2 includes the following steps:

[0141] S2.1. Design the aircraft state: The aircraft state includes the current height, speed, heading angle parameters, interceptor information, expressed as:

[0142] s = [h, v, θ, d] (7)

[0143] Where h, v, θ are the height, speed and heading angle of the aircraft at time t, respectively;

[0144] Set the range of aircraft state parameters: height h is 5000-20000 meters; speed is 1000-3000 m / s; heading angle is 0-360 degrees;

[0145] d is the interceptor information, expressed as:

[0146] d = [p I ,v I ,L I ] (8)

[0147] Where d represents the interceptor information at time t, p I , v I , L I represent the position, speed and type of the interceptor at time t, respectively;

[0148] Interceptor information parameter range: position p I is 0-50000 meters; speed v I is 500-1500 m / s; type L I is a discrete integer;

[0149] S2.2. Design action:

[0150] According to the state of the aircraft and the information of the interceptor, the action is decided through the strategy network, expressed as:

[0151] a = π(s, d) (9)

[0152] Where a is defined as:

[0153]

[0154] Where Λ is the wing sweep angle, h c is the desired height, v c is the desired speed;

[0155] S2.3. Design the reward function:

[0156] The reward function R(s t ,a r,t ) is composed of the following parts:

[0157] R(st ,a t )=w1R target -w2R intercept -w3R energy +w4R stability (11)

[0158] Among them, R target is the target reward, R intercept For interception penalty, R energy is the energy penalty, R stability Reward for stability;

[0159]

[0160] Among them, d(p r,t ,p target ) is the Euclidean distance from the aircraft position to the target point, d(p r,t ,p b,t ) is the Euclidean distance from the aircraft position to the interceptor position, d thresh is the distance threshold for interception penalty, s desired is the desired aircraft state;

[0161] R target Indicates that when the aircraft approaches the target area; R intercept Indicates the penalty for being hit: given when the distance between the aircraft and the interceptor is too close; R energy The more energy the aircraft consumes when performing maneuvers, the greater the penalty is, in order to encourage energy efficiency, R stability Indicates that when no maneuvering is required, the aircraft should be as close to its original flight state as possible;

[0162] S2.4. State transition relationship:

[0163] Let the state s at time t be s t , then the state s at time t+1 t+1 =P(s t ,a), where the process used for training is a deterministic dynamic integral P, expressed as:

[0164]

[0165] Among them, x e ,y n and h are the x-axis displacement, y-axis displacement, and z-axis displacement in the local northeast celestial coordinate system ENU, where the x-axis points to the east of the Earth's origin, the y-axis points to the north of the Earth's origin, and the z-axis is determined by the right-hand rule; V, ξ, ψ are the velocity, flight trajectory angle, and heading angle of the aircraft, respectively; g is the local standard gravity; L is the lift, and D is the drag;

[0166] The expressions of lift and drag are:

[0167] L = c L (Λ)qS ref (14)

[0168] D = c D (Λ)qS ref (15)

[0169] where S ref is the reference area of the aircraft, q is the dynamic pressure, q = 1 / 2p(h)V 2 , p(h) = p0e -βh , p(h) is the atmospheric density, p0is the atmospheric density at sea level, β is a constant, c L (Λ) is the lift coefficient of the aircraft, expressed as a function of the sweep angle, c D (Λ) is the drag coefficient of the aircraft, also expressed as a function of the sweep angle, both of which are obtained by interpolation from the aircraft aerodynamic parameter table.

[0170] Further, both the aircraft and the interceptor follow this motion equation. In particular, we call N y = L sin σ and N z = L cos σ the longitudinal and lateral overloads, which are obviously directly related to the wing sweep angle due to the c L (Λ) term;

[0171] S3. Simulate various flight scenarios in the simulation environment designed in step S2, including different initial flight conditions, states of enemy interceptors, and different target points, and collect flight data for the intelligent agent neural network model of the morphing aircraft maneuver trajectory design;

[0172] Further, an exploration weighting factor based on state uncertainty is introduced. This factor can dynamically adjust the exploration intensity, so that the intelligent agent increases exploration in states with high uncertainty, and reduces exploration in states that have been sufficiently learned.

[0173] Further, the specific implementation method of step S3 includes the following steps:

[0174] S3.1. Initialization: first, initialize the weights of the policy network, the value network, and the target value network, and set the initial value to a random number in the range of (0, 1) during the initialization phase;

[0175] S3.2. State and action sampling: in each iteration, sample the state s t at time t from the state space according to the current policy network, sample the action a t at time t from the action space, and the expression is:

[0176]

[0177] Obtain the sample flight data (s t ,a t ,r t ,s t+1 ) stored in the experience replay buffer.

[0178] S4. Train the agent neural network model for the design of the variable aircraft maneuver trajectory by using the SAC algorithm, and add an adaptive exploration mechanism in the training process to adjust the strength of exploration according to the current flight state and environmental conditions, and obtain the trained agent neural network model for the design of the variable aircraft maneuver trajectory;

[0179] Further, the specific implementation method of step S4 includes the following steps:

[0180] S4.1. Based on the SAC algorithm, set the state s t , the action a t , and the reward R t observed after the action a t is taken, and the exploration weighting factor is β(s t ), and the expression of the improved exploration strategy is:

[0181]

[0182] Where ∈ is a random variable sampled from a Gaussian distribution, and α is a temperature parameter for adjusting the balance between exploration and utilization; the exploration weighting factor β(s t ) is dynamically calculated based on the access frequency or variance of the value function of the state s t .

[0183] S4.2. Set the processing method of the delayed reward: when facing the delayed reward, use the improved reward credit distribution mechanism to set the adjusted future reward G t at time t, and use the improved time difference error to update the value function, and the expression is:

[0184] G t =R t+1 +γR 2 +γ t+2 R n-1 +...+γ t+n-1 R n +γ t+n Q(s t+n ,a t ) (18)

[0185] δ t =G ta t ) (19)

[0186] wherein γ is a discount factor, n is the number of steps ahead, R t represents the reward of the t-th step;

[0187] S4.3. Perform policy network update, using δ t Update the policy network weights, the update steps are as follows:

[0188]

[0189] wherein δ t is the time difference error;

[0190] S4.4. Perform value network update, process the sample flight data obtained in step S3, and calculate the target value y, the expression is:

[0191] y = r + γ (V (s) - a log π θ′ (a' | s')) (21)

[0192] wherein V is the value function;

[0193] Then update the network weights using the mean square error, the expression is:

[0194]

[0195] wherein E represents the expectation of the random variable b under the sample a, φ 1,2 represents the weights of the value network and the target value network used by the SAC algorithm;

[0196] Update the value network weights and the target value network weights respectively using the soft update method, the expression is:

[0197] φ i ' <- τφ i + (1-τ)φ i '(23)

[0198] wherein when i is 1, φ i ' is the value network weight, when i is 2, φ i ' is the target value network weight, and τ is the soft update coefficient;

[0199] Further, the soft update coefficient τ determines the speed of the target network parameter update. A small τ value means that the target network parameter is updated slowly, which helps to stabilize the training but may slow down the learning speed. If the target network is updated too fast, it may lead to unstable training; if it is updated too slowly, the learning response is slow. Adjust τ according to the performance of the agent in the environment.

[0200] S4.5. Adjust and optimize algorithm parameters based on training results: adjust and optimize the entropy weight a and the soft update coefficient τ according to the comprehensive training index. The expression for adjusting a is:

[0201]

[0202] wherein, is the target entropy, usually set to the value of the negative action space dimension to encourage sufficient exploration; adjust the entropy weight a, if the agent shows excessive exploration, i.e. the strategy is too random, then reduce a; if the agent does not explore enough, then increase a; then adjust τ according to the performance of the agent in the environment, to obtain the trained agent neural network model for morphing aircraft maneuver trajectory design.

[0203] S5. Deploy the trained agent neural network model for morphing aircraft maneuver trajectory design in the onboard computer for real-time trajectory adjustment and morphing decision in the flight control loop.

[0204] Further, in step S5, network pruning and quantization techniques are applied to reduce the size of the agent's model to fit the onboard computing capability. The goal is to set the model size to no more than 50MB.

[0205] Further, in actual application, the aircraft state information is obtained from the aircraft integrated navigation unit, and the interceptor information is obtained from the seeker and ground information chain, which is transmitted to the trained agent installed in the aircraft onboard computer for online decision-making to obtain the optimal wing sweep angle and desired position and speed.

[0206] When deploying the agent, first perform hardware and software configuration. Hardware requirements: processor, main frequency not less than 1.5GHz to support parallel data processing and real-time decision-making; memory, at least 4GB RAM to ensure sufficient memory to process large amounts of flight data; storage, at least 64GB for storing algorithm models and their updates. Software requirements: install and configure the operating system (RTOS), and ensure that all necessary programming environments and libraries (such as Python, TensorFlow or PyTorch) are correctly installed and configured.

[0207] Subsequently, network pruning and quantization techniques are applied to reduce the size of the agent's model to fit the onboard computing capability. The goal is to reduce the model size to no more than 50MB. Use specialized tools (such as TensorFlow Lite or ONNX) to convert the trained model into a format suitable for deployment, and use the local compiler of the onboard system for optimized compilation.

[0208] Finally, the intelligent agent is integrated into the real-time operating system of the aircraft to ensure the stability of the model running under the RTOS and the seamless cooperation with the navigation and communication systems of the aircraft.

[0209] Further, the attached Figure 4 The average reward curve with the change of the training step is shown, and the reward rise indicates that the performance of the agent in the simulation environment is getting better. If the reward does not change, it means that the training converges. It can be seen that the agent completes the training at about 8 million steps in the case of using the random strategy (SAC). As a comparison, the training convergence step of the deterministic strategy is not stable. Figure 5 The average reward curve with the change of the training step is shown, and the reward rise indicates that the performance of the agent in the simulation environment is getting better. If the reward does not change, it means that the training converges. It can be seen that the agent completes the training at about 8 million steps in the case of using the random strategy (SAC). As a comparison, the training convergence step of the deterministic strategy is not stable. Figure 6 The maneuvering overload instruction in an example scenario is shown, in which the interceptor intercepts the aircraft with proportional guidance, and the aircraft adopts intelligent maneuvering to generate the terminal maneuvering overload. Figure 7 For the corresponding three-dimensional view, it can be seen that the aircraft evades the interceptor in an S shape, so that the interceptor fails to meet the interception condition.

[0210] Embodiment 2:

[0211] An electronic device, characterized by comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm according to the embodiment 1 when executing the computer program.

[0212] The computer device of the present application can be a device comprising a processor and a memory, such as a single-chip microcomputer comprising a central processing unit. The processor is used to execute the computer program stored in the memory to implement the steps of the method for designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm.

[0213] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0214] The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by at least one function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data (such as audio data, a phone book, etc.) created according to the use of the mobile phone, and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.

[0215] Embodiment 3

[0216] A computer readable storage medium, having stored thereon a computer program, wherein the computer program is executed by a processor to implement the method of designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm according to any one of embodiments 1-3.

[0217] The computer readable storage medium of the present application can be any form of storage medium readable by a processor of a computer device, including but not limited to a non-volatile memory, a volatile memory, a ferroelectric memory, etc., and the computer readable storage medium has stored thereon a computer program, when the processor of the computer device reads and executes the computer program stored in the memory, the steps of the method of designing a maneuvering trajectory of a morphing aircraft based on a SAC algorithm can be implemented.

[0218] The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0219] It has to be noted that the terms "first", "second", and the like in connection with an entity or action refer to this entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Also, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without further constraints, exclude the presence of additional elements of the process, method, article, or apparatus.

[0220] While the application has been described with reference to specific implementations thereof, it should be understood that various modifications and substitutions can be made by those skilled in the art without departing from the scope of the present application. Especially, features of the specific implementations disclosed herein can be combined in any manner, unless structural conflicts arise, and the combinations are not exhaustively described in the present specification only for the sake of brevity and conciseness. Therefore, the application is not limited to the specific implementations disclosed herein but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for designing maneuvering trajectories of a deformable aircraft based on the SAC algorithm, characterized in that: The steps include: S1. Build an agent-based neural network model for maneuver trajectory design of a deformable aircraft, including a policy network π for generating actions and a value network Q for estimating the value of each state-action pair. S2. Model the maneuverability of a deformable aircraft under interception using a Markov process and design the simulation environment, including the aircraft state, actions, reward function, and aircraft state transition relationships. S3. Simulate various flight scenarios in the simulation environment designed in step S2, including different initial flight conditions, hostile interceptor states, and different target points, and collect flight data for the intelligent agent-based neural network model to design the maneuver trajectory of the deformable aircraft; S4. Train an agent-based neural network model for maneuvering trajectory design of a morphing aircraft using the SAC algorithm. Incorporate an adaptive exploration mechanism into the training process to adjust the exploration intensity based on the current flight state and environmental conditions, resulting in a trained agent-based neural network model for maneuvering trajectory design of a morphing aircraft. The specific implementation method of step S4 includes the following steps: S4.

1. Improve based on SAC algorithm, set the agent in state s t Take action a t The reward observed later is R t , the exploration weighting factor is β(s t ), then the expression of the improved exploration strategy is: Among them, ∈ is a random variable sampled from a Gaussian distribution, α is a temperature parameter used to adjust the balance between exploration and exploitation; the exploration weighting factor β(s t ) Based on state s t Dynamic calculation based on the access frequency or variance of the value function; S4.

2. Set the method for handling delayed rewards: When facing delayed rewards, adopt an improved reward credit allocation mechanism and set the adjusted future reward G at time t t , and use the improved temporal difference error to update the value function, the expression is: G t =R t +γR t+1 +g 2 R t+2 +…+c n-1 R t+n-1 +g n Q(s t+n ,a t+n ) (2) δ t =G t -Q(s t ,a t ) (3) Where γ is the discount factor, n is the number of steps to look ahead, and R t represents the reward at step t; S4.

3. Update the policy network using δ t Update the policy network weights. The update steps are as follows: Q(s t ,a t )←Q(s t ,a t )+αδ t (4) Among them, δ t is the timing difference error; S4.

4. Update the value network, process the sample flight data obtained in step S3, and calculate the target value y; Then use the mean squared error to update the network weights; Use soft update method to update value network weight and target value network weight respectively; S4.

5. Adjust and optimize algorithm parameters based on training results: adjust the entropy weight α and soft update coefficient τ to optimize the overall training indicators; S5. Deploy the trained agent-based neural network model for maneuver trajectory design of deformable aircraft in the onboard computer for real-time trajectory adjustment and deformation decision-making in the flight control loop.

2. The method for designing maneuvering trajectory of a deformable aircraft based on the SAC algorithm according to claim 1, characterized in that: The specific implementation method of step S1 includes the following steps: S1.

1. Build a policy network π to generate actions. The policy network consists of an input layer, a hidden layer, and an output layer. The size of the input layer is set to the dimensions of the aircraft state and interceptor information. The aircraft state is set to include three parameters: altitude, speed, and heading angle. The interceptor information includes three parameters: position, speed, and model. The size of the input layer is set to 6 layers. The hidden layer is set to two hidden layers, each of which includes 256 neurons; The output layer size is set to the dimension of the action. The action includes three parameters: the wing sweep angle, the desired position, and the desired speed. The output layer size is set to 3 layers. The policy network is expressed as: a=π(s) (5) Where a is the action, s is the aircraft state, and π represents the nonlinear function relationship expressed by the policy network; S1.

2. Construct a value network Q. The value network has the same structure as the target value network and is used to estimate the value of each state-action pair. It includes an input layer, a hidden layer, and an output layer. The size of the input layer is set to the dimensions of aircraft state, interceptor information, and action. The aircraft state includes three parameters: altitude, speed, and heading angle. The interceptor information includes three parameters: position, speed, and model. The action includes three parameters: wing sweep angle, desired position, and desired speed. The size of the input layer is set to 9 layers. The hidden layer is set to two hidden layers, each of which includes 256 neurons; The output layer size is set to the value of the state-action pair, and the output layer size is set to 1 layer; The value network is expressed as: Q=Q(s,a) (6) Among them, Q is the value of the optimal problem, Q is the nonlinear functional relationship expressed in the value network, that is, the optimality index; S1.

3. Set the calculation rules of the strategy network and value network: Calculate by the forward calculation method of the neural network, and set the output of the lth layer to be a (l) , the calculation of each layer is divided into two steps: linear transformation and nonlinear activation function, expressed as: z (l) =W (l) a (l-1) +b (l) (7) a (l) =σ (l) (z (l) ) (8) Among them, W (l) is the weight matrix of the lth layer, b (l) is the bias vector, σ (l) is the activation function of the lth layer, set to σ (l) (x) = tanh(x); The output of the policy network is expressed as: a=a (2) (9) The output of the value network is expressed as: Q=a (2) (10)。 3. The method for designing maneuvering trajectory of a deformable aircraft based on the SAC algorithm according to claim 2, characterized in that: The specific implementation method of step S2 includes the following steps: S2.

1. Design aircraft status: The aircraft status includes the current altitude, speed, heading angle parameters, and interceptor information, and is expressed as: s=[h,v,θ,d] (11) Among them, h, v, and θ are the altitude, speed, and heading angle of the aircraft at time t, respectively; Set the aircraft status parameter range: altitude h is 5000 meters to 20000 meters; speed is 1000m / s to 3000m / s; heading angle is 0 degrees to 360 degrees; d is the interceptor information, the expression is: d=[p I ,v I ,L I ] (12) Among them, d represents the interceptor information at time t, p I 、v I 、L I represent the position, velocity and model of the interceptor at time t respectively; Interceptor information parameter range: position p I 0m-50000m; speed v I 500m / s-1500m / s; Model L I is a discrete integer; S2.

2. Design action: According to the status of the aircraft and the information of the interceptor, the decision action is made through the policy network, and the expression is: a=π(s,d) (13) Where a is defined as: Where Λ is the wing sweep angle, h c is the desired height, v c is the expected speed; S2.

3. Design reward function: Reward function R(s t ,a t ) consists of the following parts: R(s t ,a t )=w1R target -w2R intercept -w3R energy +w4R stability (15) Among them, R target is the target reward, R intercept For interception penalty, R energy is the energy penalty, R stability Reward for stability; Among them, d(p r,t ,p target ) is the Euclidean distance from the aircraft position to the target point, d(p r,t ,p b,t ) is the Euclidean distance from the aircraft position to the interceptor position, d thresh is the distance threshold for interception penalty, s desired is the desired aircraft state; R target Indicates that when the aircraft approaches the target area; R intercept Indicates the penalty for being hit: given when the distance between the aircraft and the interceptor is too close; R energy The more energy the aircraft consumes when performing maneuvers, the greater the penalty is, in order to encourage energy efficiency. stability Indicates that when no maneuvering is required, the aircraft should remain close to its original flight state; S2.

4. State transition relationship: Let the state s at time t be s t , then the state s at time t+1 t+1 =P(s t ,a), where the process used for training is a deterministic dynamic integral P, expressed as: Among them, x e ,y n and h are the x-axis displacement, y-axis displacement, and z-axis displacement in the local northeast celestial coordinate system ENU, where the x-axis points to the east of the Earth's origin, the y-axis points to the north of the Earth's origin, and the z-axis is determined by the right-hand rule; V, ξ, ψ are the velocity, flight trajectory angle, and heading angle of the aircraft, respectively; g is the local standard gravity; L is the lift, and D is the drag; The expressions for lift and drag are: L=c L (Λ)qS ref (18) D=c D (Λ)qS ref (19) Among them, S ref is the reference area of ​​the aircraft, q is the dynamic pressure, q=1 / 2ρ(h)V 2 ,ρ(h)=ρ0e -βh , ρ(h) is the atmospheric density, ρ0 is the atmospheric density at sea level, β is a constant, c L (Λ) is the lift coefficient of the aircraft, expressed as a function of the sweep angle, c D (Λ) is the aircraft drag coefficient, which is also expressed as a function of the sweep angle. Both are obtained by interpolation from the aircraft aerodynamic parameter table.

4. The method for designing maneuvering trajectory of a deformable aircraft based on the SAC algorithm according to claim 3, characterized in that: The specific implementation method of step S3 includes the following steps: S3.

1. Initialization: First, initialize the weights of the policy network, the value network, and the target value network. During the initialization phase, set the initial values ​​to random numbers in the range (0, 1). S3.

2. State and action sampling: In each iteration, the state s at time t is sampled from the state space according to the current policy network t , sample the action a at time t from the action space t , the expression is: Get sample flight data (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer.

5. The method for designing maneuvering trajectory of a deformable aircraft based on the SAC algorithm according to claim 4, characterized in that: In step S4.4, the target value y is calculated as follows: y=r+γ(V(s)-αlogπ θ′ (a′|s′)) (21) in, is the value function; The expression for updating the network weights using mean square error is: J Q (φ 1,2 )=E (s,a,r,s′)~D [(Q(s,a)-y) 2 ] (22) Among them, E a (b) represents the expectation of random variable b under sample a, φ 1,2 Represents the weights of the value network and target value network used by the SAC algorithm; The expressions for updating the value network weight and target value network weight using the soft update method are: f i ′←tφ i +(1-τ)φ i ′ (23) When i is 1, φ i ′ is the value network weight, when i is 2, φ i ′ is the target value network weight, τ is the soft update coefficient; The expression for adjusting α in step S4.5 is: in, is the target entropy, set to the value of the negative action space dimension to encourage sufficient exploration; adjust the entropy weight α. If the agent exhibits over-exploration, that is, the strategy is too random, then reduce α; if the agent does not explore enough, then increase α; then adjust τ according to the agent's performance in the environment, and obtain a trained agent neural network model for maneuver trajectory design of deformable aircraft.

6. The method for designing maneuvering trajectory of a deformable aircraft based on the SAC algorithm according to claim 5, characterized in that: In step S5, network pruning and quantization techniques are applied to set the agent's model to a size suitable for onboard computing power, with the goal of setting the model size to no more than 50MB.

7. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for designing a maneuvering trajectory of a deformable aircraft based on a SAC algorithm as described in any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for designing a maneuvering trajectory of a deformable aircraft based on the SAC algorithm as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Ground target trajectory tracking method based on hybrid grid multiple models

    CN115390560A

  • Target hierarchical architecture-based deformable aircraft intelligent planning method and system

    CN117268391A