A time of flight control guidance law for a maneuvering target
Patent Information
- Application Number
- CN202411024674.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-07-29
AI Technical Summary
然而,依赖于恒定速度假设的分析方法可能面临精度问题,而基于数值积分的估计又可能受限于计算能力
[0012](1)、根据本发明提供的针对机动目标的飞行时间控制制导律,提出了一种新的双分支神经网络架构来处理拦截飞行器的状态向量和目标的轨迹序列,从而提高了速度时变模型的剩余飞行时间预测精度;
Smart Images

Figure CN121433266B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a control method for aircraft, specifically to a flight time control guidance law for maneuvering targets. Background Technology
[0002] With the gradual development of technology, it is becoming increasingly difficult for a single aircraft to penetrate enemy defense systems and hit its target. In recent years, multi-aircraft cooperative guidance has improved the penetration probability of missiles through mutual cooperation and coordination, enabling them to jointly complete strike missions. To achieve simultaneous target hits by multiple aircraft, terminal time constraints need to be introduced into the guidance law, that is, to ensure that multiple aircraft hit the target within a specified time. There are two main ways to achieve impact time control using cooperative guidance strategies: controlling the remaining flight time individually through the Impact Time Control Guidance (ITCG) method, or receiving information from neighboring aircraft to obtain a consistent remaining flight time. When intercepting maneuvering targets, the uncertainty of the target's behavior poses a significant challenge to the deployment of guidance algorithms. Using inter-aircraft communication networks to merge information from other interceptor missiles is problematic because the latency and vulnerability of the network can severely affect guidance accuracy. Therefore, the method of having the interceptor aircraft independently estimate the remaining flight time is more suitable for intercepting maneuvering targets. However, analytical methods relying on the assumption of constant velocity may face accuracy issues, while estimations based on numerical integration may be limited by computational capabilities.
[0003] Based on the above problems, the inventors have conducted in-depth research on aircraft control methods that independently estimate the remaining flight time, with the aim of designing a flight time control guidance law for maneuvering targets that can solve the above problems. Summary of the Invention
[0004] To overcome the aforementioned problems, the inventors conducted intensive research and designed a flight time control guidance law for maneuvering targets. This guidance law first establishes a two-branch neural network by simultaneously processing the interceptor missile's state vector and the target's trajectory sequence to accurately estimate the remaining flight time for intercepting maneuvering targets. Then, model-independent reinforcement learning is used to correct errors affecting the remaining flight time. This method overcomes the problems caused by target uncertainty when intercepting maneuvering targets, exhibiting excellent control performance. This guidance law can be used in various types of aircraft and has high practical value; thus, this invention is completed.
[0005] Specifically, the purpose of this invention is to provide a time-of-flight control guidance law for maneuvering targets, which controls the aircraft through the following steps:
[0006] Step 1: Obtain guidance command a in real time using the offset proportional guidance method. MThe guidance command controls the aircraft to fly towards the target, while simultaneously recording and outputting the current flight time t.
[0007] Step 2: Obtain the predicted remaining flight time in real time using a two-branch neural network model.
[0008] Step 3, using the predicted current remaining flight time Expected interception time t d And t to obtain the remaining flight time error ε t ,
[0009] Step 4, using the remaining flight time error ε t Revise the near-end strategy optimization model;
[0010] Step 5: Repeat steps 1 to 4 in real time, so that the aircraft follows guidance command a. M Under the control of [the system], according to the expected interception time t d Hit the target.
[0011] The beneficial effects of this invention include:
[0012] (1) Based on the flight time control guidance law for maneuvering targets provided by the present invention, a new dual-branch neural network architecture is proposed to process the state vector of the interceptor aircraft and the trajectory sequence of the target, thereby improving the prediction accuracy of the remaining flight time of the velocity time-varying model.
[0013] (2) According to the flight time control guidance law for maneuvering targets provided by the present invention, a bias command based on reinforcement learning is proposed to offset the remaining flight time error, which can effectively intercept maneuvering targets with time-varying speed and has broad application prospects. Attached Figure Description
[0014] Figure 1 This diagram illustrates the overall logic of the flight time control guidance law for maneuvering targets according to this application.
[0015] Figure 2 The learning curve of the proximal policy optimization model in Example 1 is shown;
[0016] Figure 3 The comparison curves between the predicted remaining flight time and the actual remaining flight time obtained by the dual-branch neural network model in Example 2 are shown.
[0017] Figure 4 A schematic diagram of the flight trajectory in Example 3 is shown;
[0018] Figure 5 A schematic diagram of flight time in Example 3 is shown;
[0019] Figure 6 A schematic diagram of the motion speed in Example 3 is shown;
[0020] Figure 7 A schematic diagram of the guidance command in Embodiment 3 is shown;
[0021] Figure 8 A schematic diagram of the flight trajectory in Example 4 is shown;
[0022] Figure 9 A schematic diagram of flight time in Example 4 is shown;
[0023] Figure 10 A schematic diagram of the motion speed in Example 4 is shown;
[0024] Figure 11 A schematic diagram of the guidance command in Embodiment 4 is shown;
[0025] Figure 12 A schematic diagram of the flight trajectory in Example 5 is shown;
[0026] Figure 13 A schematic diagram of flight time in Example 5 is shown;
[0027] Figure 14 A schematic diagram of the motion speed in Example 5 is shown;
[0028] Figure 15 A schematic diagram of the guidance command in Embodiment 5 is shown;
[0029] Figure 16 The diagram shows the structure of the Proximal Policy Optimization (PPO) model. Detailed Implementation
[0030] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become clearer and more apparent.
[0031] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments. Although various aspects of embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless specifically indicated otherwise.
[0032] According to the present invention, a time-of-flight control guidance law for maneuvering targets is provided, such as... Figure 1 As shown, the guidance law controls the aircraft through the following steps:
[0033] Step 1: Obtain guidance command a in real time using the offset proportional guidance method. M The guidance command controls the aircraft to fly toward the target, while recording and outputting the current flight time t.
[0034] In this application, steps 1 to 4 are executed in real time in a loop, once per sampling control cycle, with a frequency of 20-50Hz, which is selected and set according to the performance of the hardware equipment on the aircraft.
[0035] In step 1, the bias ratio guidance method obtains the guidance command a using the following formula (i). M :
[0036] a M =a0+a b (one)
[0037] Where a0 represents the initial guidance command,
[0038] a b This represents the bias term used to reduce interception time error, which is obtained in real time through the near-end policy optimization model.
[0039] The initial guidance command a0 is obtained by the following formula (ii):
[0040]
[0041] Among them, V I The speed of the aircraft is measured in real time by sensors on the aircraft.
[0042] θ I The trajectory inclination angle of an aircraft is indicated by the seeker on the aircraft.
[0043] The angular velocity representing the missile's line-of-sight is obtained from the seeker on the aircraft.
[0044] g represents gravitational acceleration, with a value of 9.8.
[0045] Step 2: Obtain the predicted remaining flight time in real time using a two-branch neural network model.
[0046] The dual-branch neural network model includes an MLP branch and a GRU branch.
[0047] The dual-branch neural network model uses real-time information about the aircraft itself and information about the target captured by the seeker as inputs.
[0048] In the model, the input data for the MLP branch includes: the relative velocity between the aircraft and the target, the trajectory tilt angle, the distance between the aircraft and the target along the X-axis, the distance between the aircraft and the target along the Y-axis, the straight-line distance between the aircraft and the target, and the line-of-sight angle between the missile and the target; the input data for the GRU branch is the target's historical trajectory. Taking the acquisition of the target as the starting point, the target position at this time is (0,0). The target position at that moment is collected and recorded once in each sampling control cycle, thus forming a set of target positions, which is the target's historical trajectory.
[0049] In this application, the MLP branch includes two fully connected hidden layers, each containing 100 neurons;
[0050] The GRU branch consists of 10 units, each containing 100 neurons.
[0051] Preferably, the training samples of the dual-branch neural network model are multiple sets of mapping data.
[0052] The mapping data is:
[0053]
[0054] in, Indicates the remaining flight time.
[0055] Indicates the relative speed between the aircraft and the target.
[0056] Indicates the trajectory inclination angle,
[0057] This indicates the distance between the aircraft and the target along the X-axis.
[0058] Y t I -Y t T This indicates the distance between the aircraft and the target along the Y-axis.
[0059] R t This represents the straight-line distance between the aircraft and the target.
[0060] λ t Indicates the line-of-sight angle of the bullet.
[0061] T t Represents the historical trajectory of the target;
[0062] Preferably,
[0063] That is, the historical trajectory of a target is a set of multiple historical position coordinates of the target. In the target's movement trajectory, starting from the moment when time is 0, the position coordinates of the target are captured at fixed time intervals to obtain the historical trajectory of the target.
[0064] This represents the target's coordinates in the X direction at the first extracted moment.
[0065] This indicates the target's coordinates along the X-axis at the current moment.
[0066] This represents the target's coordinates in the Y direction at the first extracted moment. t T This indicates the target's coordinates along the Y-axis at the current moment.
[0067] Preferably, during the training process of the dual-branch neural network model, R t ,λ t As the input to the MLP branch, T t As input to the GRU branch;
[0068] Preferably, during the training process of the two-branch neural network model, the network parameters are represented by β, and the loss function is represented as a function of β as shown in equation (iii) below.
[0069]
[0070] Where, N D Represents the total number of samples.
[0071] i represents the i-th sample.
[0072] This represents the remaining flight time predicted by the model based on the i-th sample.
[0073] t go,i This represents the actual remaining flight time in the i-th sample.
[0074] Preferably, the loss function is optimized using stochastic gradient descent (SGD). The update formula for the network parameters can be expressed as:
[0075]
[0076] Where, β new This indicates the updated network parameters.
[0077] β old This indicates the network parameters before the update.
[0078] α β This represents the learning rate, with a value of 0.001.
[0079] This represents the gradient of the network parameter β, which is equivalent to the loss function. The gradient is a vector whose direction points to the direction of the fastest growth of the function, and its magnitude represents the rate of growth. The specific value of the gradient is obtained by differentiating the loss function.
[0080] Step 3, using the predicted current remaining flight time Expected interception time t d And t to obtain the remaining flight time error ε t .
[0081] Among them, the remaining flight time error ε t Obtained through the following formula:
[0082] ε t =t d -(t+t go )
[0083] Step 4, using the remaining flight time error ε t The near-end strategy optimization model has been revised.
[0084] In this application, the proximal strategy optimization model is the PPO model.
[0085] Preferably, the structure of the proximal strategy optimization model, i.e., the PPO model, is as follows: Figure 16 As shown;
[0086] The agent in the PPO model consists of two neural networks: a policy network and an evaluation network. The policy network represents the mapping between the current state and instructions; the evaluation network estimates the potential value of the current state and then, combined with the reward sequence obtained from implemented instructions, calculates the advantage function of the reward for these instructions relative to their potential value. When the advantage function is positive, the probability of implemented instructions in the policy is increased; when the advantage function is negative, the probability of these instructions is decreased.
[0087] To reduce fluctuations during training, PPO limits the magnitude of policy updates, minimizing the difference between the old and new policies.
[0088] The goal of Propositional Purpose (PPO) is to optimize a policy that maximizes the total environmental reward for the agent in an unknown environment. However, the total environmental reward is generally not directly calculable. In discrete systems, the total environmental reward has the following form:
[0089]
[0090] In the formula, γ∈(0,1] represents the discount factor, i represents the future time, and r i For future single rewards, This represents the total environmental reward value for the future starting at time t.
[0091] To represent the total environmental reward under different states, s is defined in the form of mathematical expectation. t State value function V π (s t ), representing state s t Potential value
[0092]
[0093] To evaluate the merits of the behavior, this application defines a state s. t The next line is a t Behavioral value function Q π (s t ,a t ), indicating that the agent is in state s. t The following action a is to be executed t Potential value
[0094]
[0095] V π (s t Q can be used π (s t ,a t ) represents the following
[0096]
[0097] use The function representing the optimal state value. This represents the optimal behavior value function, where the agent is in state s. t The optimal strategy π * (s t That is, a strategy that maximizes both the state value function and the behavior value function.
[0098]
[0099] π * (a t |s t ) = argmaxQ π (s t ,a t )
[0100] PPO uses evaluation network estimation That is, the input and output of the evaluation network are s, respectively. t and To design the loss function for evaluating the network, we define the advantage function A. π (s t ,a t )for
[0101] A π (s t ,a t )=Q π (s t ,a t )-V π (s t )
[0102] A π (s t ,a t The definition is given in state s. t Perform a specific action t Compared to the current policy π(a) t |s t The advantages of A π (s t ,a t The smaller π(a) is, the better the strategy π(a) becomes. t |s t The closer to the optimal strategy π * (a t |s t If the current policy is the optimal policy, the corresponding value function is also the optimal value function. Therefore, the loss function of the value evaluation network in this application can be designed as follows:
[0103]
[0104] Where ξ represents the parameters of the evaluation network, and the parameter update formula is:
[0105]
[0106] In the formula, α ξ To evaluate the learning rate of the network.
[0107] Since the value function in the advantage function cannot be directly calculated, this application uses the Generalized Advantage Estimation (GAE) function to estimate the advantage function.
[0108]
[0109] In the formula, k is the estimation step size, which is related to the number of interaction samples. The results are provided by the evaluation network.
[0110] Policy networks use state s tAs input, with strategy π(a) t |s t The parameters of ) are output.
[0111] For a continuous action space, a Gaussian distribution is used to represent the probability distribution of actions, and the output of the policy network is the mean μ of this Gaussian distribution. π and standard deviation σ p i. The policy network expects its output to be the parameters of the optimal policy. However, directly obtaining the optimal policy relies heavily on expert experience, making it impossible to design a loss function using the mean squared error of the optimal policy and the current policy. Therefore, a surrogate loss function is needed to evaluate the improvement in cumulative reward after the policy parameter update. A shearing function is used to constrain the magnitude of the policy update.
[0112] In this application, the objective function of the PPO model is as follows:
[0113]
[0114] In the formula, o t (ω) is a ratio function. Here are the shearing functions, respectively.
[0115]
[0116] In the formula, Update the shearing parameters of the amplitude for the constraint strategy.
[0117] The shearing function constrains the ratio of the new and old strategies to... Within this constraint, the update range of the new policy is limited. The update formula for the policy network parameters is as follows:
[0118]
[0119] In the formula, α ω is the learning rate of the policy network.
[0120] During the flight of the aircraft, the maximum value function can be obtained in real time through this near-end policy optimization model, thereby obtaining the behavior 'a' of the near-end policy optimization model. t And it is defined as the bias term a that controls flight time. b
[0121] In this application, the state s of the near-end strategy optimization is affected by three variables, namely the remaining flight time t. go Remaining flight time error ε t and the speed V of the aircraft I By correcting the near-end strategy optimization model in real time using the remaining flight time error, the output 'a' of the near-end strategy optimization model at the next time step can be improved. MIt has the ability to correct for actual flight time errors, ultimately reducing the remaining flight time error ε. t It gradually approaches zero.
[0122] Step 5: Repeat steps 1 to 4 in real time, so that the aircraft follows guidance command a. M Under the control of [the system], according to the expected interception time t d Hit the target.
[0123] Example 1
[0124] Training the proximal policy optimization model:
[0125] The PPO model was selected for the near-end policy optimization model. The hyperparameter settings for training the PPO model are shown in Table 1 below:
[0126] Table 1
[0127]
[0128] PPO training utilizes the initial conditions described in Table 2. In each set, t d The endpoint time is set as the initial prediction plus a time randomly selected between 5 and 10 seconds.
[0129] Table 2 Initial Conditions for the Aircraft and the Target
[0130]
[0131]
[0132] PPO training utilizes scenarios with identical initial conditions as described in Table 1. In each episode, t d被 The endpoint time is set to the initial prediction time plus a randomly selected time between 5 and 10 seconds. The hyperparameters used in PPO training are shown in Table 2.
[0133] The objective function of PPO during training is as follows:
[0134]
[0135] In the formula, o t (ω) is a ratio function. Here are the shearing functions, respectively.
[0136]
[0137] In the formula, Update the shearing parameters of the amplitude for the constraint strategy.
[0138] The shearing function constrains the ratio of the new and old strategies to... Within this constraint, the update range of the new policy is limited. The update formula for the policy network parameters is as follows:
[0139]
[0140] In the formula, α ω is the learning rate of the policy network.
[0141] After running the PPO algorithm for 150 episodes, the reward value was observed to reach a stable state, and its learning curve is as follows: Figure 2 As shown, from Figure 2 It is known that the near-end policy optimization model in this application achieves a stable average reward in just 100 episodes.
[0142] Example 2
[0143] Training a two-branch neural network model:
[0144] The two-branch neural network model includes an MLP branch and a GRU branch.
[0145] The MLP branch includes two fully connected hidden layers, each containing 100 neurons;
[0146] The GRU branch consists of 10 units, each containing 100 neurons.
[0147] Based on the initial conditions of the aircraft in Table 2, at least 20,000 sets of simulation experiments were conducted under bias proportional guidance. The aircraft attacked maneuvering targets, generating training samples for the dual-branch neural network (predictor).
[0148] 20,000 sets of aircraft and target trajectories were collected during aircraft strikes against maneuvering targets, including 5,000,000 samples used to train the dual-branch neural network model. Each sample consisted of the network's input (V). I ,θ I ,X I -X T ,Y I -Y T ,R,λ,T t ) and output t go For composition, that is, in each trajectory, information at any time point can be extracted to generate a set of samples.
[0149] Will R t ,λ t As the input to the MLP branch, T t As input to the GRU branch;
[0150] These samples were divided into two groups: 98% were used as the training set, and the remaining 2% were used as the test set. The learning rate used to train the predictor was set to α.β =0.001.
[0151] The loss function is expressed as a function of β as shown in equation (iii) below.
[0152]
[0153] N D Represents the total number of samples.
[0154] i represents the i-th sample.
[0155] This represents the remaining flight time predicted by the model based on the i-th sample.
[0156] t go,i This represents the actual remaining flight time in the i-th sample.
[0157] After training all samples is complete, the performance of the two-branch neural network model is first evaluated through a single experiment:
[0158] The initial conditions are set as follows:
[0159]
[0160] and
[0161] The trained two-branch neural network model is used to generate the predicted remaining flight time. Compared with the actual remaining flight time t go In contrast, such as Figure 3 As shown, the predicted values and actual values almost completely overlap, indicating that the well-trained dual-branch neural network model has excellent performance.
[0162] The performance of the two-branch neural network model was then evaluated using a test set, where the root mean square error (RMSE) was 0.1699 and the maximum error was 1.4969. Although the maximum error was relatively large, the RMSE was small. It was observed that there was a large error in the initial stage of the trajectory prediction, but the prediction error gradually decreased as the aircraft approached the target; simulation results indicate that the two-branch neural network model has good performance.
[0163] Example 3
[0164] Three identical aircraft were selected, each equipped with the near-end strategy optimization model obtained in Example 1 and the dual-branch neural network model obtained in Example 2. The control methods for the three aircraft were identical, all including the following steps:
[0165] Step 1: Obtain guidance command a in real time using the offset proportional guidance method.M The guidance command controls the aircraft to fly towards the target, while simultaneously recording and outputting the current flight time t.
[0166] The bias ratio guidance method obtains the guidance command a through the following formula (I). M :
[0167] a M =a0+a b (one)
[0168] Where a0 represents the initial guidance command,
[0169] a b This represents the bias term used to reduce interception time error, which is obtained in real time through the near-end policy optimization model;
[0170] The initial guidance command a0 is obtained through the following equation (ii):
[0171]
[0172] Step 2: Obtain the predicted remaining flight time in real time using a two-branch neural network model.
[0173] Step 3, using the predicted remaining flight time Expected interception time t d And t to obtain the remaining flight time error ε t ,
[0174] ε t =t d -(t+t go )
[0175] Step 4, using the remaining flight time error ε t Revise the near-end strategy optimization model;
[0176] Step 5: Repeat steps 1 to 4 in real time, so that the aircraft follows guidance command a. M Under the control of [the system], according to the expected interception time t d Hit the target.
[0177] The target's maneuver amplitude is set to 0.5g, and the target's maneuver is constant, meaning the target's flight direction is fixed. The expected interception time for the first aircraft is t. d =50s, the expected interception time of the second aircraft is t d =55s, the expected interception time of the third aircraft is t d =60s;
[0178] The flight trajectories, flight times, speeds, and guidance commands of the three aircraft were simulated, and the resulting flight trajectories are as follows: Figure 4 As shown, flight time is as follows Figure 5 As shown, the speed of motion is as follows Figure 6 As shown, the guidance commands are as follows Figure 7 As shown in the figure. According to the simulation results, the time-of-flight control guidance law for maneuvering targets provided in this application can complete the interception within the expected time.
[0179] Example 4
[0180] Three identical aircraft were selected, each equipped with the near-end strategy optimization model obtained in Example 1 and the dual-branch neural network model obtained in Example 2. The control methods for the three aircraft were identical, all including the following steps:
[0181] Step 1: Obtain guidance command a in real time using the offset proportional guidance method. M The guidance command controls the aircraft to fly towards the target, while simultaneously recording and outputting the current flight time t.
[0182] The bias ratio guidance method obtains the guidance command a through the following formula (I). M :
[0183] a M =a0+a b (one)
[0184] Where a0 represents the initial guidance command,
[0185] a b This represents the bias term used to reduce interception time error, which is obtained in real time through the near-end policy optimization model;
[0186] The initial guidance command a0 is obtained through the following equation (ii):
[0187]
[0188] Step 2: Obtain the predicted remaining flight time in real time using a two-branch neural network model.
[0189] Step 3, using the predicted remaining flight time Expected interception time t d And t to obtain the remaining flight time error ε t ,
[0190] ε t =t d -(t+t go )
[0191] Step 4, using the remaining flight time error ε t The near-end strategy optimization model has been revised.
[0192] The expected interception time for all three aircraft is set to t. d =60s;
[0193] The acceleration of the first flying target is a. T =0g, the acceleration of the second flying target is a. T =0.5g, the acceleration of the third flying target is a. T =1g, the three targets have the same maneuvering direction, the three targets have the same initial conditions, and the three aircraft have the same initial conditions. The initial conditions are set as follows:
[0194]
[0195] The flight trajectories, flight times, speeds, and guidance commands of the three aircraft were simulated, and the resulting flight trajectories are as follows: Figure 8 As shown, flight time is as follows Figure 9 As shown, the speed of motion is as follows Figure 10 As shown, the guidance commands are as follows Figure 11 As shown in the image.
[0196] Simulation results show that the flight time control guidance law for maneuvering targets provided in this application can intercept targets with different maneuvering accelerations.
[0197] Example 5
[0198] Three identical aircraft were selected, each equipped with the near-end strategy optimization model obtained in Example 1 and the dual-branch neural network model obtained in Example 2. The control methods for the three aircraft were identical, all including the following steps:
[0199] Step 1: Obtain guidance command a in real time using the offset proportional guidance method. M The guidance command controls the aircraft to fly towards the target, while simultaneously recording and outputting the current flight time t.
[0200] The bias ratio guidance method obtains the guidance command a through the following formula (I). M :
[0201] a M =a0+a b (one)
[0202] Where a0 represents the initial guidance command,
[0203] a b This represents the bias term used to reduce interception time error, which is obtained in real time through the near-end policy optimization model;
[0204] The initial guidance command a0 is obtained through the following equation (ii):
[0205]
[0206] Step 2: Obtain the predicted remaining flight time in real time using a two-branch neural network model.
[0207] Step 3, using the predicted remaining flight time Expected interception time t d And t to obtain the remaining flight time error ε t ,
[0208] ε t =t d -(t+t go )
[0209] Step 4, using the remaining flight time error ε t The near-end strategy optimization model has been revised.
[0210] The expected interception time for all three aircraft is set to t. d =60s;
[0211] The first target performs a constant maneuver, meaning the aircraft's acceleration remains constant. The second target performs a square wave maneuver, and the third target performs a sinusoidal maneuver. The time periods for the square wave and sinusoidal maneuvers are 40 seconds. The initial conditions for all three targets and all three aircraft are identical, set as follows:
[0212]
[0213] The flight trajectories, flight times, speeds, and guidance commands of the three aircraft were simulated, and the resulting flight trajectories are as follows: Figure 12 As shown, flight time is as follows Figure 13 As shown, the speed of motion is as follows Figure 14 As shown, the guidance commands are as follows Figure 15 As shown in the image.
[0214] Simulation results show that the time-of-flight control guidance law for maneuvering targets provided in this application can successfully intercept targets with different maneuvering patterns. It is worth noting that because a constant maneuver is used during the collection of training samples, the prediction accuracy of the two-branch neural network model for interleaved maneuvers decreases. However, as the aircraft approaches the target, the feasible region and the magnitude of change of the target gradually decrease, and t... go The prediction error gradually converges, enabling the aircraft to achieve its target accuracy at t. d The target was successfully intercepted.
[0215] The present invention has been described above with reference to preferred embodiments; however, these embodiments are merely exemplary and illustrative. Various substitutions and modifications can be made to the present invention based on these embodiments, all of which fall within the scope of protection of the present invention.
Claims
1. A time-of-flight control guidance law for maneuvering targets, characterized in that, The guidance law controls the aircraft through the following steps: Step 1: Obtain guidance commands in real time using the offset proportional guidance method. The guidance command controls the aircraft to fly toward the target, while simultaneously recording and outputting the current flight time. , Step 2: Obtain the predicted remaining flight time in real time using a two-branch neural network model. , Step 3, using the predicted remaining flight time The expected interception time pre-filled in the aircraft and Obtain the remaining flight time error , Step 4, using the remaining flight time error Revise the near-end strategy optimization model; Step 5: Repeat steps 1 to 4 in real time to ensure the aircraft follows the guidance commands. Under control, according to the expected interception time. Hit the target; The training samples for the dual-branch neural network model consist of multiple sets of mapping data. The mapping data includes the target's historical trajectory.
2. The flight time control guidance law for maneuvering targets according to claim 1, characterized in that, In step 1, the bias ratio guidance method obtains the guidance command through the following formula (a). : (one) in, This indicates an initial guidance command. This represents the bias term used to reduce interception time error, which is obtained in real time through the near-end policy optimization model.
3. The flight time control guidance law for maneuvering targets according to claim 2, characterized in that, The initial guidance command We obtain it through the following formula (II): (two) in, Indicates the speed of the aircraft. Indicates the trajectory inclination angle of the aircraft. This indicates the angular velocity of the aircraft's line of sight. This indicates acceleration due to gravity.
4. The flight time control guidance law for maneuvering targets according to claim 1, characterized in that, The dual-branch neural network model in step 2 includes an MLP branch and a GRU branch. The MLP branch includes two fully connected hidden layers, each containing 100 neurons; The GRU branch consists of 10 units, each containing 100 neurons.
5. The flight time control method for maneuvering targets according to claim 4, characterized in that, The mapping data is: in, Indicates the remaining flight time. Indicates the relative speed between the aircraft and the target. Indicates the trajectory inclination angle, This indicates the distance between the aircraft and the target along the X-axis. This indicates the distance between the aircraft and the target along the Y-axis. This represents the straight-line distance between the aircraft and the target. Indicates the line-of-sight angle of the bullet. Represents the historical trajectory of the target; That is, the historical trajectory of a target is a set of multiple historical position coordinates of the target. In the target's movement trajectory, starting from the moment when time is 0, the position coordinates of the target are captured at fixed time intervals to obtain the historical trajectory of the target. This represents the target's coordinates in the X direction at the first extracted moment. This indicates the target's coordinates along the X-axis at the current moment. This represents the target's coordinates in the Y direction at the first extracted moment. This indicates the target's coordinates along the Y-axis at the current moment.
6. The flight time control guidance law for maneuvering targets according to claim 5, characterized in that, During the training process of the dual-branch neural network model, As the input to the MLP branch, As input to the GRU branch; During the training of a two-branch neural network model, the network parameters are represented as follows: The loss function is expressed as shown in equation (iii) below. The function, (three) in, Represents the total number of samples. Indicates the first One sample, Indicates based on the first For each sample, the model predicts the remaining flight time. Indicates the first The actual remaining flight time in each sample.