Guidance method of missile party to target party
By establishing a missile-earth motion model and adjusting the proportional guidance coefficient using the deep Q network, the problem of limited guidance accuracy during missile interception is solved, and a higher missile's precise interception capability to targets is achieved.
Patent Information
- Application Number
- CN202510350528.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
During missile interception, the existing proportional guidance law is difficult to ensure guidance accuracy when the maneuverability of missiles and targets.
By establishing a motion model of the bullet-earth motion problem and using a deep Q network for training, the proportional guidance coefficient is dynamically adjusted according to the speed of the missile party, the speed of the tracking party, and the relative distance between the missile party and the target party, to achieve higher guidance accuracy.
Through intelligent decision-making and the optimization of deep Q network, the optimal proportional guidance coefficient can be output under different motion states, improving the missile's precise interception and strike capability to targets, and enhancing guidance accuracy.
Smart Images

Figure CN120212812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of missile guidance, and particularly relates to a guidance method for a missile to a target. Background Art
[0002] The process of a missile intercepting a target consists of three stages: a launch stage of receiving the initial pose information preset by the control system; a mid-course guidance stage of flying to the target area relying on inertial navigation or ground radar according to control instructions; and a terminal guidance stage of adjusting its own flight pose, speed, etc. based on the information of the target obtained by the seeker to achieve precise interception and strike of the target. Among them, the guidance performance and accuracy are directly determined by the guidance and control system of the missile. The missile guidance law, that is, the guidance law, is the key technology to achieve precise strike. Selecting appropriate guidance law parameters is extremely important for the missile to accurately strike the target.
[0003] Currently, the commonly used guidance law is the proportional navigation guidance law. At this time, the guidance law parameter is the proportional navigation coefficient. This guidance law is derived under the condition that the target does not maneuver. In an ideal situation, the proportional navigation guidance law can achieve relatively good guidance effects.
[0004] However, in actual scenarios, considering the maneuvering characteristics of the missile and the target, for example, the moving speeds of the missile and the target may change during the movement process. At this time, the guidance accuracy of the proportional navigation guidance law will be affected to a certain extent, thus affecting the guidance accuracy of the missile to the target. Summary of the Invention
[0005] The purpose of the present invention is to provide a guidance method, device, equipment and medium for a missile to a target, which can provide the best missile proportional navigation coefficient for a specified situation, thereby improving the missile guidance accuracy.
[0006] To solve the above technical problems, an embodiment of the present invention provides a guidance method for a missile to a target, including the following steps: According to the movement process of the missile tracking the target with a proportional navigation coefficient and intercepting and striking the target, establish a motion model of the missile-target motion problem; wherein, the proportional navigation coefficient has a functional relationship with the speed of the missile, the speed of the tracking party, and the relative distance between the missile and the target; Use the speed of the missile, the speed of the tracking party, and the relative distance between the missile and the target as the state input to the deep Q-network, and use the proportional navigation coefficient as the action output by the deep Q-network to train the deep Q-network, so that the trained deep Q-network can output different proportional navigation coefficients according to different speeds of the missile, speeds of the tracking party, and relative distances between the missile and the target for the guidance of the missile to the target; Among them, the deep Q-network is iteratively trained through the following steps: using an exploratory greedy algorithm, according to the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side, as well as the functional relationship between the proportional navigation coefficient and the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side, to select the proportional navigation coefficient; according to the selected proportional navigation coefficient, solve the positions and speeds of the missile side and the tracking side through the motion model; based on a preset loss function, use the selected proportional navigation coefficient and the solved positions and speeds of the missile side and the tracking side to update the network parameters of the deep Q-network.
[0007] In some alternative embodiments, establishing a motion model for the missile-target motion problem according to the motion process of the missile side tracking the target side with a proportional navigation coefficient and intercepting and striking the target side includes: Decouple and synthesize the three-dimensional missile-target motion problem into a missile-target motion problem in the vertical plane and a missile-target motion problem in the horizontal plane; According to the motion process of the missile side and the target side in the vertical plane when the missile side tracks the target side with a proportional navigation coefficient and intercepts and strikes the target side, establish a two-dimensional motion model for the missile-target motion problem.
[0008] In some alternative embodiments, establishing a motion model for the missile-target motion problem according to the motion process of the missile side tracking the target side with an initial proportional navigation coefficient and intercepting and striking the target side includes: Define the line-of-sight coordinate system of the missile side and the line-of-sight coordinate system of the target side; Define the ballistic coordinate system of the missile side, and limit the maneuverability of the missile side and the target side respectively in the ballistic coordinate system of the missile side; Taking the initial line-of-sight coordinate system of the target side as the reference coordinate system, and according to the reference coordinate system, the line-of-sight coordinate system of the missile side, the ballistic coordinate system of the missile side, and the line-of-sight coordinate system of the target side, combined with the motion processes of the missile side and the target side, establish a two-dimensional motion model for the missile-target motion problem.
[0009] In some alternative embodiments, the step of solving the positions and speeds of the missile side and the tracking side through the motion model according to the selected proportional navigation coefficient includes: In the reference coordinate system, use the fourth-order Runge-Kutta method to solve the position and speed of the target side; According to the line-of-sight angle of the missile side relative to the target side, the relative distance between the missile side and the target side, and the selected proportional navigation coefficient, calculate the normal acceleration of the missile side in the line-of-sight coordinate system of the missile side; Based on the limit of the maneuverability of the missile side in the ballistic coordinate system, process the acceleration of the missile side; In the reference coordinate system, the position and velocity of the missile side are calculated based on the processed acceleration of the missile side.
[0010] In some alternative embodiments, the acceleration in the normal direction of the missile side in the line-of-sight coordinate system of the missile side is calculated by the following formula: ; In the formula, represents the acceleration of the missile side, represents the proportional navigation coefficient, represents the relative distance between the missile side and the target side, represents the line-of-sight angle of the missile side relative to the target side.
[0011] In some alternative embodiments, the reward function of the deep Q-network is as follows: ; ; ; ; In the formula, represents the reward function, represents the immediate reward, represents the final reward, represents the relative distance between the missile side and the target side, represents the minimum relative distance between the missile side and the target side, represents the proportional navigation coefficient, represents the line-of-sight angle of the missile side relative to the target side.
[0012] In some alternative embodiments, after updating the network parameters of the deep Q-network based on the preset loss function, using the selected proportional navigation coefficient and the calculated positions and velocities of the missile side and the tracking side, it further includes: According to the selected proportional navigation coefficient, the ballistic inclination angle, acceleration, and line-of-sight angle of the missile side and the target side are calculated through the motion model; According to the ballistic inclination angle, acceleration, and line-of-sight angle of the missile side and the target side, the miss situation of the missile side intercepting and striking the target side is determined; According to the miss situation of the missile side intercepting and striking the target side, the network parameters of the deep Q-network are updated again.
[0013] The guidance method of the missile side to the target side provided by the present invention has at least the following beneficial effects: The present invention performs intelligent decision-making through a neural network - Deep Q-Network, that is, based on the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side, it selects the proportional navigation coefficient, and can output the optimal proportional navigation coefficient under specified conditions, enabling the missile side to achieve guidance for the target side under the specified conditions through this proportional navigation coefficient.
[0014] Specifically, the entire intelligent decision-making process is based on the motion states of the missile side and the target side, that is, the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side. Taking these confrontation initial conditions as the state input into the Deep Q-Network, and the proportional navigation coefficient as the action output by the Deep Q-Network. Then, through the trained Deep Q-Network, different proportional navigation coefficients are output according to different confrontation conditions, taking into account the maneuvering characteristics of the missile side and the target side. Therefore, through this proportional navigation coefficient, precise interception and strike of the target side by the missile side under this motion state can be achieved.
[0015] Moreover, during the training process of the Deep Q-Network, the proportional navigation coefficient is selected through an exploratory greedy algorithm, enabling the exploration rate to be adjusted as the training rounds progress, thereby improving the optimization accuracy of the proportional navigation coefficient and further enhancing the guidance accuracy of the missile. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] One or more embodiments are exemplarily illustrated by the pictures in the corresponding attached drawings, and these exemplary illustrations do not constitute limitations on the embodiments.
[0017] Figure 1 is a flowchart of a guidance method for the missile side to the target side provided according to an embodiment of the present invention Figure 1 ; Figure 2 is a schematic diagram of the scenario of the missile-target motion problem provided according to an embodiment of the present invention; Figure 3 is a schematic diagram of the change of the exploration rate provided according to an embodiment of the present invention; Figure 4 is a training flowchart of a Deep Q-Network provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present invention, many technical details are provided to help readers better understand the present invention. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed by the present invention can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.
[0019] An embodiment of the present invention relates to a guidance method for a missile to target a target. The implementation details of the guidance method for the missile to target the target in this embodiment will be specifically described below. The following content is only provided for convenience of understanding and is not necessary for implementing this solution.
[0020] The specific process of the guidance method for the missile to target the target in this embodiment can be as Figure 1 shown, including: Step 101, establish a motion model of the missile-target motion problem according to the motion process of the missile tracking the target with a proportional navigation coefficient and intercepting and striking the target; wherein, the proportional navigation coefficient has a functional relationship with the speed of the missile, the speed of the tracking party, and the relative distance between the missile and the target.
[0021] Specifically, regarding the missile-target motion problem, as Figure 2 shown, when studying the relative motion of the missile and the target, since the relative distance between the two is much larger than their respective sizes, a simplified treatment is carried out, and both are regarded as particles. The two move towards each other with an initial distance . The missile selects a proportional navigation coefficient through reinforcement learning to track the target in order to achieve precise strikes. When the minimum distance during their motion process is less than the preset distance , it can be considered that the missile has successfully completed the strike. Throughout the process, the moment when the missile starts to track the target is taken as the initial moment , and the basic simulation step size of the scenario is set to 1 ms. When the distance between the two gets closer and closer, in order to avoid missing the moment when the minimum distance is generated due to the excessive speeds of both sides, the simulation step size will decrease as the distance decreases.
[0022] Generally, the interception model between the missile side and the target side needs to be modeled in a three-dimensional space. The three-dimensional missile-target motion problem can be decoupled, and the problem can be discussed through the vertical plane and the horizontal plane. Therefore, in this embodiment, the vertical plane is used for modeling. That is, the three-dimensional missile-target motion problem is decoupled and synthesized into the missile-target motion problem in the vertical plane and the missile-target motion problem in the horizontal plane. Then, according to the missile side tracking the target side with a proportional navigation coefficient and intercepting and striking the target side, the two-dimensional motion model of the missile-target motion problem is established based on the motion process of the missile side and the target side in the vertical plane.
[0023] In the specific implementation, when describing the motion conditions of the missile side and the target side and their motion relationship, introducing a suitable coordinate system can greatly simplify the motion calculation. Therefore, in this embodiment, the line-of-sight coordinate system of the missile side and the line-of-sight coordinate system of the target side are first defined, and the ballistic coordinate system of the missile side is defined. Under the ballistic coordinate system of the missile side, the maneuverability limits of the missile side and the target side are respectively limited. Then, taking the initial line-of-sight coordinate system of the target side as the reference coordinate system, and based on the reference coordinate system, the line-of-sight coordinate system of the missile side, the ballistic coordinate system of the missile side, and the line-of-sight coordinate system of the target side, combined with the motion process of the missile side and the target side, a two-dimensional motion model of the missile-target motion problem is established.
[0024] Among them, regarding the line-of-sight coordinate system, taking the missile side as an example, the line-of-sight coordinate system is defined as follows: The origin is defined as the mass center of the missile. The positive direction of the X-axis points from the mass center of the missile side to the mass center of the target side. The Y-axis is in the vertical plane of the X-axis, and the positive direction is vertically upward. The Z-axis is determined by the right-hand rule. Regarding the ballistic coordinate system, taking the missile side as an example, the ballistic coordinate system is defined as follows: Take the instantaneous mass center of the missile side as the origin of the coordinate system. The positive direction of the X-axis is consistent with its own velocity direction. The Y-axis is in the vertical plane containing the velocity and perpendicular to the X-axis, and the positive direction is upward. The Z-axis is determined by the right-hand rule.
[0025] Step 102: Use the velocity of the missile side, the velocity of the tracking side, and the relative distance between the missile side and the target side as the state input into the deep Q-network, and use the proportional navigation coefficient as the action output by the deep Q-network to train the deep Q-network, so that the trained deep Q-network can output different proportional navigation coefficients according to different velocities of the missile side, velocities of the tracking side, and relative distances between the missile side and the target side for the guidance of the missile side to the target side.
[0026] Among them, the deep Q-network is iteratively trained through the following steps: using an exploratory greedy algorithm, based on the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side, as well as the functional relationship between the proportional navigation coefficient and the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side, the proportional navigation coefficient is selected; according to the selected proportional navigation coefficient, the positions and speeds of the missile side and the tracking side are calculated through the motion model; based on a preset loss function, the network parameters of the deep Q-network are updated using the selected proportional navigation coefficient and the calculated positions and speeds of the missile side and the tracking side.
[0027] In this embodiment, the proportional navigation coefficient of the missile side is optimized based on the deep Q-network to output the optimal proportional navigation coefficient, enabling the missile to achieve precise interception and strike. The deep Q-network, i.e., DQN, is an optimization and extension of Q-learning. Based on deep learning, the Q-table that originally stored action value information is replaced by a neural network to store information, and the deep network is directly used to approximate the action value function, combined with the learning methods of double-network design and experience replay, which is more suitable for high-dimensional situations.
[0028] DQN introduces two neural networks with exactly the same structure, namely the target value network and the estimated value network. Among them, the target value network is responsible for selecting the optimal action according to the action value function, and the estimated value network is responsible for estimating the action value function. At the same time, it learns and updates the network parameters. The target value network does not learn, but copies the parameters of the estimated value network at a fixed period, that is, delayed update, which can reduce the correlation between the target Q value and the current Q value.
[0029] Among them, the update iteration of the network structure parameters is as follows. In actual use, it is directly implicitly updated by the optimizer. ; The loss function of the Q-network is: ; ; ; In the formula, represents the estimated value network parameters, represents the estimated value network parameters, represents the temporal difference target, represents the approximated value function, is the training real-time reward value, represents the discount factor, is the output of the target value network, is the output of the value network.
[0030] Continuously interacting with the environment using policy π will generate experiences. The experience replay part will set up an experience pool to store past experiences. The experience pool has its own capacity. The composition of each data is the current state , action , reward , next state . These experiences can come from different policies, and old data can still be replaced with new data after its capacity is filled. Specifically, when it comes to each training, a batch (batch-size) of experiences is randomly sampled from the experience pool to update the Q-network.
[0031] The deep Q-network designed in this embodiment introduces a target value network and an estimated value network with the same network structure. The structure adopted is as follows: Three hidden layers and one output layer are set up, and function and rectified linear unit are used as activation functions. The optimizer uses the gradient descent method for parameter update and the Adam optimizer for solution. Since the dimensions of the input and output are small, three fully connected layers are directly used for processing, as shown in Table 1:
[0032] Table 1 On the premise of a given network structure, the parameters during network training are set, as shown in Table 2: Table 2 Regarding the selection of Batch_Size, if it is too small, the convergence will be slow, taking a long time, and there will be a serious problem of gradient oscillation, which is not conducive to convergence; if it is too large, there is no obvious gradient change in different data gradient directions, and it is easy to fall into local minima. Considering comprehensively, Batch_Size is selected as 64.
[0033] During the process of reinforcement learning, an exploratory greedy algorithm is adopted. Specifically, in the initial stage of reinforcement learning, since the environment is not familiar, a relatively large exploration rate is required to explore the environment at this stage. As the exploration progresses, that is, as the number of episodes increases, the training gradually shows certain results, and the exploration rate will gradually decrease. By the end of the training, the exploration rate will be reduced to 0.05, and the neural network is relied on to make action selections. For example, the number of episodes is set to 1500, and the change of the exploration rate with the number of episodes is as Figure 3 shown.
[0034] This embodiment shows an environment for reinforcement learning. First, the settings of states and actions under reinforcement learning are as shown in Table 3: The action space in the table is the space composed of action numbers corresponding to different proportional guidance coefficients. Among them, the state space is continuous and the action space is discrete. The action number is selected through reinforcement learning and transformed into the proportional guidance coefficient adopted by the missile side in this round through a function. The number of selected actions is 10, that is, the reasonable coefficient interval range is evenly discretized into 10 numbers.
[0035] In this environment, the reward function for reinforcement learning is set as follows: ; ; ; ; In the formula, represents the reward function, represents the immediate reward, represents the final reward, represents the relative distance between the missile side and the target side, represents the minimum relative distance between the missile side and the target side, represents the proportional guidance coefficient, represents the line-of-sight angle of the missile side relative to the target side.
[0036] It can be seen from the above formula that when the distance between the two (i.e., the relative distance between the missile side and the target side) is between 1000 - 4000m, the missile side will be rewarded according to the line-of-sight angular velocity of the two, combined with the final reward, as a measure of the guidance situation in this round under this coefficient. For whether the target side is finally hit, the rewards given are very different. At the same time, in order to prevent from being too large and affecting the judgment of the guidance by the immediate reward due to attitude factors during the process, a small amount is added on the basis of . In addition, when the line-of-sight angular rate is too large, it can be considered that the tracking of the target side is temporarily lost at this time, so its impact is negative.
[0037] Under the above network structure and network environment, using the speed of the missile side, the speed of the tracking side, and the relative distance between the missile side and the target side as the state input into the deep Q network, and the proportional guidance coefficient as the action output by the deep Q network, training the deep Q network can enable the trained deep Q network to output different proportional guidance coefficients according to different speeds of the missile side, speeds of the tracking side, and relative distances between the missile side and the target side, so as to be used for the guidance of the missile side to the target side.
[0038] In one example, after updating the network parameters of the deep Q-network based on a preset loss function, using the selected proportional navigation coefficient, and the calculated positions and velocities of the missile side and the tracking side, the ballistic inclination angles, accelerations, and line-of-sight angles of the missile side and the target side are calculated through a motion model according to the selected proportional navigation coefficient; the miss situation of the missile side intercepting and striking the target side is determined based on the ballistic inclination angles, accelerations, and line-of-sight angles of the missile side and the target side; and the network parameters of the deep Q-network are updated again according to the miss situation of the missile side intercepting and striking the target side. This can further improve the optimization accuracy of the network.
[0039] In a specific implementation, the training process of the deep Q-network is as Figure 4 shown. The reinforcement learning divides the algorithm into two working states, the training state and the testing state, through the train flag bit.
[0040] In the training state, first, a suitable set of initial information of the missile side and the target side is selected for network initialization, and the of the reinforcement learning is set to 1500. At the beginning of each episode, the environment is initialized through , and the action is selected through an exploratory greedy algorithm, that is, the proportional navigation coefficient of this episode is selected. Then, motion simulation is carried out, and the positions and velocities of the missile side and the target side are calculated through the fourth-order Runge-Kutta method, and the generated data is used to fill the experience pool. When the experience pool is filled to the preset capacity, experiences are sampled, and the estimated value network is updated through the loss function of the two (here the mean square error function MSE) and the previously used optimizer. Every time the estimated value network iterates a predetermined number of times, that is, the target value network update frequency number, the network parameters are completely copied to the target value network. The experience pool is updated more carefully by judging the miss amount, and the old experiences are gradually replaced by the experiences of successful hits. After completing the expected number of episodes, the training of the network is completed.
[0041] If a ready-trained neural network is used to verify the miss amount situation, the test mode is selected. A random set of initial information is selected within a reasonable range of the initial velocities and initial positions of both sides, the network parameters saved after training are read, motion simulation is carried out, and the ballistic inclination angles, accelerations, line-of-sight angle information, etc. of both sides during this test process are plotted into images to analyze the results obtained from the simulation.
[0042] Specifically, the principle of the classical proportional navigation law is that during the process of the missile side intercepting or pursuing the target side, the ratio of the rotational angular velocity of the missile side's velocity to the line-of-sight angular velocity of both sides remains constant, and its guidance relationship equation is: In the formula, represents the missile ballistic angle, represents the missile line-of-sight angle, represents the proportional navigation coefficient.
[0043] When using an improved proportional navigation method, the acceleration of the missile in the normal direction in the line-of-sight coordinate system can be approximately expressed as: ; In the formula, represents the acceleration of the missile side, represents the proportional navigation coefficient, represents the relative distance between the missile side and the target side, represents the line-of-sight angle of the missile side relative to the target side.
[0044] Therefore, when calculating the positions and velocities of the missile side and the tracking side through the motion model according to the selected proportional navigation coefficient, first, in the reference coordinate system, the fourth-order Runge-Kutta method is used to calculate the position and velocity of the target side; according to the line-of-sight angle of the missile side relative to the target side, the relative distance between the missile side and the target side, and the selected proportional navigation coefficient (i.e., the above formula), calculate the acceleration of the missile side in the normal direction in the line-of-sight coordinate system of the missile side; based on the limit of the maneuverability of the missile side in the ballistic coordinate system, process the acceleration of the missile side; in the reference coordinate system, according to the processed acceleration of the missile side, calculate the position and velocity of the missile side.
[0045] In this embodiment, intelligent decision-making is performed through a neural network - deep Q-network, that is, according to the velocity of the missile side, the velocity of the tracking side, and the relative distance between the missile side and the target side, the proportional navigation coefficient is selected, and the optimal proportional navigation coefficient in a specified situation can be output, enabling the missile side to achieve guidance to the target side in this specified situation through this proportional navigation coefficient. Specifically, the entire intelligent decision-making process is based on the motion states of the missile side and the target side, that is, the velocity of the missile side, the velocity of the tracking side, and the relative distance between the missile side and the target side. Taking this confrontation initial condition as the state input into the deep Q-network, and the proportional navigation coefficient as the action output by the deep Q-network. Then, through the trained deep Q-network, different proportional navigation coefficients are output according to different confrontation conditions, taking into account the maneuvering characteristics of the missile side and the target side. Therefore, through this proportional navigation coefficient, precise interception and strike of the target side by the missile side in this motion state can be achieved. Moreover, during the training process of the deep Q-network, a greedy algorithm with exploration is used to select the proportional navigation coefficient, so that the exploration rate can be adjusted as the training rounds progress, thereby improving the optimization accuracy of the proportional navigation coefficient and further improving the guidance accuracy of the missile.
[0046] The step divisions of the above various methods are only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of the present invention; making insignificant modifications to the algorithm or adding insignificant designs to the process, but not changing the core design of the algorithm and process are all within the protection scope of the invention.
[0047] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present invention.
Claims
1. A method for guiding a missile to a target, characterized in that: The method comprises: According to the motion process of the missile tracking the target with the proportional guidance coefficient and intercepting and striking the target, a motion model of the missile-target motion problem is established; wherein the proportional guidance coefficient has a functional relationship with the speed of the missile, the speed of the tracking party, and the relative distance between the missile and the target; The speed of the missile, the speed of the tracking party, and the relative distance between the missile and the target are used as the states input into the deep Q network, and the proportional guidance coefficient is used as the action output by the deep Q network, so that the trained deep Q network outputs different proportional guidance coefficients for guiding the missile to the target according to different speeds of the missile, the speed of the tracking party, and the relative distance between the missile and the target; Among them, the deep Q network is iteratively trained through the following steps: using an exploratory greedy algorithm, the proportional guidance coefficient is selected according to the speed of the missile party, the speed of the tracking party, the relative distance between the missile party and the target party, and the functional relationship between the proportional guidance coefficient and the speed of the missile party, the speed of the tracking party, and the relative distance between the missile party and the target party; according to the selected proportional guidance coefficient, the position and speed of the missile party and the tracking party are solved through the motion model; based on the preset loss function, the network parameters of the deep Q network are updated using the selected proportional guidance coefficient and the solved position and speed of the missile party and the tracking party.
2. The method for guiding a missile to a target according to claim 1, characterized in that: The motion model of the missile-target motion problem is established according to the motion process in which the missile tracks the target with a proportional guidance coefficient and intercepts and strikes the target, including: Decouple the three-dimensional projectile motion problem into the projectile motion problem on the plumb plane and the projectile motion problem on the horizontal plane; According to the motion process of the missile and the target on the plumb plane when the missile tracks the target with the proportional guidance coefficient and intercepts and strikes the target, a two-dimensional motion model of the missile-target motion problem is established.
3. The method for guiding a missile to a target according to claim 1, characterized in that: The motion model of the missile-target motion problem is established according to the motion process in which the missile tracks the target with the initial proportional guidance coefficient and intercepts and strikes the target, including: Define the missile's sight coordinate system and the target's sight coordinate system; Define the missile's ballistic coordinate system, and set limits on the maneuverability of the missile and target in the missile's ballistic coordinate system; The target's initial line of sight coordinate system is taken as the reference coordinate system, and based on the reference coordinate system, the missile's line of sight coordinate system, the missile's ballistic coordinate system and the target's line of sight coordinate system, combined with the motion process of the missile and the target, a two-dimensional motion model of the missile-target motion problem is established.
4. The method for guiding a missile to a target according to claim 3, characterized in that: The method of calculating the position and velocity of the missile and the tracking party through the motion model according to the selected proportional guidance coefficient includes: In the reference coordinate system, the fourth-order Runge-Kutta method is used to solve the position and velocity of the target. According to the sight angle of the missile relative to the target, the relative distance between the missile and the target, and the selected proportional guidance coefficient, the acceleration of the missile in the normal direction in the sight coordinate system of the missile is calculated; Based on the limit of the missile's maneuverability in the ballistic coordinate system, the missile's acceleration is processed; In the reference coordinate system, the position and velocity of the missile are calculated according to the processed acceleration of the missile.
5. The method for guiding a missile to a target according to claim 4, characterized in that: The acceleration of the missile in the normal direction in the missile's line of sight coordinate system is calculated by the following formula: ; In the formula, represents the missile's acceleration, represents the proportional guidance coefficient, Indicates the relative distance between the missile and the target. Indicates the line of sight angle of the missile relative to the target.
6. The method for guiding a missile to a target according to claim 1, characterized in that: The reward function of the deep Q network is as follows: ; ; ; ; In the formula, represents the reward function, Indicates immediate reward, Represents the final reward, Indicates the relative distance between the missile and the target. Indicates the minimum relative distance between the missile and the target. represents the proportional guidance coefficient, Indicates the line of sight angle of the missile relative to the target.
7. The method for guiding a missile to a target according to claim 1, characterized in that: After the network parameters of the deep Q network are updated based on the preset loss function by using the selected proportional guidance coefficient and the calculated positions and velocities of the missile party and the tracking party, the method further includes: According to the selected proportional guidance coefficient, the trajectory inclination, acceleration and sight angle of the missile and the target are calculated through the motion model; Determine the miss of the missile when intercepting the target based on the trajectory inclination, acceleration and sight angle between the missile and the target; According to the miss situation of the missile's interception and attack on the target, the network parameters of the deep Q network are updated again.