A master-slave missile cooperative guidance method and system from a homing head
By calculating the maneuvering state of the missile through an online actor network, the guidance problem when the missile is without a seeker or is damaged is solved, achieving accurate guidance and cost savings.
Patent Information
- Application Number
- CN202310806640.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-07-03
AI Technical Summary
In cases where a missile lacks a seeker or its seeker is damaged, existing technologies cannot achieve precise guidance of the missile, resulting in the missile failing to accurately strike its target.
An online actor network is used for calculations. Based on the relevant parameters of the main missile and its own parameters, the maneuvering state of the secondary missile is determined, thereby controlling the secondary missile for guidance. This avoids using the secondary missile's seeker or improves the guidance effect when the seeker is damaged.
It enables accurate guidance even when the missile is without a seeker or the seeker is damaged, saving costs and improving guidance effectiveness.
Smart Images

Figure CN116952076B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of guided control, in particular to a master-slave missile cooperative guidance method and system without seeker of slave missile. BACKGROUND
[0002] Currently, the cooperative guidance among multiple missiles is usually performed by one missile as a master missile and other missiles as slave missiles, and the master missile leads the slave missiles to perform guidance. However, by this method, both the master missile and the slave missiles need to use seekers. If the slave missile has no seeker or the seeker of the slave missile is damaged by interference during flight, the slave missile cannot accurately attack the target. Therefore, how to guide the slave missile in the case that the slave missile has no seeker or the seeker of the slave missile is damaged becomes a problem. SUMMARY
[0003] In order to accurately guide the slave missile in the case that the slave missile has no seeker or the seeker of the slave missile is damaged, the present application provides a master-slave missile cooperative guidance method and system without seeker of slave missile.
[0004] In a first aspect, the present application provides a master-slave missile cooperative guidance method without seeker of slave missile, which adopts the following technical solution:
[0005] A master-slave missile cooperative guidance method without seeker of slave missile, comprising:
[0006] obtaining actual combat state parameters, wherein the actual combat state parameters include related parameters between the master missile and the target, related parameters between the master missile and the slave missile, related parameters of the master missile itself, and related parameters of the slave missile itself;
[0007] inputting the actual combat state parameters into a trained online actor network to perform calculation to obtain a maneuvering state of the slave missile, wherein the trained online actor network is obtained by updating and training an online network and a target network according to obtained training samples;
[0008] controlling the slave missile to operate according to the maneuvering state.
[0009] By adopting the above technical solution, the related parameters between the master missile and the target, the related parameters of the master missile itself, the related parameters of the slave missile itself, and the related parameters between the slave missile and the master missile are inputted into the trained online actor network to perform calculation, so that the maneuvering state of the slave missile can be accurately obtained, and then the slave missile is controlled to operate according to the obtained maneuvering state, thereby guiding the slave missile. Since the guidance does not use the related parameters between the target and the slave missile, the slave missile does not need to install a seeker to perform guidance, or the guidance is performed in the case that the seeker of the slave missile is interfered or damaged. Compared with the guidance performed by the slave missile installing a seeker, the cost can be saved, and the guidance effect is improved when the seeker of the slave missile is interfered or damaged.
[0010] In another possible implementation manner, the inputting the actual combat state parameter into the trained online actor network for calculation to obtain the maneuvering state of the slave projectile, and the controlling the slave projectile to run according to the maneuvering state, comprises: inputting the actual combat state parameter into the trained online actor network for calculation to obtain the maneuvering state of the slave projectile, and controlling the slave projectile to run according to the maneuvering state.
[0011] If the slave projectile does not hit the target and the target does not escape, the following loop steps are performed until a first preset condition is met:
[0012] obtaining a current actual combat state parameter;
[0013] inputting the current actual combat state parameter into an online network for state calculation to obtain a maneuvering state of the slave projectile at a next time;
[0014] controlling the slave projectile to run according to the maneuvering state at the next time;
[0015] The first preset condition comprises any one of the following:
[0016] The target escapes.
[0017] The slave projectile hits the target.
[0018] By using the above technical solution, if the slave projectile does not hit the target and the target does not escape, the maneuvering state of the slave projectile at the next time is calculated in real time according to the current actual combat state parameter, and the maneuvering state of the slave projectile at the next time is calculated in combination with actual changes of the battlefield environment, so that the slave projectile can hit the target more accurately, and the guidance success rate is improved.
[0019] In a second aspect, the present application provides a slave projectile without a seeker for cooperative guidance training of a master-slave projectile, and the following technical solution is adopted:
[0020] A slave projectile without a seeker for cooperative guidance training of a master-slave projectile, comprising:
[0021] obtaining a plurality of training samples, wherein the training samples comprise a state parameter, a maneuvering state of a slave projectile, a reward value, and a state parameter at a next time;
[0022] inputting the state parameter into an online network to evaluate a benefit to obtain a first evaluation value;
[0023] inputting the state parameter at the next time into a target network to evaluate a benefit to obtain a second evaluation value;
[0024] update the online network and the target network based on the first evaluation value, the second evaluation value and the reward value, to obtain a trained online network and a trained target network, the trained online network comprising a trained online actor network, the trained online actor network being configured to calculate the maneuvering state of the slave projectile based on the state parameters obtained from the slave projectile.
[0025] By adopting the technical solution, after obtaining the training sample, the state parameters in the training sample are input into the online network for benefit evaluation, so as to obtain the first evaluation value of the online network. The state parameters at the next moment in the training sample are input into the target network for benefit evaluation, so as to obtain the second evaluation value of the target network. The first evaluation value is used to represent the training situation of the online network, and the second evaluation value is used to represent the training situation of the target network. The online network and the target network are updated and corrected in combination with the reward value in the training sample, so as to obtain the trained online network and the trained target network. The trained online network comprises a trained online actor network, which is used to calculate the maneuvering state of the slave projectile based on the state parameters obtained from the slave projectile, so that the slave projectile can be accurately guided without installing a guidance head or the guidance head being damaged.
[0026] In another possible implementation manner, the state parameters comprise relevant parameters between the main projectile and the target, relevant parameters between the main projectile and the slave projectile, self-related parameters of the main projectile, self-related parameters of the slave projectile and target-related parameters; the training sample is input into the online network for benefit evaluation to obtain the first evaluation value; and the training sample is input into the target network for benefit evaluation to obtain the second evaluation value, which comprises:
[0027] The state parameters are input into the online actor network for calculation to obtain the target maneuvering state of the slave projectile;
[0028] The maneuvering state and the target maneuvering state are input into the online critic network for benefit evaluation to obtain the first evaluation value;
[0029] The state parameters at the next moment are input into the target actor network for calculation to obtain the target maneuvering state of the slave projectile at the next moment;
[0030] The target maneuvering state at the next moment and the maneuvering state at the next moment are input into the target critic network for benefit evaluation to obtain the second evaluation value.
[0031] By adopting the technical solution, the first evaluation value is used to evaluate the online actor network, and the second evaluation value is used to evaluate the target actor network. The quality of the online actor network and the target actor network can be more intuitively and accurately evaluated through the first evaluation value and the second evaluation value.
[0032] In another possible implementation manner, the online network is updated based on the first evaluation value, the second evaluation value and the reward value, including:
[0033] The online critic network is updated based on the first evaluation value, the second evaluation value and the reward value.
[0034] The online actor network is updated based on the updated online critic network.
[0035] The following loop steps are performed until a second preset condition is reached:
[0036] The state parameters in a next training sample are input into the current online network to evaluate the return, to obtain a current first evaluation value.
[0037] The state parameters at a next time in the next training sample are input into the current target network to evaluate the return, to obtain a current second evaluation value.
[0038] The current online critic network is updated based on the current first evaluation value, the current second evaluation value and the reward value in the next training sample.
[0039] The current online actor network is updated based on the current online critic network.
[0040] The current online critic network is the online critic network after the last update, and the current online actor network is the online actor network after the last update.
[0041] The second preset condition includes at least one of the following:
[0042] The next training sample is the last training sample.
[0043] The hit time difference is the minimum, and the hit time difference is the difference between the main projectile hit time and the slave projectile hit time.
[0044] By using the above technical solution, the first evaluation value and the second evaluation value are calculated by using multiple training samples in a loop, and then combined with the reward value in each training sample, the network can be rewarded more frequently, so that the final obtained network is more accurate.
[0045] In another possible implementation manner, the method further includes:
[0046] When the number of updates reaches a specified number of times, the target network is updated based on the online network at the specified number of times.
[0047] In another possible implementation manner, the reward value includes a first reward value and a second reward value, and the method further includes:
[0048] obtaining a state parameter;
[0049] simulating the state parameter to obtain a maneuvering state of the slave projectile;
[0050] judging whether the slave projectile hits the target based on the maneuvering state;
[0051] if the target is not hit, determining a first reward value based on the state parameter;
[0052] determining a state parameter of the slave projectile at a next time based on the state parameter and the maneuvering state;
[0053] generating a training sample based on the state parameter, the maneuvering state, the first reward value, and the state parameter at the next time;
[0054] performing the following loop steps until a third preset condition is met:
[0055] simulating the current state parameter to obtain a current maneuvering state of the slave projectile, the current state parameter being a state parameter of the slave projectile at the next time determined based on a state parameter in a last simulation period and the maneuvering state;
[0056] if the target is not hit, determining a current first reward value based on the current state parameter;
[0057] determining a state parameter of the slave projectile at the next time based on the current state parameter and the current maneuvering state;
[0058] generating a training sample based on the current state parameter, the current maneuvering state, the current first reward value, and the state parameter at the next time;
[0059] the third preset condition includes any one of the following:
[0060] the slave projectile hits the target;
[0061] the target escapes;
[0062] a number of times of determining the current first reward value reaches a preset number threshold;
[0063] if the slave projectile hits the target, determining a first hit time of the main projectile and a second hit time of the slave projectile; and calculating a second reward value based on the first hit time of the main projectile and the second hit time of the slave projectile;
[0064] generating a training sample based on the current state parameter, the current maneuvering state, the second reward value, and the state parameter at the next time.
[0065] By using the above technical solution, the state parameters of each state are used to simulate the state of the slave missile, the state parameters of the next time and the reward value, so that the corresponding training sample can be generated using the state parameters of each time, and a large number of training samples can be obtained more conveniently.
[0066] In another possible implementation manner, the state parameters include a lead angle of the main missile, a lead angle of the slave missile, a relative distance between the main missile and the target, a relative distance between the slave missile and the target, a speed of the main missile and a speed of the slave missile, and the first reward value is determined based on the state parameters, including:
[0067] predicting a first hitting time of the main missile hitting the target based on the lead angle of the main missile, the relative distance between the main missile and the target and the speed of the main missile, and predicting a second hitting time of the slave missile hitting the target based on the lead angle of the slave missile, the relative distance between the slave missile and the target and the speed of the slave missile;
[0068] determining a time difference between the first hitting time and the second hitting time;
[0069] determining the first reward value based on the time difference.
[0070] By using the above technical solution, the first hitting time and the second hitting time are calculated based on the state parameters in each training sample, and the first reward value is determined according to the time difference between the two, so that the reward value is determined at each time of the simulation movement of the slave missile, so that the network can be rewarded more densely and frequently, and the training quality of the network is improved.
[0071] In a third aspect, the present application provides a master-slave missile cooperative guidance system without a seeker for a slave missile, which adopts the following technical solution:
[0072] A master-slave missile cooperative guidance system without a seeker for a slave missile, comprising:
[0073] a main missile, which collects main missile related parameters and related parameters between the main missile and a target, and sends the main missile related parameters and the related parameters between the main missile and the target to a slave missile;
[0074] a slave missile, which acquires the main missile related parameters and the related parameters between the main missile and the target, inputs the related parameters between the main missile and the target, the related parameters between the main missile and the slave missile, main missile self-related parameters and slave missile self-related parameters into a trained online actor network for calculation to obtain a maneuvering state of the slave missile, determines the maneuvering state of the slave missile based on the state parameters, and controls the slave missile to operate according to the maneuvering state.
[0075] In a fourth aspect, the present application provides an electronic device, which adopts the following technical solution:
[0076] An electronic device, comprising:
[0077] at least one processor;
[0078] a memory;
[0079] at least one application, wherein the at least one application is stored in the memory and configured to be executed by the at least one processor, and the at least one application is configured to implement a method for training a cooperative guidance of a slave missile from a homing head according to any possible implementation of the second aspect.
[0080] In a fifth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution:
[0081] A computer-readable storage medium, when the computer program is executed in the computer, the computer executes the method for training a cooperative guidance of a slave missile from a homing head according to any one of the second aspect.
[0082] In summary, the present application includes at least one of the following beneficial technical effects:
[0083] 1. The relevant parameters of the target to the master missile, the relevant parameters of the master missile itself, the relevant parameters of the slave missile itself, and the relevant parameters of the slave missile to the master missile are input into the trained online actor network for calculation, so that the maneuvering state of the slave missile can be accurately obtained, and then the slave missile is controlled to operate according to the obtained maneuvering state, thereby guiding the slave missile. Since the guidance does not use the relevant parameters of the target to the slave missile, the slave missile does not need to install a homing head to guide, or guides in the case that the homing head of the slave missile is disturbed or damaged. Compared with the guidance of the slave missile with the homing head, the cost can be saved, and the guidance effect is improved when the homing head of the slave missile is disturbed or damaged.
[0084] 2. After obtaining the training sample, the state parameters in the training sample are input into the online network for benefit evaluation, so as to obtain a first evaluation value of the online network. The state parameters of the next moment in the training sample are input into the target network for benefit evaluation, so as to obtain a second evaluation value of the target network. The first evaluation value is used to represent the training situation of the online network, and the second evaluation value is used to represent the training situation of the target network. The online network and the target network are updated and corrected in combination with the reward value in the training sample, so as to obtain the trained online network and the target network. The trained online network includes a trained online actor network, which is used to calculate the maneuvering state of the slave missile according to the state parameters obtained by the slave missile, so that the slave missile can be accurately guided without installing a homing head or in the case that the homing head is damaged. BRIEF DESCRIPTION OF DRAWINGS
[0085] Figure 1Figure 1 is a flowchart of a method for cooperative guidance of a slave missile by a master missile according to an embodiment of the present application.
[0086] Figure 2 Figure 2 is a flowchart of a method for cooperative guidance of a slave missile by a master missile according to an embodiment of the present application.
[0087] Figure 3 Figure 3 is a flowchart of a method for training of a slave missile by a master missile according to an embodiment of the present application.
[0088] Figure 4 Figure 4 is a schematic diagram of an overall framework for updating an online network and a target network according to an embodiment of the present application.
[0089] Figure 5 Figure 5 is a flowchart of a method for training of a slave missile by a master missile according to an embodiment of the present application.
[0090] Figure 6 Figure 6 is a schematic diagram of a system for cooperative guidance of a slave missile by a master missile according to an embodiment of the present application.
[0091] Figure 7 Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present application.
[0092] BRIEF DESCRIPTION OF DRAWINGS: 11, master missile; 12, slave missile; 13, target; 30, electronic device; 301, processor; 302, bus; 303, memory; 304, transceiver. DETAILED DESCRIPTION
[0093] The present application will be further described below in conjunction with the accompanying drawings.
[0094] Any modifications made by those of ordinary skill in the art based on the description of the present application are within the scope of the present application, and are protected by the patent law.
[0095] In order to make the objectives, technical solutions, and superiorities of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.
[0096] In addition, the term "and / or" in this document merely describes an associated relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " in this document generally represents an "or" relationship between the front and rear associated objects unless otherwise specified.
[0097] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0098] The embodiments of the present application provide a master-slave missile cooperative guidance method from a homing head, which is executed by a slave missile, as shown in the formula (1): Figure 1 The method comprises steps S101, S102 and S103, wherein,
[0099] S101, acquiring a real combat state parameter.
[0100] The real combat state parameter comprises a related parameter between the master missile and the target, a related parameter between the master missile and the slave missile, a self-related parameter of the master missile, and a self-related parameter of the slave missile.
[0101] Specifically, the related parameter between the master missile and the target is collected by the homing head of the master missile, and the self-related parameter of the master missile can be collected by sensors on the master missile, such as a speed sensor, a gyroscope and an angle sensor. The master missile and the slave missile are linked through wireless communication, and the master missile sends the self-related parameter of the master missile and the related parameter between the master missile and the target to the slave missile through wireless transmission, so that the slave missile acquires the above-mentioned parameters.
[0102] The self-related parameter of the slave missile can be collected by sensors on the slave missile, such as a speed sensor, a gyroscope and an angle sensor, so that the slave missile acquires the self-related parameter of the slave missile. The related parameter between the slave missile and the master missile can be collected by sensors on the master missile or by sensors on the slave missile.
[0103] In the embodiments of the present application, the real combat state parameter s comprises:
[0104] s = (r LT , q LT , ψ L , v L , r LF , q LF , ψ F , v F , v T )
[0105] Wherein, r LT is the relative speed of the master missile to the target; q LT is the line-of-sight angle of the master missile to the target; ψ L is the heading angle of the master missile; vL Vp is the velocity of the primary missile; r LF is the distance between the primary missile and the secondary missile; q LF is the line-of-sight angle between the primary missile and the secondary missile; ψ F is the heading angle of the target; v F is the velocity of the secondary missile; v T is the velocity of the target.
[0106] The interception guidance law of the primary missile is designed by taking the horizontal plane as an example, and the interception guidance law of the primary missile is as follows:
[0107]
[0108] wherein, η T is the target heading angle; η L is the lead angle of the primary missile to the target; ψ T is the target heading angle; a L is the acceleration of the primary missile. For the primary missile, the interception guidance law of the primary missile is as follows: a L = f L (r LT , q LT , ψ L , v L , v T ) = Nv L q LT . Wherein, N is a proportional guidance coefficient. Therefore, after the primary missile collects the relevant parameters of the primary missile itself, the relevant parameters of the target and the relevant parameters between the primary missile and the target, the acceleration a L of the primary missile can be calculated according to the interception guidance law of the primary missile. Since the acceleration a F of the secondary missile needs to be calculated according to the relevant parameters between the primary missile and the target, the missile with a larger remaining time to hit the target is selected as the primary missile, and other missiles are selected as the secondary missiles, so as to avoid the situation that the secondary missile cannot cooperate.
[0109] S102, inputting the actual combat state parameters into the trained online actor network for calculation to obtain the maneuvering state of the secondary missile.
[0110] Wherein, the trained online actor network is obtained by updating and training the online network and the target network according to the obtained training samples.
[0111] After the secondary missile obtains the state parameters, the state parameters are input into the trained online actor network for state calculation, so as to obtain the maneuvering state of the secondary missile. Specifically, the maneuvering state can be the acceleration of the secondary missile, and the acceleration is a vector. According to the acceleration, the secondary missile can be controlled to make corresponding maneuvering actions.
[0112] Specifically, the trained online actor network is obtained by updating the online network and the target network according to the obtained training sample, the online network including the online actor network, and the online network and the target network are trained and updated, so as to obtain the trained online actor network.
[0113] In the embodiment of the present application, s = (r LT , q LT , ψ L , v L , r LF , q LF , ψ F , v F , v) T are input into the trained online actor network for state calculation, that is, the maneuvering state of the seeker is obtained, that is, the acceleration a F of the seeker is obtained, that is, a F = f RL (r LT , q LT , ψ L , v L , r LF , q LF , ψ F , v F , v T ).
[0114] S103, controlling the seeker to run according to the maneuvering state.
[0115] Specifically, after obtaining the maneuvering state of the seeker, the seeker sends instructions to the engine and the tail wing of the seeker itself, so as to control the guidance.
[0116] In one possible implementation of the embodiment of the present application, in steps S102 and S103, the state parameters are input into the trained online actor network for calculation to obtain the maneuvering state of the seeker, and the seeker is controlled to run according to the maneuvering state, which specifically includes steps S1a (not shown in the figure), S1b (not shown in the figure) and S1c (not shown in the figure), wherein,
[0117] S1a, inputting the actual combat state parameters into the trained online actor network for calculation to obtain the maneuvering state of the seeker.
[0118] S1b, controlling the seeker to run according to the maneuvering state.
[0119] S1c, if the seeker does not hit the target and the target does not escape, the following loop steps are executed until the first preset condition is met:
[0120] obtaining the current actual combat state parameters;
[0121] The current combat status parameters are input into the online actor network for status calculation to obtain the current maneuver status of the missile; the missile is then controlled to operate according to the current maneuver status.
[0122] The first preset condition includes any one of the following:
[0123] The target escaped;
[0124] The bullet hit the target.
[0125] For the embodiments of this application, refer to Figure 2 , Figure 2 Before coordinating guidance, the parameters are initialized, and then the main missile calculates its maneuvering state a. L The a value of the missile is calculated from the missile through an actor network. F The missile sends commands to its maneuvering components, such as the engine and tail fins, to induce maneuvers and thus provide guidance. This guides the missile according to a... F After operation, it is necessary to determine whether the missile hits the target. If it hits the target, it means that the guidance is successful. If it misses the target, it is necessary to determine whether the target escapes. Specifically, the relative distance between the main missile and the target at the current moment can be compared with the relative distance between the main missile and the target at the previous moment. If the relative distance increases, it means that the target is getting farther and farther away from the main missile, which further indicates that the missile is getting farther and farther away from the target, and the target has escaped and is no longer guided.
[0126] If it is determined that the missile missed the target and the target did not escape, the missile acquires its own information according to a. F The state parameters after the maneuver, the main missile according to a L State parameters after maneuver and target velocity v T This refers to the current combat state parameters. These parameters are then input into a trained online actor network for state calculation, yielding the required maneuver state 'a' for the missile. F Get the current maneuver state a F Then, the missile controls its own maneuvering components, such as the engine and tail fins, to move in accordance with a F The missile then performs a maneuver. It continues to determine whether the missile hits the target and whether the target escapes. If the missile misses and the target does not escape, it continues to acquire the current combat status parameters and calculate the missile's current maneuver state, controlling the missile to maneuver according to this state. If the missile still misses the target after the maneuver and the target does not escape, it continues to acquire the current combat status parameters and calculate the missile's current maneuver state... and so on, until the missile hits the target or the target escapes, at which point the calculation of the missile's maneuver state stops.
[0127] Similarly, when it is judged that the main missile does not hit the target and the target does not escape, the main missile obtains its own state parameters according to a L the state parameters after the maneuver and the target speed v T Then, the relevant parameters of the main missile and the relevant parameters of the target are input into the interception guidance law of the main missile to calculate the maneuver state a L of the main missile required at present, and the current state parameters are obtained. L After that, the maneuvering part of the main missile such as the engine and the tail wing is controlled to maneuver according to a L . Then, it is continuously judged whether the main missile hits the target and whether the target escapes. If the main missile still does not hit the target and the target does not escape, the current state parameters are continuously obtained and the current maneuver state of the main missile is calculated to control the main missile to maneuver according to the current maneuver state. After the maneuver, the main missile still does not hit the target and the target does not escape, and the current state parameters are still obtained and the current maneuver state of the main missile is calculated through the interception guidance law, and so on until the main missile hits the target or the target escapes.
[0128] The embodiment of the present application provides a master-slave missile cooperative guidance training method from a missile without a seeker, which is executed by an electronic device. The electronic device can be a server or a terminal device. The server can be an independent physical server, a server cluster composed of multiple physical servers or a distributed system, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited to this. The terminal device and the server can be directly or indirectly connected through wired or wireless communication, and the present application is not limited in this regard. As shown in the figure, the method comprises steps S201, S202, S203 and S204, wherein, Figure 3
[0129] S201, a plurality of training samples are obtained.
[0130] The training sample comprises state parameters, a maneuver state of the slave missile, a reward value and state parameters at a next moment. First, a Markov decision model <s, a, r, s'> for the slave missile is established, i.e., a training sample, wherein s is the state parameters at a current moment; a is the acceleration of the slave missile; r is the reward value; and s' is the state parameters at a next moment.
[0131] For the embodiment of the present application, a person can connect the electronic device through a mobile storage device or the like, and write a sample set comprising a plurality of training samples into the electronic device, so that the electronic device trains and updates the network using the plurality of training samples.
[0132] S202, input the state parameter into the online network to evaluate the benefit, and obtain a first evaluation value.
[0133] For the embodiment of the present application, the first evaluation value obtained by inputting a training sample into the online network for benefit evaluation is used to evaluate the training quality of the online network.
[0134] S203, input the state parameter of the next moment into the target network to evaluate the benefit, and obtain a second evaluation value.
[0135] For the embodiment of the present application, taking step S202 as an example, the target network evaluates the benefit according to the state parameter of the next moment in the training sample, and the second evaluation value obtained is used to evaluate the training quality of the target network.
[0136] S204, update the online network and the target network based on the first evaluation value, the second evaluation value and the reward value, and obtain the trained online network and the trained target network.
[0137] The trained online network includes a trained online actor network, and the trained online actor network is used to calculate the maneuvering state of the sub-missile according to the state parameter obtained from the sub-missile.
[0138] For the embodiment of the present application, after obtaining the first evaluation value and the second evaluation value, the online network and the target network are updated in combination with the reward value in the training sample currently used, so as to obtain the trained online network and the trained target network. The trained online network includes a trained online actor network, and the trained online actor is used to calculate the maneuvering state of the sub-missile according to the state parameter obtained in the actual combat scene.
[0139] In one possible implementation of the embodiment of the present application, the state parameter includes the related parameter between the main missile and the target, the related parameter between the main missile and the sub-missile, the related parameter of the main missile itself, the related parameter of the sub-missile itself and the related parameter of the target; the step S202 and the step S203 include steps S2a (not shown in the figure), S2b (not shown in the figure), S2c (not shown in the figure) and S2d (not shown in the figure), wherein S2a, the state parameter is input into the online actor network for calculation, and the target maneuvering state of the sub-missile is obtained.
[0140] For the embodiment of the present application, refer to Figure 4The state parameter s in the training sample is input into the online actor network for calculation to obtain the target maneuver state a of the interceptor. Specifically, the online actor network can include a basic interception guidance law, the online actor network is updated based on the basic interception guidance law, and thus a trained online actor network is obtained. Further, when the online actor network outputs the target maneuver state of the interceptor, random noise can also be added to the target maneuver state, and thus a policy mapping is formed.
[0141] S2b, the state parameter and the target maneuver state are input into the online critic network to evaluate the reward, and a first evaluation value is obtained.
[0142] For the embodiment of the present application, as shown in Figure 4 the state parameter in the training sample and the target maneuver state output by the online actor network are input into the online critic network for reward evaluation, and a first evaluation value about the online actor network is obtained.
[0143] Further, as shown in Figure 4 the online critic network can include two critic networks, critic1 network and critic2 network, and thus two evaluation values can be obtained. In order to prevent overestimation of the online actor, the minimum evaluation value of the two evaluation values is determined as the first evaluation value. Specifically, the structures of the critic1 network and the critic2 network are the same.
[0144] S2c, the state parameter at the next time is input into the target actor network for calculation to obtain the target maneuver state of the interceptor at the next time.
[0145] For the embodiment of the present application, as shown in Figure 4 the state parameter s' at the next time in the training sample is input into the target actor network for calculation to obtain the target maneuver state a' of the interceptor at the next time. Specifically, the structure of the target actor network can be the same as that of the online actor network, that is, the target actor network is updated based on the basic interception guidance law, and thus a trained target actor network is obtained.
[0146] S2d, the state parameter at the next time and the target maneuver state at the next time are input into the target critic network to evaluate the reward, and a second evaluation value is obtained.
[0147] For the embodiment of the present application, as shown in Figure 4 the state parameter at the next time in the training sample and the target maneuver state at the next time output by the online actor network are input into the target critic network for reward evaluation, and a second evaluation value about the target actor network is obtained.
[0148] Further, the target critic network can include two critic networks, a critic3 network and a critic4 network, so that two evaluation values can be obtained, and in order to prevent the target actor from being evaluated too high, the minimum evaluation value of the two evaluation values is determined as the second evaluation value. Specifically, the critic3 network and the critic4 network have the same structure, and further, the structure of the target critic network can be the same as that of the online critic network.
[0149] In an embodiment of the present application, the online network and the target network are updated based on the first evaluation value, the second evaluation value and the reward value in step S204 to obtain trained online network and target network, specifically including steps S2041 (not shown in the figure), S2042 (not shown in the figure), S2043 (not shown in the figure) and S2044 (not shown in the figure), wherein,
[0150] S2041, updating the online critic network based on the first evaluation value, the second evaluation value and the reward value.
[0151] Specifically, the online critic network is updated by minimizing the loss function, and the loss function is:
[0152]
[0153] wherein N is the number of training samples; r i is the reward value of the i-th training sample; γ is the reward discount factor; Q' is the second evaluation value; s i+1 is the state parameter of the next time of the i-th sample; a i+1 is the maneuver state of the next time corresponding to the i-th sample; ω' is the network weight parameter of the target critic network; Q is the first evaluation value; s i is the state parameter of the current time of the i-th sample; a i is the maneuver state of the current time of the i-th sample; ω is the network weight parameter of the online critic network.
[0154] The electronic device calculates γQ'(s i+1 ,a i+1 |ω'), that is, the second evaluation value of the next time maneuver state input by the target critic network in the next time state parameter, and then calculates the product of the second evaluation value and the reward discount factor. And calculate Q(s i ,a i|ω), i.e. the first evaluation value of the maneuver state at the current state parameter under the input of the online critic network. Then the sum of (r i + γQ'(s i+1 ,a i+1 |ω') - Q(s i ,a i |ω) 2 is summed up and divided by the number of training samples to obtain the loss value.
[0155] S2042, updating the online actor network based on the updated online critic network.
[0156] Specifically, the online actor network update adopts a policy gradient update, and the specific policy gradient is represented as:
[0157]
[0158] Wherein, N is the number of training samples; θ is the network weight parameter of the online actor network; μ is the online actor network.
[0159] The electronic device calculates i.e. the gradient of the online actor network at the current state parameter s i with respect to the online actor network weight parameter, and then calculates i.e. the gradient of the first evaluation value with respect to the output of the online actor network and the current state parameter s i of the maneuver state, and then the sum of of the N training samples used in a training batch is summed up and divided by the number of training samples, so as to obtain the updated policy gradient.
[0160] S2043, performing the following loop steps until the second preset condition is reached:
[0161] Inputting the state parameter in the next training sample into the current online network to evaluate the return to obtain the current first evaluation value;
[0162] Inputting the state parameter at the next time in the next training sample into the current target network to evaluate the return to obtain the current second evaluation value;
[0163] Updating the current online critic network based on the current first evaluation value, the current second evaluation value and the reward value in the next training sample;
[0164] Updating the current online actor network based on the current online critic network.
[0165] The current online critic network is the online critic network after the last update, and the current online actor network is the online actor network after the last update.
[0166] The second preset condition comprises at least one of the following:
[0167] The next training sample is the last training sample.
[0168] The hit time difference is the difference between the hit time of the primary projectile and the hit time of the secondary projectile.
[0169] For the embodiments of the present application, refer to Figure 5 Before training the online network and the target network, relevant parameters are set, for example, the number of layers of the online network and the target network respectively, the dimension of the output feature map of each intermediate layer and the used activation function, the learning rate of the online network and the target network, the number N (batch size) of samples used for each batch of training, the size of the experience pool, the network noise, the network noise decay rate, and the like. The speed and overload of the missile are set, and the speed, overload and initial position of the target are set.
[0170] After the above parameters are set, the parameters are initialized, and the target controller controls the target flight in the training environment according to the set relevant parameters. The primary projectile calculates the maneuvering state according to the set interception guidance law and performs guidance.
[0171] When the new training sample is used, it is equivalent to using the new training sample to perform simulation in the training environment. The simulation process is similar to the real cooperative guidance process. In the simulation process, it is determined whether the primary projectile or the secondary projectile hits the target, and then it is determined whether the primary projectile or the secondary projectile escapes. If the primary projectile or the secondary projectile hits the target or the target escapes, it is determined whether the second preset condition is met. In the simulation process, the new training sample can be processed according to the manner in steps S2a to S2b to obtain the current first evaluation value and the current second evaluation value. Then, the current online critic network and the current online actor network can be updated according to the manner in steps S2041 to S2042. The specific process is not described again. Until the N training samples are used up and / or the hit time difference is minimum.
[0172] In the embodiments of the present application, the minimum hit time difference can be represented as: Wherein, t 0end is the time when the primary projectile hits the target, and t iend is the time when the secondary projectile hits the target. If the hit time difference is minimum, it means that the closer the hit times of the primary projectile and the secondary projectile are, the better the cooperative guidance effect of the primary projectile and the secondary projectile is.
[0173] Specifically, in the training process, the state parameters in the training samples can be used to predict the hit time of the main projectile and the hit time of the sub-projectile.
[0174] In a possible implementation of the embodiment of the application, the method further includes updating the target network based on the online network at the specified time when the number of updates reaches the specified number.
[0175] For the embodiment of the application, the DDPG algorithm updates the target network in a soft update manner to improve the stability of the learning process. The update amplitudes of the target actor network and the target critic network are respectively:
[0176]
[0177]
[0178] wherein μ' is the network weight parameter of the target actor network; ω' is the network weight parameter of the target critic network; τ represents the update speed, τ << 1. That is, the target network is updated in a soft update manner, as shown in the following formula: Figure 4 That is, when the specified number is reached, the target network is updated according to the online network at the specified time.
[0179] In a possible implementation of the embodiment of the application, the reward value includes a first reward value and a second reward value, and the method further includes steps S1 (not shown in the figure), S2 (not shown in the figure), S3 (not shown in the figure), S4 (not shown in the figure), S5 (not shown in the figure), S6 (not shown in the figure), S7 (not shown in the figure), S8 (not shown in the figure), S9 (not shown in the figure) and S10 (not shown in the figure), wherein,
[0180] S1, obtaining a state parameter.
[0181] The state parameter can be a state parameter input by a person to an electronic device through a mouse, a keyboard, a touch screen or the like.
[0182] S2, simulating the state parameter to obtain a maneuvering state of a sub-projectile.
[0183] For the embodiment of the application, after obtaining the state parameter, the maneuvering state of the sub-projectile is obtained by prediction through an online actor network or a classical interception guidance law.
[0184] S3, determining whether the sub-projectile hits a target based on the maneuvering state.
[0185] For the embodiment of the application, whether the sub-projectile hits the target is determined after simulating the maneuvering state in a training environment.
[0186] S4, if the target is not hit, determining a first reward value based on the state parameter.
[0187] For the embodiment of the present application, if the target is not hit, it indicates that the first reward value corresponding to the current state parameter needs to be determined. Specifically,
[0188] R1=e-|(t Lgo -t Fgo )|-1
[0189] wherein R1 is the reward when the target is not hit, and is related to the predicted main projectile hit time t Lgo and the sub-projectile hit time t Fgo ; e is the natural logarithm, which can be set by the staff according to the situation.
[0190] Since there is a time difference between the main projectile hit time and the sub-projectile hit time, a negative reward needs to be added to the training process, so that the negative reward between [-1, 0) can be obtained through the calculation formula of R1.
[0191] In order to accelerate the convergence and ensure the convergence accuracy, a positive reward related to the distance between the sub-projectile and the target can be added to avoid the local optimal solution of rapid ending of training caused by the accumulation of time negative reward. The guide reward designed by the distance between the sub-projectile and the target is as follows:
[0192]
[0193] wherein R2 is the guide reward, e is the natural logarithm, and R FT is the relative distance between the sub-projectile and the target; k1 is the proportional coefficient, which can be set by the staff according to the situation.
[0194] Through the calculation formula of R2, the positive reward between [0, 1) can be obtained.
[0195] Therefore, the reward value before hitting is calculated as:
[0196]
[0197] wherein, The role of the term is to constrain the amplitude of the output in the training process, so as to avoid the emergence of too large or saturated instructions.
[0198] S5, determining the state parameter of the sub-projectile at the next time based on the state parameter and the maneuvering state.
[0199] For the embodiment of the present application, the maneuvering state of the sub-projectile is simulated through the training environment, and the state parameter at the next time is obtained.
[0200] S6, generating a training sample based on the state parameter, the maneuvering state, the first reward value, and the state parameter at the next time.
[0201] For the embodiment of the present application, since the Markov decision model <s, a, r, s'> is established for the sub-munition, a training sample can be obtained according to the obtained state parameter, the simulated maneuvering state, the first reward value corresponding to the state parameter and the state parameter at the next moment.
[0202] S7, the following loop steps are executed until a third preset condition is met:
[0203] The current maneuvering state of the sub-munition is simulated based on the current state parameter, which is the state parameter of the sub-munition at the next moment determined based on the state parameter and the maneuvering state in the last simulation cycle;
[0204] If the target is not hit, the current first reward value is determined based on the current state parameter;
[0205] The state parameter of the sub-munition at the next moment is determined based on the current state parameter and the current maneuvering state;
[0206] The training sample is generated based on the current state parameter, the current maneuvering state, the current first reward value and the state parameter at the next moment.
[0207] The third preset condition includes any one of the following:
[0208] The sub-munition hits the target;
[0209] The target escapes;
[0210] The number of times of determining the current first reward value reaches a preset number threshold.
[0211] For the embodiment of the present application, the maneuvering state of the sub-munition is simulated in the training environment, so that the state parameter of the sub-munition after maneuvering, i.e. the current state parameter, can be obtained, and then the training sample can be generated according to the manner in steps S2 to S6. When the sub-munition hits the target in the simulation process in the training environment, the simulation ends and the training sample does not need to be generated; or the target escapes in the simulation process, and if the training sample is still generated after the target escapes, the training sample after the escape does not have reference significance; assuming that the preset number threshold is 1000 times, the number of times of determining the current first reward value reaches 1000 times, i.e. 1000 training samples are generated, and the preset number threshold is a number representing that the training samples are sufficient, i.e. sufficient training samples are generated.
[0212] S8, if the sub-munition hits the target, the first hit time of the main munition and the second hit time of the sub-munition are determined.
[0213] For the embodiment of the present application, when simulating in the training environment, if the projectile hits the target, the time when the projectile hits the target is recorded, and after the projectile hits the target, the main projectile continues to simulate in the training environment, and if the main projectile hits the target, the time when the main projectile hits the target is recorded.
[0214] S9, calculating a second reward value based on the first hitting time of the main projectile and the second hitting time of the projectile.
[0215] For the embodiment of the present application, the ideal state is that the first hitting time is the same as the second hitting time, that is, the projectile and the main projectile hit the target at the same time, but during the simulation process, it is not the ideal state, so it is necessary to calculate the second reward value according to the two hitting times, and the second reward value can be calculated through the following formula:
[0216] R3 = -k2(t Lend -t Fend ) 2
[0217] Wherein, R3 is the reward value when the projectile and the main projectile hit the target; k2 is a proportional coefficient, which can be set by the staff according to the situation; t Lend is the first hitting time of the main projectile; t Fend is the second hitting time of the projectile.
[0218] Therefore, the reward function for training can be obtained by combining R1, R2 and R3 as follows:
[0219]
[0220] S10, generating a training sample based on the current state parameter, the current maneuvering state, the second reward value and the state parameter at the next moment.
[0221] For the embodiment of the present application, the training sample of the projectile when hitting the target can be obtained by obtaining the state function when hitting the target, the maneuvering state when hitting the target, the second reward value when hitting the target, and the state parameter at the next moment after hitting the target.
[0222] Specifically, since the projectile has hit the target, the state parameters related to the projectile itself in the state parameter at the next moment remain unchanged, and since the main projectile has not hit the target, the related state parameters of the main projectile itself, the related parameters of the main projectile to the target and the related parameters of the main projectile to the projectile are still calculated, so as to obtain the state parameter at the next moment.
[0223] In a possible implementation of the embodiment of the application, the state parameters include a lead angle of the main missile, a lead angle of the follower missile, a relative distance between the main missile and the target, a relative distance between the follower missile and the target, a speed of the main missile, and a speed of the follower missile. The determination of the first reward value based on the state parameters in step S4 includes steps S41 (not shown in the figure), S42 (not shown in the figure), and S43 (not shown in the figure), in which
[0224] In step S41, a first hit time of the main missile hitting the target is predicted based on the lead angle of the main missile, the relative distance between the main missile and the target, and the speed of the main missile. A second hit time of the follower missile hitting the target is predicted based on the lead angle of the follower missile, the relative distance between the follower missile and the target, and the speed of the follower missile.
[0225] For the embodiment of the application, the time at which the main missile or the follower missile hits the target at each moment can be predicted according to the lead angle, the relative distance to the target, and the time of hitting the target. In the embodiment of the application, the hit time can be predicted by the following formula:
[0226]
[0227] where t is the time at which the main missile or the follower missile hits the target, r is the relative distance between the main missile or the follower missile and the target, V is the relative speed of the main missile or the follower missile to the target, η is the lead angle of the main missile or the follower missile to the target, and N is a proportional guidance coefficient. go
[0228] In step S42, a time difference between the first hit time and the second hit time is determined.
[0229] For the embodiment of the application, the time difference can be obtained by subtracting the first hit time from the second hit time after the first hit time and the second hit time of the main missile and the follower missile are calculated by the formula in step S41. Since the hit times of the main missile and the follower missile have a time difference, in order to minimize the time difference and make it closest to the ideal situation, the reward value needs to be determined according to the time difference.
[0230] In step S43, the first reward value is determined based on the time difference.
[0231] Specifically, the first reward value can be calculated by the calculation formula of R1 in step S4, and details are not described herein.
[0232] The above embodiment introduces a main-follower missile cooperative guidance method and a training method of the follower missile without a seeker from the perspective of a method flow. The following embodiment introduces a main-follower missile cooperative guidance device and a training device of the follower missile without a seeker from the perspective of a virtual module or a virtual unit. Details are shown in the following embodiment.
[0233] The embodiment of the application provides a master-slave missile cooperative guidance device of a slave missile without a seeker, and the master-slave missile cooperative guidance device of the slave missile without the seeker specifically can include:
[0234] A parameter acquisition module is configured to acquire a real combat state parameter, and the real combat state parameter includes a related parameter between the master missile and a target, a related parameter between the master missile and the slave missile, a related parameter of the master missile itself, and a related parameter of the slave missile itself;
[0235] A calculation module is configured to input the real combat state parameter into a trained online actor network for calculation to obtain a maneuvering state of the slave missile, and the trained online actor network is obtained by updating and training an online network and a target network according to an acquired training sample;
[0236] A control module is configured to control the slave missile to run according to the maneuvering state.
[0237] In a possible implementation of the embodiment of the application, the calculation module inputs the real combat state parameter into the trained online actor network for calculation to obtain the maneuvering state of the slave missile, and the control device, when controlling the slave missile to run according to the maneuvering state, is specifically configured to:
[0238] input the real combat state parameter into the trained online actor network for calculation to obtain the maneuvering state of the slave missile;
[0239] control the slave missile to run according to the maneuvering state;
[0240] If the slave missile does not hit the target and the target does not escape, the following loop steps are executed until a first preset condition is met:
[0241] acquire a current real combat state parameter;
[0242] input the current real combat state parameter into the online actor network for state calculation to obtain a current maneuvering state of the slave missile;
[0243] control the slave missile to run according to the current maneuvering state;
[0244] The first preset condition includes any one of the following:
[0245] the target escapes;
[0246] the slave missile hits the target.
[0247] The embodiment of the application provides a master-slave missile cooperative guidance training device of a slave missile without a seeker, and the master-slave missile cooperative guidance training device of the slave missile without the seeker specifically can include:
[0248] The sample obtaining module is configured to obtain a plurality of training samples, the training samples comprising state parameters, a maneuvering state of the slave projectile, a reward value, and state parameters at a next time point;
[0249] The first evaluation module is configured to input the state parameters into the online network to evaluate a return to obtain a first evaluation value;
[0250] The second evaluation module is configured to input the state parameters at the next time point into the target network to evaluate the return to obtain a second evaluation value; and the updating module is configured to update the online network and the target network based on the first evaluation value, the second evaluation value, and the reward value to obtain a trained online network and a trained target network, the trained online network comprising a trained online actor network, the trained online actor network being configured to calculate the maneuvering state of the slave projectile according to the state parameters obtained from the slave projectile.
[0251] In a possible implementation of the embodiments of the present application, the state parameters comprise relevant parameters between the main projectile and the target, relevant parameters between the main projectile and the slave projectile, relevant parameters of the main projectile itself, relevant parameters of the slave projectile itself, and relevant parameters of the target; when the first evaluation module inputs the training samples into the online network to evaluate the return to obtain the first evaluation value, and the second evaluation module inputs the training samples into the target network to evaluate the return to obtain the second evaluation value, the first evaluation module and the second evaluation module are specifically configured to:
[0252] input the state parameters into the online actor network to calculate a target maneuvering state of the slave projectile;
[0253] input the state parameters and the target maneuvering state into the online critic network to evaluate the return to obtain the first evaluation value;
[0254] input the state parameters at the next time point into the target actor network to calculate a target maneuvering state of the slave projectile at the next time point; and input the state parameters at the next time point and the target maneuvering state at the next time point into the target critic network to evaluate the return to obtain the second evaluation value.
[0255] In a possible implementation of the embodiments of the present application, when the updating module updates the online network based on the first evaluation value, the second evaluation value, and the reward value, the updating module is specifically configured to:
[0256] update the online critic network based on the first evaluation value, the second evaluation value, and the reward value.
[0257] update the online actor network based on the updated online critic network.
[0258] perform the following loop steps until a second preset condition is reached:
[0259] inputting the state parameter in the next training sample into the current online network to evaluate the return to obtain a current first evaluation value;
[0260] inputting the state parameter in the next training sample into the current target network to evaluate the return to obtain a current second evaluation value;
[0261] updating the current online critic network based on the current first evaluation value, the current second evaluation value and the reward value in the next training sample;
[0262] updating the current online actor network based on the current online critic network;
[0263] The current online critic network is the online critic network after the last update, and the current online actor network is the online actor network after the last update.
[0264] The second preset condition includes at least one of the following:
[0265] The next training sample is the last training sample.
[0266] The time difference between the hits is the smallest, and the time difference between the hits is the difference between the main projectile hit time and the sub-projectile hit time.
[0267] The updating device is also used for:
[0268] When the number of updates reaches a specified number of times, the target network is updated based on the online network at the specified number of times.
[0269] In one possible implementation of the embodiments of the present application, the reward value includes a first reward value and a second reward value, and the device further includes:
[0270] The parameter acquisition module is configured to acquire the state parameter.
[0271] The simulation module is configured to simulate the state parameter to obtain the maneuvering state of the sub-projectile.
[0272] The determination module is configured to determine whether the sub-projectile hits the target based on the maneuvering state.
[0273] The determination module is configured to determine the first reward value based on the state parameter when the target is not hit.
[0274] The parameter determination module is configured to determine the state parameter of the sub-projectile at the next time based on the state parameter and the maneuvering state.
[0275] The first sample generation module is configured to generate a training sample based on the state parameter, the maneuvering state, the first reward value and the state parameter at the next time.
[0276] A cycle module is configured to perform the following cycle steps until a preset condition is met:
[0277] A current maneuvering state of the slave missile is obtained based on the simulation of the current state parameters, the current state parameters being state parameters of the slave missile at the next time point determined based on the state parameters in the last simulation cycle and the maneuvering state;
[0278] If the target is not hit, a current first reward value is determined based on the current state parameters;
[0279] State parameters of the slave missile at the next time point are determined based on the current state parameters and the current maneuvering state;
[0280] A training sample is generated based on the current state parameters, the current maneuvering state, the current first reward value and the state parameters at the next time point; the preset condition includes any one of the following:
[0281] The slave missile hits the target;
[0282] The target escapes;
[0283] The number of times of determining the current first reward value reaches a preset number threshold;
[0284] A time determination module is configured to determine a first hitting time of the primary missile and a second hitting time of the slave missile when the slave missile hits the target; a reward value calculation module is configured to calculate a second reward value based on the first hitting time of the primary missile and the second hitting time of the slave missile; and a second sample generation module is configured to generate a training sample based on the current state parameters, the current maneuvering state, the second reward value and the state parameters at the next time point.
[0285] In a possible implementation of the embodiment, the state parameters include a primary missile lead angle, a slave missile lead angle, a relative distance between the primary missile and the target, a relative distance between the slave missile and the target, a primary missile speed and a slave missile speed, and the determination module is specifically configured to:
[0286] predict the first hitting time of the primary missile hitting the target based on the primary missile lead angle, the relative distance between the primary missile and the target and the primary missile speed, predict the second hitting time of the slave missile hitting the target based on the slave missile lead angle, the relative distance between the slave missile and the target and the slave missile speed, and determine a time difference between the first hitting time and the second hitting time;
[0287] determine the first reward value based on the time difference.
[0288] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described primary-slave missile cooperative guidance device without a seeker of the slave missile and the primary-slave missile cooperative guidance training device without a seeker of the slave missile can refer to the corresponding process in the foregoing method embodiments, which will not be described herein.
[0289] This application provides a master-slave cooperative guidance system for a slave missile without a seeker, where the guidance is executed by the slave missile. Figure 6 As shown, the system includes:
[0290] Main missile 11 collects relevant parameters of main missile 11 and relevant parameters between main missile 11 and target 2, and sends the relevant parameters of main missile 11 and relevant parameters between main missile 11 and target 2 to secondary missile 12;
[0291] The system obtains relevant parameters of the main missile 11 and the relevant parameters between the main missile 11 and the target 2 from missile 12. The relevant parameters between the main missile and the target 2, the relevant parameters between the main missile and the target 2, the relevant parameters between the main missile and the target 12, the relevant parameters of the main missile 11 itself, and the relevant parameters of the target 12 itself are input into a trained online actor network for calculation to obtain the maneuvering state of the target missile 12. The system determines the maneuvering state of the target missile 12 based on the state parameters and controls the target missile 12 to operate according to the maneuvering state.
[0292] In this embodiment, the primary missile 11 and the secondary missile 12 communicate wirelessly during flight. The primary missile 11 collects its own parameters and the parameters between itself and the target 2 using its seeker and sensors, and transmits these parameters to the secondary missile 12 via wireless communication. The secondary missile 12 is equipped with its own sensors to collect its own parameters. Therefore, after obtaining the primary missile 11's parameters and the parameters between itself and the target 2, the secondary missile 12 combines these parameters with its own parameters to form state parameters. The secondary missile 12 has a trained actor network internally, so it can calculate its maneuvering state based on the state parameters. In this embodiment, the maneuvering state is the acceleration 'a' of the secondary missile 12. F This enables the missile to be guided even when it is without a seeker or when the seeker is interfered with or damaged during flight.
[0293] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the bullet 12 described above can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here.
[0294] This application provides an electronic device, such as... Figure 7 As shown, Figure 7 The illustrated electronic device 30 includes a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, for example, via a bus 302. Optionally, the electronic device 30 may also include a transceiver 304. It should be noted that in practical applications, the transceiver 304 is not limited to one type, and the structure of this electronic device 30 does not constitute a limitation on the embodiments of this application.
[0295] The processor 301 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in connection with the disclosure. The processor 301 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0296] The bus 302 can include a path for transmitting information between the above-mentioned components. The bus 302 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 302 can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 7 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or one type of bus.
[0297] The memory 303 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto.
[0298] The memory 303 is configured to store application program codes for implementing the solutions of the present application, and the processor 301 is configured to control the execution of the application program codes stored in the memory 303. The processor 301 is configured to execute the application program codes stored in the memory 303 to implement the content shown in the foregoing method embodiments.
[0299] The electronic device includes, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. It can also be a server or the like. Figure 7 The electronic device shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0300] The computer readable storage medium provided in the embodiments of the present application stores a computer program, and when the computer program runs on a computer, the computer can execute the corresponding content in the foregoing method embodiments. Compared with the related art, after obtaining the training sample, the state parameters in the training sample are input into the online network to evaluate the revenue, so as to obtain the first evaluation value of the online network. The state parameters at the next moment in the training sample are input into the target network to evaluate the revenue, so as to obtain the second evaluation value of the target network. The first evaluation value is used to represent the training situation of the online network, and the second evaluation value is used to represent the training situation of the target network. The online network and the target network are updated and corrected in combination with the reward value in the training sample, so as to obtain the trained online network and the trained target network. The trained online network includes a trained online actor network, which is used to calculate the maneuvering state of the sub-missile according to the state parameters obtained from the sub-missile, so that the sub-missile can be accurately guided without installing a seeker or when the seeker is damaged.
[0301] It should be understood that, although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other sequences. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.
[0302] The above merely describes some embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. A method for cooperative guidance of a master missile and a slave missile from a homing head, characterized in that, The method comprises the following steps: acquiring battle state parameters, the battle state parameters comprising relevant parameters between a main projectile and a target, relevant parameters between the main projectile and a sub-projectile, relevant parameters of the main projectile itself, relevant parameters of the sub-projectile itself, and a target speed, the relevant parameters between the main projectile and the target comprising a relative speed of the main projectile to the target and a line-of-sight angle of the main projectile to the target, the relevant parameters between the main projectile and the sub-projectile comprising a distance between the main projectile and the sub-projectile and a line-of-sight angle between the main projectile and the sub-projectile, the relevant parameters of the main projectile itself comprising a heading angle of the main projectile and a speed of the main projectile, and the relevant parameters of the sub-projectile itself comprising a heading angle of the sub-projectile and a speed of the sub-projectile; inputting the battle state parameters into a trained online actor network for calculation to obtain a maneuvering state of the sub-projectile, the trained online actor network being obtained by updating and training an online network and a target network according to acquired training samples; controlling the sub-projectile to operate according to the maneuvering state.
2. The method according to claim 1, wherein, The method comprises the following steps: inputting the battle state parameters into a trained online actor network for calculation to obtain a maneuvering state of the sub-projectile; controlling the sub-projectile to operate according to the maneuvering state, which comprises the following steps: inputting the battle state parameters into a trained online actor network for calculation to obtain a maneuvering state of the sub-projectile; controlling the sub-projectile to operate according to the maneuvering state; if the sub-projectile does not hit the target and the target does not escape, performing the following loop steps until a first preset condition is met: acquiring current battle state parameters; inputting the current battle state parameters into an online actor network for state calculation to obtain a current maneuvering state of the sub-projectile; controlling the sub-projectile to operate according to the current maneuvering state; the first preset condition comprising any one of the following: the target escapes; 3. A method for training cooperative guidance of a master missile and a slave missile from a homing head, characterized in that, the sub-projectile hits the target. The method comprises the following steps: acquiring a plurality of training samples, the training samples comprising state parameters, a maneuvering state of a sub-projectile, a reward value, and next-time state parameters, the state parameters comprising relevant parameters between a main projectile and a target, relevant parameters between the main projectile and a sub-projectile, relevant parameters of the main projectile itself, relevant parameters of the sub-projectile itself, and target parameters; inputting the state parameters into an online network to evaluate a benefit to obtain a first evaluation value; inputting the next-time state parameters into a target network to evaluate a benefit to obtain a second evaluation value; updating the online network and the target network based on the first evaluation value, the second evaluation value, and the reward value to obtain a trained online network and a trained target network, the trained online network comprising a trained online actor network, the trained online actor network being used to calculate a maneuvering state of the sub-projectile according to state parameters acquired by the sub-projectile; the reward value comprising a first reward value and a second reward value, and the method further comprises the following steps: acquiring state parameters; simulating the state parameters to obtain a maneuvering state of the sub-projectile; judging whether the sub-projectile hits a target based on the maneuvering state; if the sub-projectile does not hit the target, determining a first reward value based on the state parameters; determining next-time state parameters of the sub-projectile based on the state parameters and the maneuvering state; generate a training sample based on the state parameter, the maneuver state, the first reward value, and the state parameter of the next time; perform the following loop steps until a third preset condition is met: simulate to obtain the current maneuver state of the slave projectile based on the current state parameter, the state parameter of the next time being determined based on the state parameter in the last simulation cycle and the maneuver state; if the target is not hit, determine a current first reward value based on the current state parameter; determine the state parameter of the next time of the slave projectile based on the current state parameter and the current maneuver state; generate a training sample based on the current state parameter, the current maneuver state, the current first reward value, and the state parameter of the next time; the third preset condition includes any of the following: the slave projectile hits the target; the target escapes; the number of times of determining the current first reward value reaches a preset number threshold; if the slave projectile hits the target, determine a first hit time of the main projectile and a second hit time of the slave projectile; calculate a second reward value based on the first hit time of the main projectile and the second hit time of the slave projectile; generate a training sample based on the current state parameter, the current maneuver state, the second reward value, and the state parameter of the next time.
4. The method of claim 3, wherein, input the state parameter into an online network to evaluate a reward to obtain a first evaluation value; input the state parameter of the next time into a target network to evaluate a reward to obtain a second evaluation value, including: input the state parameter into an online actor network to calculate a target maneuver state of the slave projectile; input the state parameter and the target maneuver state into an online critic network to evaluate a reward to obtain the first evaluation value; input the state parameter of the next time into a target actor network to calculate a target maneuver state of the slave projectile of the next time; input the state parameter of the next time and the target maneuver state of the next time into a target critic network to evaluate a reward to obtain the second evaluation value.
5. The method of claim 4, wherein the master-slave missile cooperative guidance training method is characterized in that, update the online network based on the first evaluation value, the second evaluation value, and the reward value, including: update the online critic network based on the first evaluation value, the second evaluation value, and the reward value; update the online actor network based on the updated online critic network; perform the following loop steps until a second preset condition is met: input the state parameter in the next training sample into the current online network to evaluate a reward to obtain a current first evaluation value; input the state parameter of the next time in the next training sample into the current target network to evaluate a reward to obtain a current second evaluation value; update the current online critic network based on the current first evaluation value, the current second evaluation value, and the reward value in the next training sample; update the current online actor network based on the current online critic network. The current online critic network is an online critic network updated last time, and the current online actor network is an online actor network updated last time. The second preset condition includes at least one of the following: The next training sample is the last training sample. The time difference between the first hitting time of the main missile and the second hitting time of the slave missile is the smallest.
6. The training method of claim 5, further comprising: When the number of updates reaches a specified number, updating the target network based on the online network at the specified number.
7. The method of claim 3, wherein the master-slave missile cooperative guidance training method is characterized in that, The state parameters include the lead angle of the main missile, the lead angle of the slave missile, the relative distance between the main missile and the target, the relative distance between the slave missile and the target, the speed of the main missile, and the speed of the slave missile. The first hitting time of the main missile hitting the target is predicted based on the lead angle of the main missile, the relative distance between the main missile and the target, and the speed of the main missile. The time difference between the first hitting time and the second hitting time is determined. The first reward value is determined based on the time difference.
8. A master-slave missile cooperative guidance system of a slave missile without a homing head, used to perform the master-slave missile cooperative guidance method of any one of claims 1-2, characterized in that, It includes: The main missile collects the main missile related parameters and the related parameters between the main missile and the target, and sends the main missile related parameters and the related parameters between the main missile and the target to the slave missile. The slave missile obtains the main missile related parameters and the related parameters between the main missile and the target, inputs the related parameters between the main missile and the target, the related parameters between the main missile and the slave missile, the main missile itself related parameters and the slave missile itself related parameters into the trained online actor network for calculation to obtain the maneuvering state of the slave missile. The maneuvering state of the slave missile is determined based on the state parameters.
9. An electronic device, comprising: The slave missile is controlled to run according to the maneuvering state. It includes: At least one processor; Memory; At least one application program, wherein the at least one application program is stored in the memory and is configured to be executed by the at least one processor, and the at least one application program is used to execute the training method of claim 3-7.
Citation Information
Patent Citations
Cooperative guidance method applied to guided weapon without seeker
CN109631687A
Multi-missile cooperative attack guidance law design method based on reinforcement learning
CN112799429A