An intelligent collaborative trajectory online planning method for multiple missiles

By generating a lateral range domain motion model and Markov decision process for a multi-missile cooperative system and combining it with reinforcement learning, the complexity of cooperative trajectory optimization in the coordinated interception of multiple interceptor missiles is solved, and efficient and accurate online trajectory planning and target coverage are achieved.

CN119916817BActive Publication Date: 2025-10-03AIR FORCE UNIV PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411992480.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-03
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

During the interception of high-altitude, high-speed, and highly maneuverable targets, a single interceptor missile is difficult to fully cover the high-probability hit prediction range. The coordinated interception of multiple interceptor missiles has the complexity of coordinated ballistic optimization and the performance indicators cannot meet the requirements of online optimization.

Method used

Generate a time domain motion model of the multi-missile cooperative system and convert it into a lateral range domain motion model. Combined with Markov decision process and reinforcement learning, through convexification and discretization processing, an intelligent cooperative trajectory planning model is generated, and the cooperative trajectory of multiple interceptor missiles is optimized online.

Benefits of technology

The computational efficiency and accuracy of the initial solution of the coordinated trajectory are improved, and online optimization before the mid-guidance stage is achieved, ensuring complete coverage and effective interception of the target by multiple interceptor missiles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119916817B_ABST
    Figure CN119916817B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for online planning of intelligent collaborative trajectories of multiple missiles, which relates to the technical field of multiple missile interception. The method includes: generating a time domain motion model of multiple interceptor missiles in a multi-missile collaborative system; converting the time domain motion model into a lateral distance domain motion model, generating constraints and objective functions for collaborative trajectory planning; generating a Markov decision process based on the online planning of intelligent collaborative trajectories of multiple missiles based on the lateral distance domain motion model, constraints and objective functions, and generating an intelligent collaborative trajectory planning model of multiple missiles through reinforcement learning training; using the intelligent collaborative trajectory planning model of multiple missiles as the initial solution for online generation of multi-missile collaborative trajectories, and optimizing online the collaborative trajectories of multiple interceptor missiles based on the convex and discretized lateral distance domain motion model, constraints and objective functions, interceptor missile parameters, and maneuvering target parameters before the mid-range guidance phase. Based on this technical solution, the performance of online optimization of collaborative trajectories can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of multi-missile interception, and in particular to an online trajectory planning method for intelligent coordinated multi-missiles. Background Art

[0002] During the interception of high-altitude, high-speed, and highly maneuverable targets, the defender needs to predict the High Probability Hit Prediction Range (H2PR) based on the target information provided by the early warning detection system to ensure that the seeker can intercept the target as soon as it is turned on during the handover between mid-range and terminal guidance. In other words, the seeker's detection field of view completely covers the H2PR.

[0003] Due to the highly maneuverable characteristics of the target, it is very difficult for the defender to fully cover and effectively intercept the target H2PR by relying on a single interceptor missile. Therefore, multiple interceptor missiles are needed to coordinately intercept the target.

[0004] Currently, the coordinated interception of targets by multiple interceptor missiles mainly refers to multi-autonomous vehicle systems, multi-machine systems, and multi-robot systems. However, the complex constraints and performance indicators of collaborative trajectory optimization based on related technologies cannot meet the needs of collaborative trajectory online optimization.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0006] In order to overcome the problems existing in the related art, the embodiments of the present disclosure provide an intelligent collaborative trajectory online planning method for multiple missiles, which can improve the performance of collaborative trajectory online optimization.

[0007] According to a first aspect of an embodiment of the present disclosure, a method for online intelligent collaborative trajectory planning of multiple missiles is provided, the method comprising: generating a time domain motion model of multiple interceptor missiles in a multi-missile collaborative system; converting the time domain motion model into a lateral range domain motion model, generating constraints and an objective function for collaborative trajectory planning, the objective function indicating the terminal distance error reached by the multiple interceptor missiles; generating a Markov decision process for online intelligent collaborative trajectory planning of multiple missiles based on the lateral range domain motion model, the constraints and the objective function; and generating an intelligent collaborative trajectory planning model of multiple missiles through reinforcement learning training based on the Markov decision process; convexifying and discretizing the lateral range domain motion model, the constraints and the objective function; using the intelligent collaborative trajectory planning model of multiple missiles as the initial solution for online generation of the collaborative trajectory of multiple missiles, and optimizing online before the mid-range guidance phase the collaborative trajectory of multiple interceptor missiles based on the convexified and discretized lateral range domain motion model, the constraints and the objective function, the interceptor missile parameters, and the maneuvering target parameters.

[0008] According to a second aspect of an embodiment of the present disclosure, a device for online intelligent collaborative trajectory planning of multiple missiles is provided, and the device comprises: a time domain motion model generation module, a distance domain motion model conversion module, an MDP generation module, an intelligent collaborative trajectory training module, a convexification module, a discretization module and an online optimization collaborative trajectory generation module; a time domain motion model generation module for generating a time domain motion model of multiple interceptor missiles in a multi-missile collaborative system; a distance domain motion model conversion module for converting the time domain motion model into a lateral distance domain motion model, generating constraints and an objective function for collaborative trajectory planning, wherein the objective function indicates the terminal distance error reached by the multiple interceptor missiles; an MDP generation module for generating a time domain motion model of multiple interceptor missiles based on the lateral distance domain motion model, the constraints and the objective function; The system is composed of a standard function, which generates a Markov decision process for online planning of intelligent collaborative trajectories of multiple missiles; an intelligent collaborative trajectory training module, which is used to generate an intelligent collaborative trajectory planning model of multiple missiles through reinforcement learning training based on the Markov decision process; a convexification module, which is used to convexify the lateral distance domain motion model, constraints and objective function; a discretization module, which is used to discretize the lateral distance domain motion model, constraints and objective function; an online optimized collaborative trajectory generation module, which is used to use the intelligent collaborative trajectory planning model of multiple missiles as the initial solution for online generation of multi-missile collaborative trajectories, and based on the convexed and discretized lateral distance domain motion model, constraints and objective function, interceptor missile parameters, and maneuvering target parameters, online optimizes and generates collaborative trajectories of multiple interceptor missiles before the mid-range guidance phase.

[0009] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for online intelligent collaborative trajectory planning of multiple missiles as described in the first aspect is implemented.

[0010] According to a fourth aspect of an embodiment of the present disclosure, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the computer-readable instructions, when executed by the processor, implement the method for online intelligent collaborative trajectory planning of multiple missiles as described in the first aspect.

[0011] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0012] In the disclosed embodiment, first, a time domain motion model of multiple interceptor missiles in a multi-missile cooperative system is generated, the time domain motion model of the multiple interceptor missiles is converted into a lateral range domain motion model, and constraints and objective functions for cooperative trajectory planning are generated; then, a Markov decision process for online planning of intelligent cooperative trajectories of multiple missiles is generated based on the lateral range domain motion model, constraints and objective functions, and an intelligent cooperative trajectory planning model of multiple missiles is generated through reinforcement learning training; further, the lateral range domain motion model, constraints and objective functions are convexified and discretized, and then the generated intelligent cooperative trajectory planning model of multiple missiles is used as the initial solution for online generation of cooperative trajectories of multiple missiles. Based on the convexified and discretized lateral range domain motion model, constraints and objective functions, interceptor missile parameters, and maneuvering target parameters, the cooperative trajectories of multiple interceptor missiles are generated through online optimization before the mid-range guidance stage. Since the time domain motion model is converted into a distance domain motion model, the MDP of collaborative trajectory planning can be accurately generated based on the distance domain motion model, constraints and objective function. Through training through reinforcement learning training strategy, an intelligent collaborative trajectory planning model for multiple missiles can be quickly and accurately generated. Compared with related technologies, the computational efficiency and accuracy of the initial solution of the collaborative trajectory can be improved. Not only can the subsequent multi-missile collaborative trajectory convex optimization stage be quickly entered, but also based on the collaborative trajectory planning model, the collaborative trajectory can be efficiently and accurately generated online, realizing online trajectory planning of multiple missiles before the mid-range guidance stage, thereby realizing the online optimization generation of intelligent collaborative trajectories for multiple missiles in mid-range guidance, and providing a guarantee for multiple interceptor missiles to complete full coverage of the target H2PR and effectively intercept it.

[0013] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0015] Figure 1 A schematic diagram of the system architecture of a multi-missile coordination system provided in an embodiment of the present disclosure.

[0016] Figure 2 A flowchart of a method for online trajectory planning of multiple missiles using intelligent collaborative ballistics is provided in an embodiment of the present disclosure.

[0017] Figure 3 A schematic diagram of the detection field of view of a radar seeker provided in an embodiment of the present disclosure.

[0018] Figure 4 A schematic diagram of collaborative coverage of a detection field of view provided in an embodiment of the present disclosure.

[0019] Figure 5 A schematic diagram of a training strategy provided in an embodiment of the present disclosure.

[0020] Figure 6 This is one of the simulation result diagrams provided in the embodiment of the present disclosure.

[0021] Figure 7 This is the second schematic diagram of the simulation results provided by the embodiment of the present disclosure.

[0022] Figure 8 This is the third schematic diagram of the simulation results provided by the embodiment of the present disclosure.

[0023] Figure 9 This is the fourth schematic diagram of the simulation results provided by the embodiment of the present disclosure.

[0024] Figure 10 This is the fifth schematic diagram of the simulation results provided in the embodiment of the present disclosure.

[0025] Figure 11 This is a hardware structure diagram of the computer device where the intelligent collaborative trajectory online planning system for multiple missiles in the embodiment of the present disclosure is located.

[0026] Figure 12 A schematic diagram of the structure of an intelligent collaborative trajectory online planning device for multiple missiles provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0028] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the embodiments of the present disclosure refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0029] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0030] Next, the embodiments of the present disclosure are described in detail.

[0031] Figure 1 This is a schematic diagram of the system architecture of a multi-missile cooperative system provided by an embodiment of the present disclosure. Figure 1 As shown in the figure, three interceptors of the same model, Interceptor 1, Interceptor 2, and Interceptor 3, are considered as an integrated system for online coordinated trajectory planning. After the coordinated phase is completed, the seeker is released during the terminal intercept phase. The seeker's field of view for each of the three interceptors must cover the target HP2R, determined based on the predicted trajectory of the maneuvering target, to ensure successful impact within the predicted high-probability hit area. The target HP2R is a circular area with a radius of r, centered at the predicted impact point, in the hoz plane of the East-North-Up (ENU) coordinate system.

[0032] like Figure 2 As shown, Figure 2 A flowchart of a method for online trajectory planning of multiple missiles using intelligent collaborative trajectory planning according to an embodiment of the present disclosure includes the following steps S201 to S205:

[0033] S201. Generate time-domain motion models of multiple interceptor missiles in a multi-missile coordination system.

[0034] Optionally, embodiments of the present disclosure can be applied to scenarios where multiple interceptor missiles of the same or similar models coordinate to intercept a high-altitude, highly maneuvering target. Each of the multiple interceptor missiles includes a seeker. Specifically, during the mid-range guidance phase, after the multiple interceptor missiles' seekers are activated, the planned coordinated trajectory enables their seeker fields of view to fully cover the maneuvering target's HP2R.

[0035] Figure 3 A schematic diagram of the detection field of view of a radar seeker provided in an embodiment of the present disclosure. Taking the interceptor missile seeker as a platform-type radar seeker as an example, the maximum detection range of the seeker is R max , the detection field of view of the seeker is The maximum rotation angle is χ. Figure 3 As shown in (a), the seeker detection field of view indicates: a cone area with the interceptor missile seeker as the vertex; the cone half-vertex angle is The length of the cone busbar is R max , the cone height is R max The radius of the cone base is R max The detection field of view of the seeker can rotate around the vertex in space with a maximum of no more than x degrees. Figure 3 As shown in (b), full coverage of the target HP2R means that the combination of the illumination circles of the detection fields of view of the seekers of multiple interceptor missiles can completely cover the target HP2R of the maneuvering target.

[0036] In order to facilitate calculation, dimensionless processing is performed on the motion model in the time domain, dimensionless time domain parameters are used, and the values ​​are normalized.

[0037] For example, the time domain motion models of multiple interceptor missiles in the multi-missile coordinated system can be generated based on the position coordinates, velocity, ballistic inclination, ballistic deviation, angle of attack and roll angle of each interceptor missile in the northeast celestial coordinate system.

[0038] Specifically, the time domain motion model of the i-th interceptor missile in the multi-missile cooperative system can be generated based on the following formula (1).

[0039]

[0040] Among them, x i 、h i 、z i They represent the dimensionless northeast celestial coordinates of the center of mass of the i-th interceptor missile in the northeast celestial coordinate system, v i represents the dimensionless velocity of the i-th interceptor missile, θ i represents the ballistic inclination angle of the i-th interceptor missile, ψ vi represents the ballistic deviation angle of the i-th interceptor missile, α i represents the attack angle of the i-th interceptor missile, ζ i represents the roll angle of the i-th interceptor missile, D αi represents the dimensionless drag acceleration of the i-th interceptor missile, L αi represents the dimensionless lift acceleration of the i-th interceptor missile, g irepresents the dimensionless gravitational acceleration of the i-th interceptor missile, D αi =C Di q i S i / m i g0,L αi =C Li q i S i / m i g0, g i =(1 / (1+h i )) 2 , C Di represents the drag coefficient of the i-th interceptor missile, C Di =c d1 +c d2 α i +c d3 α i 2 , c d1 , c d2 , c d3 Represents the resistance parameter, C Li represents the lift coefficient of the i-th interceptor missile, C Li =c l1 +c l2 α i , c l1 , c l2 represents the lift parameter, q i represents the dynamic pressure of the i-th interceptor missile, ρ i represents the air density where the i-th interceptor missile is located, ρ0 represents the air density at sea level, ρ0 = 1.226 km / m 3 , H represents the reference height, H = 7254.24m, r e represents the radius of the Earth, r e =6371.2km, S i represents the characteristic area of ​​the i-th interceptor missile, m i represents the mass of the i-th interceptor missile, and g0=9.81 represents the gravitational acceleration at sea level.

[0041] S202: Convert the time domain motion model of multiple interceptor missiles into a lateral range domain motion model to generate constraint conditions and objective functions for collaborative trajectory planning.

[0042] Among them, the objective function indicates the terminal distance error reached by multiple interceptor missiles.

[0043] Optionally, the constraints of collaborative trajectory planning may include at least one of the following: process constraints, control quantity constraints, terminal constraints and trust region constraints; process constraints may include at least one of the following: heat flux density constraints, dynamic pressure constraints, overload constraints; control quantity constraints may include at least one of the following: attack angle constraints, roll angle constraints, and safety distance constraints between interceptor missiles; terminal constraints include at least one of the following: trajectory inclination terminal value constraints and trajectory deviation terminal value constraints.

[0044] S203. Based on the lateral range domain motion model, constraints, and objective function, a Markov decision process for online planning of intelligent collaborative trajectories of multiple missiles is generated; and based on the Markov decision process, an intelligent collaborative trajectory planning model for multiple missiles is generated through reinforcement learning training.

[0045] In this disclosed embodiment, the objective function is the difference between the actual terminal distance reached by the interceptor missile using the initial trajectory solution generated by reinforcement learning (i.e., the multi-missile intelligent collaborative trajectory planning model generated above) and the predicted ideal terminal distance. The performance of the model training can be judged by minimizing the terminal distance error as a performance indicator.

[0046] Exemplarily, the reinforcement learning training method adopted in the embodiments of the present disclosure may be a deep reinforcement learning training method, such as DDPG (Deep Deterministic Policy Gradient).

[0047] S204. Convex and discretize the lateral distance domain motion model, constraints and objective function.

[0048] Illustratively, in the embodiments of the present disclosure, convexification is achieved based on linearization processing.

[0049] S205. The intelligent collaborative trajectory planning model of multiple missiles is used as the initial solution for online generation of the collaborative trajectory of multiple missiles. Based on the convex and discretized lateral range domain motion model, constraints and objective function, interceptor missile parameters, and maneuvering target parameters, the collaborative trajectory of multiple interceptor missiles is generated through online optimization before the mid-range guidance phase.

[0050] It should be noted that the missile's flight process is divided into the initial guidance stage, the intermediate guidance stage and the terminal guidance stage.

[0051] In the disclosed embodiments, online trajectory planning is primarily performed prior to the mid-range guidance phase, primarily for the missile's mid-range guidance phase. This online planning indicates that coordinated trajectory planning for the mid-range guidance phase is performed online simultaneously during the missile's initial guidance phase. During the missile's actual mid-range guidance flight, it is sufficient to track the mid-range guidance trajectory planned online during the initial guidance phase. In other words, from the perspective of the missile's entire flight sequence, trajectory planning for the mid-range guidance phase is performed online.

[0052] The disclosed embodiments provide a method for online planning of intelligent collaborative trajectories of multiple missiles. First, a time-domain motion model of multiple interceptor missiles in a multi-missile collaborative system is generated, the time-domain motion model of the multiple interceptor missiles is converted into a lateral range domain motion model, and constraints and objective functions for collaborative trajectory planning are generated. Then, a Markov decision process for online planning of intelligent collaborative trajectories of multiple missiles is generated based on the lateral range domain motion model, constraints and objective functions, and an intelligent collaborative trajectory planning model of multiple missiles is generated through reinforcement learning training. Further, the lateral range domain motion model, constraints and objective functions are convexified and discretized, and then the generated intelligent collaborative trajectory planning model of multiple missiles is used as the initial solution for online generation of the collaborative trajectory of multiple missiles. Based on the convexified and discretized lateral range domain motion model, constraints and objective functions, interceptor missile parameters, and maneuvering target parameters, the collaborative trajectory of the multiple interceptor missiles is generated through online optimization before the mid-range guidance stage. Since the time domain motion model is converted into a distance domain motion model, the MDP of collaborative trajectory planning can be accurately generated based on the distance domain motion model, constraints and objective function. Through training through reinforcement learning training strategy, an intelligent collaborative trajectory planning model for multiple missiles can be quickly and accurately generated. Compared with related technologies, the computational efficiency and accuracy of the initial solution of the collaborative trajectory can be improved. Not only can the subsequent multi-missile collaborative trajectory convex optimization stage be quickly entered, but also based on the collaborative trajectory planning model, the collaborative trajectory can be efficiently and accurately generated online, realizing online trajectory planning of multiple missiles before the mid-range guidance stage, thereby realizing the online optimization generation of intelligent collaborative trajectories for multiple missiles in mid-range guidance, and providing a guarantee for multiple interceptor missiles to complete full coverage of the target H2PR and effectively intercept it.

[0053] Optionally, in the method for online intelligent collaborative trajectory planning of multiple missiles provided in an embodiment of the present disclosure, the above S205 may include the following S205a and S205b:

[0054] S205a. Determine the mid-range guidance terminal positions of the multiple interceptor missiles based on the target HP2R determined by the maneuvering target parameters and the number of the multiple interceptor missiles.

[0055] In the disclosed embodiment, combined with guidance principles, the aforementioned "coordinated coverage of the target HP2R by the detection fields of the seekers of multiple interceptor missiles" can be converted into a theory of "small circles covering a large circle." According to the optimal small circle coverage principle, if the radius of the small circle is determined, the optimal number and position of the small circles can be determined by the radius and position of the large circle. The optimal small circle coverage principle is: Ω represents a unit circle, and assuming there are w circles with a radius ratio R to Ω. Ω The same small circle Ω1, Ω2, ...Ω w , if Ω1, Ω2, …Ω k Can completely cover Ω, then Table 1 shows RΩ The corresponding relationship with the number of small circles w.

[0056] Table 1

[0057]

[0058] For example, Figure 4 This is a schematic diagram of a detection field of view cooperative coverage provided by an embodiment of the present disclosure. Taking the multi-missile cooperative system composed of three interceptor missiles to intercept a maneuvering target as an example, the detection field of view cooperative coverage of multiple seekers is as follows: Figure 4 As shown in the figure, the three interceptor missiles form a regular triangle and can evenly and completely cover the radius of The three interceptor missiles' detection fields of view intersect the target HP2R at three points: D1, D2, and D3. T is the predicted impact point, and M1, M2, and M3 are the center points of the three interceptor missiles' detection fields of view, located at the midpoints of D1D2, D1D3, and D2D3, respectively. Based on geometric deduction, we can determine: When the geometric situation of the three interceptor missiles and the target HP2R is determined, the ideal position of the guidance terminal of the three interceptor missiles can be determined based on the target preset impact point based on formula (1): z2=z T ,

[0059] S205b. Based on the convexed and discretized lateral range domain motion model, constraints and objective function, the learned intelligent collaborative trajectory planning model for multiple missiles, the parameters of multiple interceptor missiles in the multi-missile collaborative system, and the mid-range guidance terminal positions of multiple interceptor missiles, the collaborative trajectory of multiple interceptor missiles is optimized online before the mid-range guidance stage.

[0060] Specifically, based on the above-mentioned determination of the mid-range guidance terminal positions of multiple interceptor missiles, the initial solutions of the coordinated trajectories of the multiple interceptor missiles can be generated through offline training of reinforcement learning. Based on the initial solutions of the coordinated trajectories and combined with the coordinated trajectory convex optimization method (a related technology), the mid-range guidance coordinated trajectories of the multiple interceptor missiles can be generated so that the multiple interceptor missiles can reach the determined mid-range guidance terminal positions as much as possible.

[0061] Optionally, in the method for online trajectory planning of multiple missiles provided in the embodiment of the present disclosure, the above-mentioned S202 may be performed by the following S202a:

[0062] S202a: Based on a preset conversion relationship, convert the time domain motion model into a lateral range domain motion model.

[0063] It should be noted that, in the embodiments of the present disclosure, the lateral distance domain may also be referred to as the range domain.

[0064] Specifically, in the embodiment of the present disclosure, the interception strategy is set so that the lateral distance (x-axis) of each interceptor missile is consistent.

[0065] For example, based on the preset conversion relationship shown in formula (2), the time domain motion model of the i-th interceptor missile in the multi-missile cooperative system shown in formula (1) can be converted into the lateral range domain motion model of the i-th interceptor missile in the multi-missile cooperative system shown in formula (3).

[0066] dt / dx=1 / vcosθcosψ v Formula (2)

[0067]

[0068] It can be understood that by converting the time domain motion model into the distance domain motion model, calculation can be facilitated and accuracy of calculation can be improved.

[0069] Optionally, in the method for online intelligent collaborative trajectory planning of multiple missiles provided in the embodiment of the present disclosure, the above-mentioned S202 may include the following S202b:

[0070] S202b. Based on the motion parameters of each interceptor missile in the time domain motion model of the multiple interceptor missiles, generate the first state equation, the first constraint condition and the first objective function of the multi-missile coordination system in the lateral range domain.

[0071] The first state equation includes: the state equation of each interceptor missile, the control vector matrix of each interceptor missile, and the control vector of each interceptor missile.

[0072] The first constraint condition includes at least one of the above-mentioned constraints of collaborative trajectory planning.

[0073] Specifically, based on formula (4), the state equation of each interceptor missile in the multi-missile cooperative system, the control vector matrix of each interceptor missile, and the control vector of each interceptor missile, the first state equation of the multi-missile cooperative system is generated.

[0074]

[0075] Where s=(s1,…,s i ,…,s N ) T , N represents the number of interceptor missiles in the multi-missile coordination system, x f It represents the ideal terminal position of each interceptor missile at the end of the mid-guidance phase, s i represents the state vector of the i-th interceptor missile, s i =(h i ,z i ,v i ,θ i ,ψvi ,α i ,σ i ) T ,f(s,x)=(f1(s1,x),…,f i (s i ,x),…f N (s N ,x)) T , f i (s i ,x) represents the state equation of the i-th interceptor missile. B represents the system control vector matrix of the multi-missile coordination system, B=(B1,…,B i ,…,B N ) T , B i represents the control vector matrix of the i-th interceptor missile in the multi-missile coordinated system, u represents the system control vector matrix of the multi-missile cooperative system, u=(u1,…,u i ,…,u N ) T ,u i represents the control vector of the i-th interceptor missile, u i =(α i ,σ i ) T .

[0076] Exemplarily, the state equation of the i-th interceptor missile is generated based on formula (5) in combination with the lateral range domain motion model shown in formula (3).

[0077]

[0078] For example, taking the multi-missile coordinated system including three interceptor missiles as an example, in the first state equation, the above s=(s1,s2,s3) T , u=(u1,u2,u3) T , B=(B1,B2,B3) T ,

[0079] Generate constraints:

[0080] Example 1: Based on formula (6), the constraint functions of the heat flux density, dynamic pressure, and overload of the i-th interceptor missile are set to ensure long-term high-altitude and high-speed flight during the guidance phase.

[0081]

[0082] Among them, C Q represents the heat flux calculation constant, CQ =1.291×10 -4 , Q max represents the maximum value of heat flux constraint, q max represents the maximum value of dynamic pressure constraint, n max Indicates the maximum value of the overload constraint.

[0083] Example 2: Based on formula (7), set the attack angle constraint and roll angle constraint of the interceptor missile to ensure the flight stability of each interceptor missile.

[0084]

[0085] Among them, α max Indicates the maximum limit of the angle of attack, ζ max Indicates the maximum limit of the roll angle.

[0086] Example 3: Based on formula (8), set the constraint function of the safe distance between the i-th interceptor missile and the j-th interceptor missile.

[0087]

[0088] Among them, d min It is the safety distance threshold between each interceptor missile.

[0089] Example 4: Based on formula (12), constraints are set on the ballistic inclination angle and the terminal value of the ballistic deviation angle of the i-th interceptor missile to ensure the stability of the flight state of each interceptor missile.

[0090]

[0091] It should be noted that the introduction of the maximum rotation angle of the seeker can ensure that the detection field of the seeker can illuminate the target.

[0092] Generate the objective function:

[0093] Based on the system capture situation of each interceptor missile against the maneuvering target, that is, each interceptor missile tries to reach its own predicted handover point, the first objective function is set based on formula (10).

[0094]

[0095] Among them, c1, c2 are weight coefficients, h i (t f ), x i (t f ) and z i (t f ) represents the ideal position of the mid-range guidance terminal of the i-th interceptor missile, h if x if and z ifIndicates the actual arrival position of the mid-range guidance terminal of the i-th interceptor missile.

[0096] Based on the constraints and objective function, the model performance indicator is selected as the minimum distance error between the interceptor missile and the ideal terminal, which satisfies the coordinated capture of maneuvering targets by multiple interceptor missiles. The online optimization task of multi-missile coordinated interception is:

[0097] P0: min(10), subject to: (4)-(9), s(x=0)=[h i0 ,z i0 ,v i0 ,θ i0 ,ψ vi0 ,α i0 ,σ i0 ].

[0098] Specifically, based on the above approach, a collaborative trajectory online optimization model, constraint function, and objective function can be established in the distance domain.

[0099] Motion model convexification processing:

[0100] It should be noted that the disclosed embodiment convexifies the model using a linearized approach to facilitate subsequent convex optimization. The non-convex first equation of state and the non-convex second constraint in the first constraint are convexified. The second constraint includes at least one of the following: a heat flux constraint, a dynamic pressure constraint, an overload constraint, and a safety distance constraint between interceptor missiles.

[0101] Optionally, in the method for online intelligent collaborative trajectory planning of multiple missiles provided in the embodiment of the present disclosure, the above-mentioned S204 can be performed by the following S204a1 to S204a4:

[0102] S204a1, based on the reference trajectory, linearize the first state equation into a second state equation and linearize the second constraint condition. For example, the first state equation shown in the above formula (4) is linearized based on the reference trajectory s k , perform Taylor expansion and obtain the equation of state: make The second state equation shown in formula (11) can be obtained.

[0103]

[0104] Taking three interceptor missiles as an example, in the second state equation after linearization, for A i(i=1,2,3) The elements in the order from top to bottom are:

[0105] a31 =r e D αi / (Hv i cosθ i cosψ vi )+2tanθ i / ((1+h i ) 3 v i cosψ vi );

[0106]

[0107] a 14 =1 / (cos 2 θ i cosψ vi );

[0108] a 34 =-D αi tanθ i / (v i cosθ i cosψ vi )-1 / ((1+h i ) 2 v i cos 2 θ i cosψ vi );

[0109]

[0110] a 35 =-D αi tanψ vi / (v i cosθ i cosψ vi )-tanθ i tanψ vi / ((1+h i ) 2 v i cosψ vi );

[0111]

[0112]

[0113] Convexification of constraints:

[0114] For example, the above formula (6) is applied to the reference trajectory s k Performing Taylor expansion, we obtain the following formula (12).

[0115]

[0116] Among them, the terms in formula (12) are:

[0117] Based on the reference trajectory s k The safety distance constraint function between each interceptor missile shown in formula (8) is convexified to obtain the safety distance constraint function between the i-th interceptor missile and the j-th interceptor missile shown in formula (13).

[0118]

[0119] S204a2: Set a first trust region function.

[0120] The first trust region function is used to adjust the linearization error.

[0121] Specifically, in order to reduce the error caused by the linearization of the motion model, the first trust region function is set based on formula (14).

[0122] |ss k |≤δ Formula (14)

[0123] Where δ is the trust region constraint radius, which is the same as the dimension of s.

[0124] S204a3: Generate a second objective function based on the first objective function and the preset additional items.

[0125] The preset additional term is used to adjust the linearization error of the first objective function.

[0126] For example, by adding a preset additional risk to formula (10), the second objective function shown in formula (15) is obtained, so that the control amount is smoother and the oscillation amplitude is avoided to be too large.

[0127]

[0128] Among them, c3 is the weight coefficient.

[0129] Furthermore, the above-mentioned multi-missile coordinated interception online optimization task P0 can be transformed into a multi-missile coordinated interception online optimization convex task P1: min(15); subject to(7),(9)(11)-(14); s(x=0)=[h i0 ,z i0 ,v i0 ,θ i0 ,ψ vi0 ,α i0 ,ζ i0 ].

[0130] Based on this scheme, the objective function, inequality constraints, equality constraints (i.e., motion model), etc. can be converted into convex and linear. Since the above-mentioned dynamic equations and path constraints are highly nonlinear, the above-mentioned motion model can be linearized through the above-mentioned method to facilitate convex optimization processing.

[0131] Discretization processing:

[0132] It can be understood that task P1 is a convex optimization problem on a continuous distance domain. Both the state and control variables are continuous. It can be converted into a discrete mathematical programming problem for optimizing finite parameters, so that numerical methods can be used to quickly solve it.

[0133] Optionally, in the method for online trajectory planning of multiple missiles provided in the embodiment of the present disclosure, the above-mentioned S204 can be performed by the following S204b1 to S204b3:

[0134] S204b1. Discrete the second state equation, target constraints and second objective function.

[0135] The target constraint condition includes at least one of the third constraint conditions and / or at least one of the discretized second constraint conditions.

[0136] The third constraint condition includes: constraint on the angle of attack, constraint on the roll angle, constraint on the terminal value of the trajectory inclination angle, and constraint on the terminal value of the trajectory deviation angle.

[0137] For example, the Euler method can be used to discretize formula (11) to obtain the state equation shown in formula (16).

[0138]

[0139] Among them, x p represents the grid point position of the generated collaborative ballistic trajectory in the lateral range domain, p represents the position sequence of the grid point, p=1,...,l, x1=0≤…≤x p ≤x p+1 ≤…≤x l =x f , Δx p Represents the difference in horizontal distance position between grid point p and grid point p+1.

[0140] It can be understood that when the Euler method is used to discretize the second state equation, the Euler method can also be used to discretize the constraint conditions and the objective function.

[0141] S204b2. Generate a second trust region function based on the first trust region function and the trust region relaxation coefficient.

[0142] Among them, the trust region relaxation coefficient is used to adjust the error caused by discretization.

[0143] Exemplarily, since the trust region required for each iterative optimization is different, a trust region relaxation coefficient is used in the embodiment of the present disclosure to adjust the trust region function.

[0144]

[0145] Where λ is the trust region relaxation coefficient.

[0146] It should be noted that the use of a variable trust region processing method can ensure that the trust region required for each iterative optimization solution is different, thereby improving the computational efficiency of the algorithm optimization.

[0147] S204b3. Generate a third objective function based on the discretized second objective function and the trust region relaxation coefficient.

[0148] Furthermore, the second objective function of task P1 can be transformed into the third objective function shown in formula (18).

[0149]

[0150] Among them, c4 is the weight coefficient.

[0151] It should be noted that the J0 term is used to ensure that the terminal position constraint is satisfied. Expanding the trust region relaxation coefficient λ into the objective function can reduce the optimization space and improve the convergence efficiency. The actual part that needs to be optimized is Integral term, this processing method can increase the feasible solution space of the problem and avoid the situation where convergence cannot be achieved due to the inability to find a feasible solution.

[0152] Optionally, in the method for online intelligent collaborative trajectory planning of multiple missiles provided in the embodiment of the present disclosure, the above-mentioned S205 may be executed by the following S205c:

[0153] S205c. Before the mid-guidance phase, based on the discretized second state equation, determine the initial solution trajectory of the multi-missile coordination system that satisfies the target constraint condition, the second trust region function, and the third objective function.

[0154] In the embodiment of the present disclosure, the above non-convex problem 1 is first converted into a convex problem P1, and then converted into a discretized problem P2: min(18); subject to: (7), (9), (12), (16), (17); s(x=0)=[h0,z0,v0,θ0,ψ v0 ,α0,ζ0].

[0155] Generate the initial solution of the collaborative trajectory based on the above P2.

[0156] It should be noted that, in order to facilitate reinforcement learning, the embodiments of the present disclosure provide a method for generating an MDP (Markov Decision Process) for a collaborative trajectory planning task.

[0157] Optionally, before the above S204, the following S206 may be further included. Furthermore, the above S204 may be executed through the following S204c:

[0158] S204c, generating a state set of the multi-missile cooperative system, an action set of the multi-missile cooperative system, a reward function and a discount coefficient of the multi-missile cooperative system.

[0159] The state set of the multi-missile coordination system includes at least one of the following items of the interceptor missile: position coordinates, speed, ballistic inclination, ballistic deviation, angle of attack, roll angle, remaining distance, component of the velocity lead angle in the longitudinal plane and component of the transverse plane.

[0160] The action set of the multi-missile coordination system includes the interceptor missile's: angle of attack and roll angle.

[0161] The reward function of the multi-missile cooperative system includes at least one of the following: a step-length feedback reward based on the remaining distance, a step-length feedback reward based on the velocity lead angle, a step-length feedback reward based on the consistency of the remaining time, a node target reward based on a single interceptor missile, and a final target reward based on the multi-missile cooperative system.

[0162] It should be noted that, when the MDP is generated in the above S204, other relevant parameters of the MDP may also be generated, such as setting the discount coefficient to a preset value.

[0163] For example, in the disclosed embodiment, the collaborative trajectory planning task training of the multi-missile collaborative system can be performed based on the DDPG (Deep Deterministic Policy Gradient) algorithm in deep reinforcement learning to generate an initial collaborative trajectory solution.

[0164] Typically, an MDP includes: a set of states, a set of actions, a reward function, a discount factor, and a transition probability. The disclosed embodiment uses a model RL approach that does not consider transition probabilities.

[0165] The corresponding MDP of the multi-missile coordinated trajectory planning problem is generated in the following way.

[0166] Specifically, according to the guidance mechanism of the interceptor missile, the remaining distance s of the i-th interceptor missile can be generated based on formula (19): togoi .

[0167]

[0168] Based on the interceptor missile’s longitudinal plane sight angle, transverse plane sight angle and formula (20), the component η of the velocity lead angle of the i-th interceptor missile in the longitudinal plane is generated. xhi and the transverse plane component η xzi .

[0169]

[0170] in, represents the longitudinal plane sight angle of the i-th interceptor missile, is the line of sight angle of the lateral plane of the i-th interceptor missile,

[0171] Generate multi-missile coordinated system state set: [h i ,z i ,v i ,θ i ,ψ vi ,α i ,ζ i ,s togoi ,η xhi ,η xzi ].

[0172] The input for generating guidance control of multi-missile coordinated system is an action set:

[0173] Generate reward function:

[0174] The goal of trajectory planning in the coordinated guidance phase of a multi-missile coordinated system is to ensure that multiple interceptor missiles can reach the predetermined handover position, forming a favorable situation for target containment and interception, and ensuring stable target killing in the terminal guidance phase. The setting of the reward function affects the training process of trajectory planning reinforcement learning.

[0175] For example, in order to enable the multi-missile cooperative system to achieve its goal faster through training, the reward function in the embodiment of the present disclosure may include: step feedback reward, node target reward, and final target reward. In order to encourage the multi-missile cooperative system to collaboratively achieve the overall planning goal, the reward function can be set based on formula (21).

[0176]

[0177] in, represents the interceptor missile’s step-length feedback reward based on the remaining distance, R η It represents the interceptor missile’s step feedback reward based on the velocity lead angle. Indicates that the interceptor missile is based on the remaining time t go Consistent step size feedback reward, R Δf represents the node target reward of the interceptor missile, Rf Represents the final target reward based on the multi-missile coordinated system.

[0178] For example, the step-length feedback reward of the i-th interceptor missile based on the remaining distance can be generated based on formula (22):

[0179]

[0180] in, represents the lateral remaining distance of the i-th interceptor missile after the y-th step, = represents the remaining horizontal distance of the i-th interceptor missile after the y-1th step, and Δx represents the step size of each step. Based on this approach, by setting the reward for each interceptor missile based on the step size of the remaining distance, the problem of sparse rewards can be avoided.

[0181] For example, the step feedback reward of the i-th interceptor missile based on the velocity lead angle is generated based on formula (23):

[0182]

[0183] Based on this scheme, by setting the step-size feedback reward of each interceptor missile based on the velocity lead angle, it is possible to ensure that each interceptor missile flies as far as possible towards its respective terminal position.

[0184] For example, the i-th interceptor missile is set based on the remaining time t based on formula (24) go Consistent step size feedback reward

[0185]

[0186] in,

[0187] It should be noted that coordinated trajectory planning requires a multi-missile coordinated system to ensure consistency in coordination time. It is easier for a single interceptor missile to reach the expected terminal position than for multiple interceptors to reach their respective expected terminal positions. A single interceptor missile reaching the expected terminal position is also a prerequisite for the multi-missile coordinated system to achieve the planned goal. The node target reward value based on a single interceptor missile can be set. Final target reward based on multi-missile coordinated system (A reward for multiple interceptor missiles reaching the final target position). Represents the node target reward of the i-th interceptor missile. The node target reward of each interceptor missile is equal to K and is a preset constant.

[0188] For example, setting R f =1000,if max(s togoi )≤η.

[0189] in, represents the node target reward of the i-th interceptor missile, R f It represents the final target reward for multiple interceptor missiles to reach the predetermined terminal position. η is a constant that can be set according to the training accuracy requirements.

[0190] Discount factor: The discount factor γ ranges from 0 to 1 and indicates the degree to which future rewards influence current rewards. Generally speaking, the further away future behavior is from the present, the less impact it should have on the present. When γ = 0, the current reward depends only on the reward at the next moment. When γ = 1, future rewards are equally important to the current state, regardless of how far away they are from the present.

[0191] Since it is difficult to guarantee the convergence of the formula when γ=1, and the future behavior is uncertain, 1 is generally not selected in practical applications. In the embodiment of the present disclosure, it is set to 0.99 based on experience so that the obtained coordinated trajectory can comprehensively consider the overall performance.

[0192] Collaborative trajectory planning training strategy

[0193] Optionally, in the system trajectory planning method for a multi-missile system provided in an embodiment of the present disclosure, the above-mentioned S204 may further include the following S204d:

[0194] S204d. The progressive training strategy based on curriculum learning generates an intelligent collaborative trajectory planning model for multiple missiles through reinforcement learning training.

[0195] The progressive training strategy includes: a first training task and a second training task, the training order of the first training task is earlier than the training order of the second training task, and the convergence distance of the first training task is greater than the convergence distance of the second training task.

[0196] Generally, Course Learning (CL) refers to a method in machine learning that gradually increases the difficulty of tasks to speed up learning. In the embodiment of the present disclosure, a progressive training strategy for multi-projectile coordinated trajectory planning based on CL is adopted.

[0197] Task 1: The target prediction hit point is at any location, the distance convergence condition is set to τ = 5km, and the average reward convergence threshold R ave is 1000;

[0198] Task 2: The target prediction hit point is at any location, and the distance convergence condition is set to η = 1km, and the average reward convergence threshold R ave is 1000.

[0199] It should be noted that the target prediction impact point position in Task 1 and Task 2 is arbitrary and is used to reflect the z-axis caused by the target's lateral maneuver. T The process constraints are reflected in the algorithm as termination conditions, that is, if the heat flux exceeds the constraint threshold, or the dynamic pressure exceeds the constraint threshold, or the overload exceeds the threshold constraint, or the safety distance between the interceptors is less than the threshold constraint, then the collaborative trajectory planning is terminated, as shown in formula (25).

[0200]

[0201] For example, Figure 5 A training strategy diagram provided by the embodiment of the present disclosure is shown in FIG. Figure 5 As shown in Figure 2, the specific implementation process of the training strategy is as follows:

[0202] (1) Generate training networks for two different mission requirements of a multi-missile coordination system;

[0203] (2) Train the multi-missile coordination system network in Task 1 and obtain the parameters of the network;

[0204] (3) Migrate the parameters of the task 1 training network to the task 2 training network to complete the final training of the multi-missile coordination system and derive the multi-missile coordination trajectory planning strategy.

[0205] Among them, the training difficulty of Task 1 is less than that of Task 2.

[0206] Based on this scheme, during the collaborative trajectory planning task training process of the disclosed embodiment, by narrowing the distance convergence condition and setting learning tasks from easy to difficult, the collaborative guided ballistic trajectory planning capability of the multi-missile collaborative system is progressively trained, which can effectively improve the generalization of the model.

[0207] Example of the method for generating the initial solution of the determined collaborative trajectory

[0208] Step 1: Convert the multi-interceptor coordinated trajectory planning model into the state transition sequence of the MDP corresponding to DDPG.

[0209] Step 2: Based on the DDPG model, train according to the following offline training process.

[0210] 2-1: Initialize network parameters (ω μ ,ω μ′ ,ω Q ,ω Q′ ) and memory capacity (e.g. 1×10 6 ), start the loop;

[0211] 2-2: Set the initial state of the multi-ballistic system (e.g., Table 3), randomly select the target point position, and start a single coordinated ballistic cycle;

[0212] 2-3: Select the action of the multi-projectile system based on the current strategy (the above-mentioned progressive training strategy) and exploration noise;

[0213] 2-4: Execute the action and obtain the corresponding state and reward value, and store the state transition sequence in the memory bank;

[0214] 2-5: Randomly sample a small batch of training data from the memory bank (e.g., batch_size = 256), update the network parameters, and the single collaborative trajectory cycle ends;

[0215] 2-6: Determine whether the coordinated trajectory has achieved the training goal of formula (25). If so, proceed to the next step. If not, return to step 2-2.

[0216] 2-7: The loop ends and the optimal network parameters and collaborative trajectory planning strategy are output.

[0217] Step 3: Based on the above training strategy, perform offline training.

[0218] Step 4: Generate the initial solution of the collaborative trajectory online based on the trained network parameters.

[0219] Experimental example:

[0220] Experimental parameters: Assume that the multi-missile coordination system has completed the trajectory planning of the primary guidance phase before the mid-range guidance phase, the interceptor missiles are of the same model, the interceptor missile mass m = 900 kg, and the interceptor missile characteristic area S C =0.7m 2 , the maximum detection range of the seeker R max =80km, detection field of view The maximum rotation angle χ = 6°, the target HP2R radius r = 6 km, and the initial state settings of the three interceptor missiles are shown in Table 2.

[0221] Table 2

[0222] <![CDATA[x0 / (km)]]> <![CDATA[h0 / (km)]]> <![CDATA[z0 / (km)]]> <![CDATA[v0 / (m×s -1 )]]> <![CDATA[θ0 / (°)]]> <![CDATA[Ψ v0 / (°)]]> <![CDATA[M1]]> 0 70 -3 3000 -5 0 <![CDATA[M2]]> 0 70 0 3000 -5 0 <![CDATA[M3]]> 0 70 3 3000 -5 0

[0223] Specifically, the number of grid points in the iterative solution of the convex optimization algorithm is set to 200, and the maximum number of iterations is 500. Among them, the process constraints of each interceptor missile are: Q max =1×106W / m2,q max =1×105Pa,n min =8; the control quantity constraints of each interceptor missile are: α max =30°,ζ max=85°, the minimum spacing between each interceptor missile is constrained to be d min =1×103m; Trust region constraint radius of each interceptor missile Convergence threshold

[0224] Reinforcement learning offline training uses a Python language simulation platform, using Pytorch modules for training and corresponding "CUDA" devices. The network parameter optimizer is Adam Optimizer, and the convex optimization uses the ECOS-BB solver for sequence planning. Finally, MATLAB software is used for simulation and drawing. The network learning rate is set to 0.0001. The memory capacity is set to 1×10 6 , the number of random training samples is 256, and the maximum number of training times is set to 1×10 5 The step size of the embodiment of the present disclosure is

[0225] Simulation Results

[0226] Experiment 1: Convergence Verification

[0227] In order to verify the convergence of the CL-based progressive training strategy for multi-missile coordinated trajectory planning proposed in the embodiment of the present disclosure, the embodiment of the present disclosure defines the average reward (the average value of the total reward of 100 rounds, and when the number of rounds is less than 100 at the beginning, the average value of the total reward of only the rounds is used) to observe the convergence of the training method, and sets the average reward convergence threshold to 1000. Figure 6 This is one of the simulation results provided by the embodiment of the present disclosure, such as Figure 6 As shown. Figure 6 It can be seen that the progressive training strategy designed in the embodiment of the present disclosure can make the algorithm converge and achieve the expected effect during training. In particular, the algorithm reaches the average reward convergence threshold in Task 1 after 4353 training iterations (see Figure 6 (a) Progressive task training 1), in task 2, training iterations 23157 times to reach the average reward convergence threshold (see Figure 6 (b) Progressive task training 2). When the direct training strategy is used, the algorithm training iterations are repeated until the maximum number of training times (1×10 5 times) still does not reach the convergence threshold (see Figure 6 (c) Progressive task training 2), and the algorithm's average reward eventually shows a divergent trend. Simulation results show that the training strategy designed by the disclosed embodiment has stronger convergence and better training results for complex, high-dimensional multi-projectile coordinated trajectory planning tasks. It also proves that direct training strategies are unsuitable for complex, high-dimensional training tasks, and can easily cause the algorithm to fail to find the optimal solution.

[0228] Experiment 2: Validation

[0229] To verify the effectiveness of the optimal network parameters obtained through training using the disclosed embodiments, the target predicted impact point location parameters were randomly assigned, and 1,000 Monte Carlo tests were performed on the multi-missile coordination system training network. The simulation data is shown in Table 3.

[0230] Table 3

[0231] Success rate / % Reward mean (maximum, minimum) Average terminal position error / m Average time / s 98.9 1028.38(1265.45,426.92) 638.64 0.39

[0232] Table 3 shows that the DDPG-based collaborative trajectory initial solution generation method has a success rate of 98.9%. The average reward for the test is 1028.38, exceeding the average reward convergence threshold (1000). The average error in the terminal position of each trajectory is 638.64 meters, which is less than the training convergence condition (1 km). The average time required to generate the collaborative trajectory initial solution is 0.39 seconds. Simulation results show that the collaborative trajectory initial solution generated by the method of the disclosed embodiment not only meets basic guidance requirements but also can be generated quickly online. This also proves that the optimal network parameters trained by the disclosed embodiment are effective.

[0233] Experiment 3: Performance Verification

[0234] In order to verify that the collaborative trajectory convex optimization method (CVX-1) based on the collaborative trajectory initial solution generation method (DDPG) provided by the embodiment of the present disclosure has better performance, the target prediction impact point position parameters are randomly selected and a simulation comparison experiment is conducted with the collaborative trajectory convex optimization method (CVX-2) based on the traditional initial solution generation method (TISG). Figure 7 This is the second schematic diagram of the simulation results provided by the embodiment of the present disclosure, and the data comparison results are shown in Table 4. It should be noted that the collaborative trajectory convex optimization method based on the collaborative trajectory initial solution is a related mature technology.

[0235] Table 4

[0236]

[0237] From the simulation results, we can see that the initial solution of the collaborative trajectory generated by the DDPG method is better than that of the TISG method. Figure 7 The coordinated trajectory planning of CVX-1 shown in (a) is as follows: Figure 7 As shown in the collaborative trajectory planning of CVX-1 in (b) and Table 4, the terminal position deviations of each interceptor missile using the DDPG method (626.29m, 303.45m, 551.74m) are significantly smaller than those using the TISG method (12571.07m, 8920.14m, 11935.12m). The coverage of the collaborative detection field of view of the target HP2R by the DDPG method is significantly higher than that of the TISG method (see Figure 4). Figure 7 The collaborative detection field of view of CVX-1 shown in (c) Figure 7 The DDPG method can satisfy all the process constraints (see Figure 8 (a), (b), (c) and (d) in ), while the TISG method exceeds the dynamic pressure constraint limit (see Figure 9 (a), (b), (c) and (d) in Fig. 3. In addition, the DDPG algorithm has a good performance under the terminal angle constraint (see Figure 10 The control amount of CVX-1 and the control amount of CVX-2 shown in (a), (b), (c) and (d) in the figure also have obvious advantages over the TISG method in terms of control amount and coordination time error (see Table 4). Due to the high-quality coordinated trajectory initial solution provided by the DDPG method, the number of iterations of the CVX-1 method (2 times) is significantly less than that of the CVX-2 method (6 times). As shown in Table 4, the optimization time is only 1.08s, which is nearly 3 times higher than the solution efficiency of the CVX-2 method (3.24s). The final simulation time of this "multi-projectile coordinated trajectory convex optimization method combined with deep reinforcement learning" provided by the embodiment of the present disclosure is very short, and online planning can be achieved. At the same time, the terminal position deviations of each interceptor missile using the CVX-1 method (0.01m, 0.04m, 0.09m) are significantly reduced compared to the CVX-2 method (0.11m, 0.48m, 0.32m). This means that, under the condition that both methods can satisfy all constraints, the CVX-1 method significantly outperforms the CVX-2 method in terms of solution accuracy. Simulation results demonstrate that the collaborative trajectory convex optimization method of the disclosed embodiment, while generating high-quality collaborative trajectory initial solutions, not only exhibits excellent optimization performance but also significantly improves optimization solution efficiency, enabling it to meet the requirements of collaborative trajectory online optimization.

[0238] In the multi-projectile intelligent collaborative trajectory online planning method provided by the embodiments of the present disclosure, the concept of lateral distance domain is first proposed, the motion model is converted from the time domain to the lateral distance domain, and a discrete trajectory convex optimization model is established in the lateral distance domain; then, the corresponding MDP is generated according to the characteristics of the multi-projectile collaborative trajectory planning model, a CL-based progressive training strategy for multi-projectile collaborative trajectory planning is designed, and a DDPG-based method for quickly generating the collaborative trajectory initial solution is proposed; and combined with the multi-projectile collaborative trajectory convex optimization method of deep reinforcement learning, on the basis of obtaining a high-quality collaborative trajectory initial solution, the solution efficiency and accuracy of the algorithm can be improved, so that it meets the needs of collaborative trajectory online optimization.

[0239] Corresponding to the embodiments of the aforementioned method, the present disclosure also provides an embodiment of an intelligent collaborative trajectory online planning device for multiple projectiles and a computer device used therein.

[0240] The embodiment of the intelligent collaborative trajectory online planning device for multiple missiles disclosed in the present invention can be applied to computer equipment, such as servers or terminal devices. The embodiment of the intelligent collaborative trajectory online planning device for multiple missiles can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically sensed intelligent collaborative trajectory online planning device for multiple missiles, it is formed by the processor of the project projection in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, such as Figure 11 The figure shows a hardware structure diagram of the computer device where the intelligent collaborative trajectory online planning device for multiple missiles in the embodiment of the present disclosure is located. Figure 11 In addition to the processor 1110, memory 1130, network interface 1120, and non-volatile memory 1140 shown, the server or electronic device where the device 1131 is located in the embodiment may also include other hardware according to the actual function of the computer device, which will not be described in detail.

[0241] like Figure 12 As shown, Figure 12 The present invention provides a schematic structural diagram of an intelligent collaborative trajectory online planning device for multiple missiles provided in an embodiment of the present invention. The intelligent collaborative trajectory online planning device 1200 for multiple missiles includes: a time domain motion model generation module 1201, a distance domain motion model conversion module 1202, a Markov decision process generation module 1203, an intelligent collaborative trajectory training module 1204, a convexification module 1205, a discretization module 1206 and an online optimization collaborative trajectory generation module 1207; the time domain motion model generation module 1201 is used to generate a time domain motion model of multiple interceptor missiles in a multi-missile collaborative system; the distance domain motion model conversion module 1202 is used to convert the time domain motion model into a lateral distance domain motion model, generate constraints and an objective function for collaborative trajectory planning, and the objective function indicates the terminal distance error reached by multiple interceptor missiles; the Markov decision process generation module 1203 is used to Based on the lateral distance domain motion model, constraints and objective function, a Markov decision process for online planning of intelligent collaborative trajectories of multiple missiles is generated; an intelligent collaborative trajectory training module 1204 is used to generate an intelligent collaborative trajectory planning model of multiple missiles through reinforcement learning training based on the Markov decision process; a convexification module 1205 is used to convexify the lateral distance domain motion model, constraints and objective function; a discretization module 1206 is used to discretize the lateral distance domain motion model, constraints and objective function; an online optimized collaborative trajectory generation module 1207 is used to use the intelligent collaborative trajectory planning model of multiple missiles as the initial solution for online generation of the collaborative trajectory of multiple missiles, and based on the convexed and discretized lateral distance domain motion model, constraints and objective function, interceptor missile parameters, and maneuvering target parameters, online optimization is used to generate the initial solution of the collaborative trajectory of multiple interceptor missiles before the mid-range guidance stage.

[0242] The embodiment of the present disclosure provides an online intelligent collaborative trajectory planning device for multiple missiles. First, a time domain motion model of multiple interceptor missiles in a multi-missile collaborative system is generated, the time domain motion model of the multiple interceptor missiles is converted into a lateral range domain motion model, and constraints and objective functions for collaborative trajectory planning are generated. Then, a Markov decision process for online intelligent collaborative trajectory planning of multiple missiles is generated based on the lateral range domain motion model, constraints and objective functions, and an intelligent collaborative trajectory planning model for multiple missiles is generated through reinforcement learning training. Further, the lateral range domain motion model, constraints and objective functions are convexified and discretized, and then the generated intelligent collaborative trajectory planning model for multiple missiles is used as the initial solution for online generation of the collaborative trajectory of multiple missiles. Based on the convexified and discretized lateral range domain motion model, constraints and objective functions, interceptor missile parameters, and maneuvering target parameters, the collaborative trajectory of the multiple interceptor missiles is generated through online optimization before the mid-range guidance stage. Since the time domain motion model is converted into a distance domain motion model, the MDP of collaborative trajectory planning can be accurately generated based on the distance domain motion model, constraints and objective function. Through training through reinforcement learning training strategy, an intelligent collaborative trajectory planning model for multiple missiles can be quickly and accurately generated. Compared with related technologies, the computational efficiency and accuracy of the initial solution of the collaborative trajectory can be improved. Not only can the subsequent multi-missile collaborative trajectory convex optimization stage be quickly entered, but also based on the collaborative trajectory planning model, the collaborative trajectory can be efficiently and accurately generated online, realizing online trajectory planning of multiple missiles before the mid-range guidance stage, thereby realizing the online optimization generation of intelligent collaborative trajectories for multiple missiles in mid-range guidance, and providing a guarantee for multiple interceptor missiles to complete full coverage of the target H2PR and effectively intercept it.

[0243] Correspondingly, the present disclosure also provides an intelligent collaborative trajectory online planning device for multiple missiles, which includes a processor; a memory for storing processor executable instructions; wherein the processor is configured to: generate a time domain motion model of multiple interceptor missiles in a multi-missile collaborative system; convert the time domain motion model into a lateral range domain motion model, generate constraints and an objective function for collaborative trajectory planning, and the objective function indicates the terminal distance error reached by the multiple interceptor missiles; generate a Markov decision process based on the intelligent collaborative trajectory online planning of multiple missiles based on the lateral range domain motion model, the constraints and the objective function; and generate an intelligent collaborative trajectory planning model for multiple missiles through reinforcement learning training based on the Markov decision process; convexify and discretize the lateral range domain motion model, the constraints and the objective function; use the intelligent collaborative trajectory planning model of multiple missiles as the initial solution for the online generation of the collaborative trajectory of multiple missiles, and based on the convexified and discretized lateral range domain motion model, the constraints and the objective function, the interceptor missile parameters, and the maneuvering target parameters, online optimize and generate the collaborative trajectory of multiple interceptor missiles before the mid-range guidance stage.

[0244] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the various steps in the embodiment of the above-mentioned method for online intelligent collaborative trajectory planning of multiple projectiles.

[0245] The present disclosure also provides a computer device, which includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the computer-readable instructions, when executed by the processor, implement the various steps in the above-mentioned embodiment of the method for online intelligent collaborative trajectory planning of multiple missiles.

[0246] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0247] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the disclosed solution. Those of ordinary skill in the art can understand and implement it without paying any creative work.

[0248] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0249] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the inventions claimed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0250] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0251] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A multi-projectile intelligent collaborative trajectory online planning method, characterized in that: The method comprises: Generate time domain motion models of multiple interceptor missiles in a multi-missile coordinated system; Converting the time-domain motion model to a lateral range-domain motion model to generate constraints and an objective function for collaborative trajectory planning, wherein the objective function indicates the terminal range error reached by the multiple interceptor missiles; Based on the lateral range domain motion model, the constraints, and the objective function, a Markov decision process for online planning of intelligent collaborative trajectories of multiple missiles is generated; and based on the Markov decision process, a reinforcement learning training is used to generate the intelligent collaborative trajectory planning model of the multiple missiles; Convexifying and discretizing the lateral range domain motion model, the constraint conditions, and the objective function; The intelligent collaborative trajectory planning model of multiple missiles is used as the initial solution for online generation of the collaborative trajectory of multiple missiles. Based on the convexed and discretized lateral range domain motion model, the constraints and the objective function, the interceptor missile parameters, and the maneuvering target parameters, the collaborative trajectory of the multiple interceptor missiles is generated by online optimization before the mid-range guidance stage.

2. The method according to claim 1, characterized in that The constraints and objective functions for generating collaborative trajectory planning include: generating a first state equation, a first constraint condition, and a first objective function of the multi-missile coordination system in a lateral range domain based on motion parameters of each interceptor missile in the time-domain motion model of the multiple interceptor missiles; Among them, the first state equation includes: the state equation of each interceptor missile, the control vector matrix of each interceptor missile and the control vector of each interceptor missile; the first constraint condition includes at least one of the following: process constraint, control quantity constraint, terminal constraint and trust region constraint; the process constraint includes at least one of the following: heat flux constraint, dynamic pressure constraint, overload constraint; the control quantity constraint includes at least one of the following: attack angle constraint, roll angle constraint, and safety distance constraint between each interceptor missile; the terminal constraint includes at least one of the following: ballistic inclination terminal value constraint and ballistic deviation terminal value constraint.

3. The method according to claim 2, characterized in that The convexifying and discretizing the lateral range domain motion model, the constraint conditions and the objective function includes: Based on the reference trajectory, linearize the first state equation into a second state equation and linearize the second constraint condition; Setting a first trust region function, wherein the first trust region function is used to adjust the linearization error; generating a second objective function based on the first objective function and a preset additional term, wherein the preset additional term is used to adjust a linearization error of the first objective function; The second constraint condition includes at least one of the following: a constraint on heat flux density, a constraint on dynamic pressure, a constraint on overload, and a constraint on the safe distance between each interceptor missile.

4. The method according to claim 3, characterized in that The convexifying and discretizing the lateral range domain motion model, the constraint conditions and the objective function includes: discretizing the second state equation, the second constraint condition, and the second objective function; generating a second trust region function based on the first trust region function and the trust region relaxation coefficient; A third objective function is generated based on the discretized second objective function and the trust region relaxation coefficient, wherein the trust region relaxation coefficient is used to adjust the error caused by discretization.

5. The method according to claim 4, characterized in that The online optimization and generation of the coordinated trajectories of the multiple interceptor missiles before the mid-range guidance phase includes: Before the mid-range guidance phase, determining an initial solution trajectory of the multi-missile coordination system that satisfies target constraints, the second trust region function, and the third objective function based on the discretized second state equation; Among them, the target constraint condition includes at least one of the third constraint conditions, and / or at least one of the second constraint conditions after discretization; the third constraint condition includes: the constraint of the angle of attack, the constraint of the roll angle, and the constraint of the terminal value of the ballistic inclination angle and the constraint of the terminal value of the ballistic deviation angle.

6. The method according to any one of claims 1 to 5, characterized in that The generating of a Markov decision process for online planning of intelligent collaborative trajectories of multiple missiles based on the lateral range domain motion model, the constraint conditions, and the objective function comprises: Generate a state set of the multi-missile coordinated system, an action set of the multi-missile coordinated system, and a reward function of the multi-missile coordinated system; Among them, the state set of the multi-missile cooperative system includes at least one of the following items of the interceptor missile: position coordinates, speed, ballistic inclination angle, ballistic deviation angle, angle of attack, roll angle, remaining distance, component of the velocity lead angle in the longitudinal plane and component of the transverse plane; the action set of the multi-missile cooperative system includes the interceptor missile: angle of attack and roll angle; the reward function of the multi-missile cooperative system includes at least one of the following items: step feedback reward based on remaining distance, step feedback reward based on velocity lead angle, step feedback reward based on remaining time consistency, node target reward based on single interceptor missile and final target reward based on the multi-missile cooperative system.

7. The method according to claim 6, characterized in that No. The step feedback reward for each interceptor missile based on the remaining distance is: ; said The step-size feedback reward of an interceptor missile based on the velocity lead angle is: ; said The step-size feedback reward for an interceptor missile based on the consistency of the remaining time is: ; The ultimate goal reward of the multi-missile coordination system = ; in, Indicates the The remaining lateral distance of the interceptor missile after the yth step, Indicates the The remaining lateral distance of the interceptor missile after the y-1th step, Indicates the step size of each step, , Indicates the Interceptor missile node target rewards, each interceptor missile node target reward is equal to , and is a preset constant.

8. The method according to claim 7, characterized in that The generation of the multi-projectile intelligent collaborative trajectory planning model through reinforcement learning training includes: The progressive training strategy based on curriculum learning generates the intelligent collaborative trajectory planning model of the multiple missiles through reinforcement learning training; wherein, the progressive training strategy includes: a first training task and a second training task, the training order of the first training task is earlier than the training order of the second training task, and the convergence distance of the first training task is greater than the convergence distance of the second training task.

9. The method according to any one of claims 1 to 5, characterized in that The online optimization and generation of the coordinated trajectories of the multiple interceptor missiles before the mid-range guidance phase based on the convexified and discretized lateral range domain motion model, the constraint conditions and the objective function, the interceptor missile parameters, and the maneuvering target parameters includes: determining mid-range guidance terminal positions of the plurality of interceptor missiles based on the target HP2R determined by the maneuvering target parameter and the number of the plurality of interceptor missiles; Based on the convexed and discretized lateral range domain motion model, the constraints and the objective function, the parameters of the multiple interceptor missiles in the multi-missile coordination system, and the mid-range guidance terminal positions of the multiple interceptor missiles, the coordinated trajectories of the multiple interceptor missiles are generated by online optimization before the mid-range guidance stage.

10. The method according to claim 1, characterized in that The lateral range domain motion model includes: ; in, 、 、 They represent the first The dimensionless northeast celestial coordinate of the interceptor missile's center of mass, Indicates the The dimensionless velocity of the interceptor missile, Indicates the The ballistic inclination of the interceptor missile, Indicates the The ballistic deviation angle of the interceptor missile, Indicates the The angle of attack of the interceptor missile, Indicates the The roll angle of the interceptor missile, Indicates the The dimensionless drag acceleration of the interceptor missile, Indicates the The dimensionless lift acceleration of the interceptor missile, Indicates the The dimensionless gravitational acceleration of the interceptor missile, , , , Indicates the The drag coefficient of the interceptor missile, , , and represents the resistance parameter, Indicates the The lift coefficient of the interceptor missile, , and represents the lift parameter, Indicates the The dynamic pressure of the interceptor missile, , Indicates the The density of the air in which the interceptor missile is located, , represents the air density at sea level, Indicates the reference altitude, represents the radius of the Earth, Indicates the The characteristic area of ​​the interceptor missile, Indicates the The mass of the interceptor missile, is the acceleration due to gravity at sea level.

Citation Information

Patent Citations

  • Intelligent vehicle multi-scene trajectory planning method based on Frenet coordinate system

    CN113886764A

  • Multi-missile cooperative hunting and intercepting method and device based on online iterative optimization algorithm, medium and product

    CN118444570A