Interception guidance strategy acquisition method and device, equipment and storage medium

By combining reinforcement learning and differential game theory, the Bellman equation for nonlinear interception dynamics and game value function is obtained, which solves the problem of insufficient accuracy of traditional interception guidance strategies under highly maneuverable targets and realizes optimal interception control under unknown dynamic information.

CN119739039BActive Publication Date: 2025-11-25BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411896177.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-11-25
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Traditional interception and guidance strategies suffer from insufficient guidance accuracy when facing highly maneuverable targets, and require a known and accurate dynamic model.

Method used

An interception guidance strategy acquisition method combining reinforcement learning and differential game theory is adopted. The optimal interception guidance control strategy is obtained by iteratively solving the Bellman equation based on nonlinear interception dynamics equation, game value function and iterative control strategy.

Benefits of technology

The optimal control strategy is generated without relying on specific dynamic information, which improves the interception accuracy and guidance performance of maneuvering escape targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739039B_ABST
    Figure CN119739039B_ABST
Patent Text Reader

Abstract

The application provides an interception guidance strategy acquisition method and device, equipment and a storage medium. The method comprises the following steps: obtaining a first Bellman equation, which is an off-policy integral reinforcement learning Bellman equation obtained based on a nonlinear interception dynamics equation, a game value function and an iterative control strategy; obtaining control input data and state output data; and iteratively solving the first Bellman equation based on the control input data and the state output data to obtain an optimal interception guidance control strategy. In this way, the optimal control strategy is found in a data-driven manner without relying on specific dynamics information, so that the optimal control strategy is used to achieve the purpose of high-precision interception of a high-maneuvering escape party by a pursuit party.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of aircraft interception guidance, and in particular to an interception guidance strategy acquisition method and device, equipment and a storage medium. BACKGROUND

[0002] As a key link in the air defense system, interception guidance strategy involves a series of complex technical operations and strategic decisions for accurate tracking and interception of incoming targets. Its core goal is to ensure that the missile system can quickly and accurately lock onto the target and complete the interception task in the shortest time. In the face of the increasing mobility and uncertainty of targets in modern warfare, the importance of interception guidance strategy is increasingly prominent.

[0003] During the interception process, the missile not only needs to track the initial trajectory of the target, but also needs to adjust the interception decision in real time according to the target's maneuvering strategy. For example, when the target suddenly changes its flight direction or speed, the interception missile must be able to quickly respond and adjust its flight trajectory and speed to ensure that it always remains in the best interception position.

[0004] However, the traditional interception guidance strategy has the problem of insufficient guidance accuracy when facing high-maneuvering targets. Therefore, how to achieve high-precision interception of missiles when facing high-maneuvering targets has become a top priority. SUMMARY

[0005] Therefore, the present application provides an interception guidance strategy acquisition method, device, equipment and storage medium to solve the technical problems in the background art.

[0006] In a first aspect, the present application provides an interception guidance strategy acquisition method, comprising: obtaining a first Bellman equation, the first Bellman equation being an off-policy integral reinforcement learning Bellman equation obtained based on a nonlinear interception dynamics equation, a game value function and an iterative control strategy; obtaining control input data and state output data; based on the control input data and the state output data, iteratively solving the first Bellman equation to obtain an optimal interception guidance control strategy.

[0007] In a second aspect, the present application provides an interception guidance strategy acquisition device, comprising: a first equation obtaining module for obtaining a first Bellman equation, the first Bellman equation being an off-policy integral reinforcement learning Bellman equation obtained based on a nonlinear interception dynamics equation, a game value function and an iterative control strategy; an input and output data obtaining module for obtaining control input data and state output data; a control strategy generating module for iteratively solving the first Bellman equation based on the control input data and the state output data to obtain an optimal interception guidance control strategy.

[0008] In a third aspect, an electronic device is provided, which includes a processor and a memory storing computer program instructions, wherein the computer program instructions, when executed by the processor, implement the steps of the method for obtaining an interception guidance strategy according to any of the embodiments of the present application.

[0009] In a fourth aspect, a computer readable storage medium is provided, which stores computer executable programs, wherein the computer executable programs, when executed by a processor, implement the steps of the method for obtaining an interception guidance strategy according to any of the embodiments of the present application.

[0010] To sum up, the method, device, electronic device and storage medium for obtaining an interception guidance strategy provided by the embodiments of the present application have at least the following beneficial effects: the embodiments of the present application first obtain a first Bellman equation, which is an off-policy integral reinforcement learning Bellman equation obtained based on a nonlinear interception dynamics equation, a game value function and an iterative control strategy, then obtain control input data and state output data, and then iteratively solve the first Bellman equation based on the control input data and the state output data to obtain an optimal interception guidance control strategy. The embodiments of the present application can obtain the first Bellman equation based on the nonlinear interception dynamics equation, the game value function and the iterative control strategy, realize the generation of an optimal control strategy in a data-driven manner without relying on specific dynamics information, make the control input optimal, and better realize the guidance interception task of a maneuvering escape object with high interception accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0012] Figure 1 A flowchart of a method for obtaining an interception guidance strategy provided by an embodiment of the present application is shown;

[0013] Figure 2 A flowchart of another method for obtaining an interception guidance strategy provided by an embodiment of the present application is shown;

[0014] Figure 3 A convergence graph of a neural network weight parameter evaluation method provided by an embodiment of the present application is shown;

[0015] Figure 4 ​A pursuit neural network weight parameter provided by an embodiment of the application is shown a convergence graph

[0016] Figure 5 An escape neural network weight parameter provided by an embodiment of the application is shown a convergence graph

[0017] Figure 6 A three-dimensional diagram showing a missile intercepting a target according to an optimal interception control strategy provided by an embodiment of the application is shown

[0018] Figure 7 A structural diagram of an interception guidance strategy acquisition device provided by an embodiment of the application is shown

[0019] Figure 8 A structural diagram of an electronic device provided by an embodiment of the application is shown. DETAILED DESCRIPTION

[0020] In order to make the above and other features and advantages of the present application clearer, the present application will be further described below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explanation and are only illustrative and are not restrictive.

[0021] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. It will be apparent, however, to one skilled in the art that the present application can be practiced without the specific details. In other instances, well-known steps or operations are not described in detail in order to avoid obscuring the present application.

[0022] The traditional guidance method has the following problems when facing a high-maneuvering target: (1) the pursuit guidance method cannot achieve omnidirectional interception of the target, and is generally used for missiles attacking low-speed moving or stationary targets; (2) the implementation of the parallel approach method requires accurate measurement of the state information of the missile and the target, and strictly satisfies the motion relationship of the parallel approach method, which requires strict guidance system and is difficult to implement in engineering; (3) the proportional guidance method has poor guidance accuracy when the target appears to be rapidly, highly maneuvering and launched at a large off-axis angle; (4) the traditional guidance method lacks performance indicators in the guidance process, and the control input is not optimal; (5) the optimal guidance method has optimality, but the solving process requires accurate dynamic models of both parties.

[0023] To solve the above problems, the inventors provide an interception guidance strategy acquisition method, which is a one-to-one interception guidance strategy acquisition method combining reinforcement learning and differential game, and can be applied to the scenario of interceptor intercepting a high-maneuvering target, thereby breaking through the defects of the existing various guidance methods, such as insufficient guidance accuracy, lack of optimality, and the need for an accurate known dynamic model, and improving the missile interception guidance performance.

[0024] It should be noted that in the embodiments of the present application, the pursuit side interception scenario can be regarded as a pursuit side and escape pursuit scenario, and the pursuit side is regarded as the pursuit side and the escape is regarded as the escape side.

[0025] Before introducing the interception guidance strategy acquisition method provided by the embodiments of the present application, the coordinate system involved in the interception guidance strategy acquisition method is first described. The ground inertial coordinate system body coordinate system and velocity coordinate system (airflow coordinate system) are used to describe the motion of the pursuit side and the escape side.

[0026] The ground inertial coordinate system E I takes the launch point of the pursuit side as the origin O I , the axis points to the projection of the launch direction in the horizontal plane, that is, the intersection line of the trajectory plane and the horizontal plane, the axis is located in the vertical plane containing the axis is perpendicular to the axis points vertically upward, and the axis points according to the right-hand rule.

[0027] The origin O B of the body coordinate system E B is the center of mass of the pursuit side or the escape side, the axis points forward along the longitudinal axis of the pursuit side or the escape side, is perpendicular to the axis points upward, and the axis points according to the right-hand rule.

[0028] The origin O A of the velocity coordinate system (airflow coordinate system) E A is consistent with the body coordinate system, the axis points to the direction of the center of mass velocity, the axis is located in the longitudinal symmetry plane of the pursuit side or the escape side and is perpendicular to the axis points upward, and the axis points according to the right-hand rule.

[0029] The embodiments of the present application provide an interception guidance strategy acquisition method, which can be implemented by an interception guidance strategy acquisition device.Figure 1 A flowchart of a method for obtaining an interception guidance strategy is shown in FIG. 1. The method for obtaining the interception guidance strategy can include the following steps. Figure 1

[0030] S11, obtaining a first Bellman equation.

[0031] In an embodiment, the first Bellman equation is an off-policy integral reinforcement learning Bellman equation obtained based on a nonlinear interception dynamics equation, a game value function, and an iterative control strategy. The first Bellman equation describes the relationship between the update of the game value function and the current reward and the subsequent state value. Through the first Bellman equation, the optimal game control input can be obtained under the condition of completely unknown system dynamics parameters.

[0032] wherein the nonlinear interception dynamics equation describes the nonlinear relationship between the state change of the pursuer and the evader and the control input and the current state. The game value function can be equal to the game performance function constructed according to the zero-sum differential game theory. The iterative control strategy can include the update strategy of the interceptor control strategy and the update strategy of the evader control strategy.

[0033] S12, obtaining control input data and state output data.

[0034] In an embodiment, the control input data can include the initial control strategy of the interceptor and the evader, denoted as The state output data can be multi-dimensional interception state space data, including but not limited to the position vector difference and the velocity vector difference between the pursuer and the evader, the speed of the pursuer, the flight path angle of the pursuer, the flight path azimuth angle of the pursuer, the speed of the evader, the flight path angle of the evader, and the flight path azimuth angle of the evader, etc. The state output data can be the system output state data recorded at a time interval T. Wherein, T can be set according to system requirements.

[0035] S13, iteratively solving the first Bellman equation based on the control input data and the state output data to obtain an optimal interception guidance control strategy.

[0036] In an embodiment, the iterative solution can mean that starting from an arbitrary initial control strategy, the game value function of each state and the control strategy of both parties are repeatedly updated by applying the first Bellman equation based on the state output data until the optimal control strategy is converged. The optimal interception guidance control strategy can mean the optimal control input strategy that satisfies the Nash equilibrium condition. The optimal interception guidance control strategy has optimality.

[0037] wherein the Nash equilibrium condition can be expressed as:

[0038]

[0039] wherein J represents a game performance function, V * (x(t)) represents an optimal game value function.

[0040] It should be noted that the convergence of the control strategies of both parties to the optimal control strategy can be determined according to a convergence condition. The convergence condition can be that the game value function of the kth iteration converges to the optimal game value function, or the positions of both parties are close, i.e., the distance between both parties is less than a preset threshold. Wherein k is not less than 0.

[0041] In the above embodiment, the first Bellman equation is obtained, which is an off-policy integral reinforcement learning Bellman equation obtained based on the nonlinear interception dynamics equation, the game value function and the iterative control strategy, so that the interception problem can be more accurately described and solved by using the first Bellman equation based on the nonlinear interception dynamics equation, the game value function and the iterative control strategy, and a framework suitable for off-policy integral reinforcement learning is formed. The control input data and the state output data are obtained, and based on the control input data and the state output data, the first Bellman equation is iteratively solved to obtain the optimal interception guidance control strategy, so that by using the off-policy integral reinforcement learning method, the optimal control strategy can be found in a data-driven manner without relying on specific dynamics information, so that the control input is optimal, and the optimal control strategy can better achieve the guidance interception task of the maneuvering escape party, and the interception precision is high.

[0042] In some embodiments, the first Bellman equation can be represented as the following equation.

[0043]

[0044] wherein u M and u T represent the initial input control strategy, and is the kth iteration optimal input control strategy, and is the k+1th optimal input control strategy, x represents the state output data, T represents the time interval, a represents the attack angle, t represents the state sampling start time, Q C represents the augmented state weight coefficient, R M represents the first control weight coefficient, R T represents the second control weight coefficient, and t represents the integral time.

[0045] In the above formula (2), the dynamics information f, g T , g M is not included, only the fixed control input u M , u Tand the output state information x, the game value function V can be updated simultaneously by solving equation (2) k and the optimal control input of both sides The initial control input u M , u T and the updated control strategy may be different. Among them, the initial control strategy needs to be a stable control strategy, and with exploration noise. The initial control strategy is used to collect different state output data.

[0046] The off-policy integral reinforcement learning algorithm proposed by the inventor can be divided into two stages. First, in the first stage, a fixed initial exploratory control strategy is applied to the system, and the output state data of the system in the time interval T is collected, and the output state data is processed into a data matrix form. Secondly, in the second stage, without the need for system dynamics knowledge, the information collected in the first stage is repeatedly used to update the control strategy sequence until it converges to the optimal control strategy. In addition, since the first Bellman equation is a scalar equation, after sufficient system data is collected, a linear matrix solving method can be used for solving.

[0047] Therefore, in some embodiments, S13, based on the control input data and the state output data, iteratively solving the first Bellman equation to obtain the optimal interception guidance control strategy, can include: converting the input control strategy and the state output data into a first data matrix and a second data matrix, respectively; based on the evaluation neural network, the pursuit side neural network and the escape side neural network, fitting the game value function to be updated in the first Bellman equation, the pursuit side control strategy and the escape side control strategy, respectively, to obtain the fitted game value function to be updated, the fitted pursuit side control strategy and the fitted escape side control strategy; based on the fitted game value function to be updated, the fitted pursuit side control strategy and the fitted escape side control strategy, converting the first Bellman equation into a fitted equation; based on the first data matrix, the second data matrix, the fitted game value function to be updated, the fitted pursuit side control strategy and the fitted escape side control strategy, converting the fitted equation into a linear matrix equation; solving the linear matrix equation to obtain the optimal interception guidance control strategy.

[0048] Here, the first data matrix can be a matrix representation of the control input strategy, can include a control data matrix of the pursuit side and a control data matrix of the escape side, and can be represented by a basis function of the pursuit side neural network or a basis function of the escape side neural network. The first data matrix can be specifically represented by the following formula.

[0049]

[0050] wherein, represents the control data matrix of the pursuit side, denotes the control data matrix of the evader, e denotes an index, t1 to t q denote the start or end time of the integral of each element in the data matrix respectively, τ denotes the integral time, φ(x(·)) denotes the basis function of the pursuer neural network, denotes the basis function of the evader neural network, u M denotes the control strategy of the pursuer, u T denotes the control strategy of the evader.

[0051] The second data matrix can be a matrix representation of the state output data, and can include δ xx , I xx , I φφ and The second data matrix can be represented by the following formula.

[0052]

[0053] wherein δ xx denotes the first component of the second data matrix, I xx denotes the second component of the second data matrix, I φφ denotes the third component of the second data matrix, denotes the fourth component of the second data matrix, α denotes the attack angle, T denotes the time interval, φ(x(·)) denotes the basis function of the pursuer neural network, denotes the basis function of the evader neural network, σ(x(·)) denotes the basis function of the evaluation neural network.

[0054] The "evaluation neural network", "pursuer neural network" and "evader neural network" involved in the embodiments of the present application are all neural networks designed according to specific task requirements. In the embodiments of the present application, the three neural networks of the evaluation neural network, the pursuer neural network and the evader neural network are used to approximate the to-be-updated game value function V k , the pursuer control strategy and the evader control strategy to obtain the fitted to-be-updated game value function the fitted pursuer control strategy and the fitted evader control strategy

[0055] The fitted to-be-updated game value function can be represented by formula (9).

[0056]

[0057] Fitted pursuer control policy which can be represented as equation (10).

[0058]

[0059] Fitted evader control policy which can be represented as equation (11).

[0060]

[0061] wherein σ(x), φ(x), are basis functions of the evaluation neural network, the pursuer neural network and the evader neural network respectively, are constant weight matrices of the evaluation neural network, the pursuer neural network and the evader neural network respectively, and are objects to be updated in the iteration process.

[0062] The “fitted equation” involved in the embodiments of the present application can be based on the fitted to-be-updated game value function Fitted pursuer control policy and fitted evader control policy represent the Bellman equation corresponding to the first Bellman equation. That is, substituting equations (9) to (11) into the first Bellman equation can obtain the following fitted equation.

[0063]

[0064] wherein each term in the fitted equation can be represented by a Kronecker product.

[0065]

[0066] wherein R T represents the first control weight coefficient, R M represents the second control weight coefficient, I φn represents the data matrix of φ(x), represents the data matrix of , and vec(·) represents a vector.

[0067] The “linear matrix equation” involved in the embodiments of the present application can be represented by the first data matrix, the second data matrix, the fitted to-be-updated game value function, the fitted pursuer control policy and the fitted evader control policy. Substituting the first data matrix, the second data matrix, the fitted to-be-updated game value function, the fitted pursuer control policy and the fitted evader control policy into the fitted equation can obtain the following linear equation.

[0068]

[0069] Among them, H ik Matrix and Γ ik The matrices can be represented by the following formulas.

[0070]

[0071] In this embodiment, since formula (18) is a scalar equation, it can be solved using a linear matrix equation after collecting sufficient system data. It should be noted that there are various methods for solving linear matrix equations, such as the least squares method, ridge regression, and other linear regression methods.

[0072] In the above embodiments, by combining the Bellman equation with three neural networks, complex nonlinear problems and uncertainties can be handled, making the interception guidance strategy acquisition method more flexible and versatile. Utilizing the solution method of linear matrix equations can significantly improve computational efficiency, especially when dealing with large-scale data, significantly reducing computation time and resource consumption, and lowering the difficulty of solving the Bellman equation. Through a data-driven approach, it can automatically adapt to changes in different scenarios and conditions without requiring tedious manual adjustments to the system model. Through iterative optimization and neural network fitting, the optimal solution can be gradually approximated, improving the accuracy and performance of the interception guidance strategy.

[0073] In some embodiments, before obtaining the first Bellman equation in S11, the interception guidance strategy acquisition method may further include: constructing a nonlinear interception dynamic equation based on the motion model of the pursuer and the motion model of the escaper, and obtaining the game value function of zero-sum game theory.

[0074] In one implementation, both the pursuing motion model and the escaping motion model can be six-degree-of-freedom dynamic models. The pursuing motion model and the escaping motion model can be represented by a general formula. This general formula can be expressed as follows: (20)

[0075]

[0076] Where the subscript i can be either the pursuing party M or the escaping party T, and T i V represents the constant thrust of the pursuing or escaping party. i γi represents the speed of the pursuing or escaping party, ψi represents the track inclination angle of the pursuing or escaping party, αi represents the angle of attack of the pursuing or escaping party, φi represents the sideslip angle of the pursuing or escaping party, mi represents the mass parameter of the pursuing or escaping party, and g represents the local gravitational angular velocity constant.

[0077] L i and D iLet the lift or drag of the pursuing or escaping party be represented respectively, as shown in the following formula (21).

[0078]

[0079] Among them, c Li c is the lift coefficient. Di Where ρ is the drag coefficient, ρ is the atmospheric density, and s is the atmospheric density. i Indicates the reference cross-sectional area.

[0080] Let p This represents the position vector of either the pursuing or escaping party. Let represent the velocity vector of the pursuing or escaping party. Therefore, the interception state space is as follows (22).

[0081]

[0082] Subsequently, the interception guidance control strategy u i It can be expressed as formula (23).

[0083]

[0084] Among them, T i α represents the constant thrust of the pursuing or escaping force. i φ represents the angle of attack of the pursuing or fleeing party. i It indicates the sideslip angle of the pursuing or escaping party.

[0085] The nonlinear interception dynamics equation can be a nonlinear equation, which can be expressed as formula (24).

[0086]

[0087] in, Let f represent the control strategies of the pursuer and the escapee after k iterations. Let f represent functions f(x) and g(x). T Represents the function g T (x), and g M G represents M (x).

[0088]

[0089]

[0090] in, This represents the atmospheric density used by the pursuing party to calculate drag. The atmospheric density used to calculate drag on the escape side. The atmospheric density used to calculate thrust for the escape side. This indicates the atmospheric density used by the pursuing party to calculate drag.

[0091] The "game value function" involved in this application can be equal to the game performance function in zero-sum game theory. The game performance function can be expressed as formula (28).

[0092]

[0093] V(x(t))=J(u M ,u T (29)

[0094] Where V(x(t)) represents the game value function, Q represents the state weight matrix, R M R represents the first control weight coefficient. T α represents the second control weight coefficient, and α represents the discount factor.

[0095] In some embodiments, constructing a nonlinear interception dynamic equation based on the motion model of the pursuer and the motion model of the escaper may include: constructing an initial nonlinear interception dynamic equation based on the motion model of the pursuer and the motion of the escaper; and converting the initial nonlinear interception dynamic equation into a nonlinear interception dynamic equation based on the time derivative of the initial nonlinear interception dynamic equation.

[0096] Here, the initial nonlinear interception dynamic equation can be derived from formulas (20) to (22), and the initial nonlinear interception dynamic equation can be expressed as formula (30) below.

[0097]

[0098] Where f(x) can be derived from formula (25), g T (x) can be derived from formula (26) and g M (x) can be represented by formula (27). to Let each of these terms represent a term in formula (22).

[0099] The initial nonlinear interception dynamics equation is differentiated in time, thereby transforming the initial nonlinear interception dynamics equation (30) into the nonlinear interception dynamics equation (24).

[0100] In the above embodiments, by constructing initial nonlinear interception dynamic equations and transforming them into final nonlinear interception dynamic equations, an important theoretical foundation and mathematical model can be provided for the subsequent design of interception guidance and control strategies. This helps to achieve more accurate and efficient interception missions and improve the performance and reliability of interception guidance strategies.

[0101] In some embodiments, S11, obtaining the first Bellman equation may include: obtaining the second Bellman equation based on the nonlinear interception dynamics equation and the game value function; and transforming the second Bellman equation into the first Bellman equation based on an iterative control strategy.

[0102] In some embodiments, the second Bellman equation can be obtained by differentiating the game value function and substituting the nonlinear interception dynamics equation into the differentiated game value function. The game value function can be a zero-sum game theory game value function.

[0103] The second Bellman equation can be expressed as follows:

[0104]

[0105] The meanings of the parameters in formula (30) are the same as those in the first Bellman equation.

[0106] The iterative control strategy can refer to the change strategy between the k-th interception control strategy and the (k+1)-th interception control strategy. The iterative control strategy can be expressed as the following formula (32).

[0107]

[0108] in, This represents the control input policy of the escape side in the (k+1)th iteration. This represents the control input strategy of the pursuing side in the (k+1)th iteration. Let represent the game value function for the k-th round.

[0109] In some embodiments, the transformation of the second Bellman equation into the first Bellman equation based on an iterative control strategy includes: transforming the second Bellman equation into an intermediate equation based on an iterative control strategy; multiplying both sides of the intermediate equation by an exponential function and integrating to obtain the first Bellman equation.

[0110] In some embodiments, substituting the iterative control strategy into the second Bellman equation yields an intermediate equation. This intermediate equation can be expressed as formula (33).

[0111]

[0112] The "exponential function" involved in the embodiments of this application can be represented as e -α(τ-t) The integral can refer to the integral of the differentiated game value function over the time interval T. Thus, both sides of the intermediate equation are multiplied by e. -α(τ-t) By integrating over time intervals, the second Bellman equation is obtained.

[0113] It should be noted that the first Bellman equation is a key equation within the policy-integral reinforcement learning framework, combining the advantages of nonlinear interception dynamics, game theory, and iterative control policies. Compared to the second Bellman equation, the first Bellman equation focuses more on finding the optimal control policy during the iterative process and describes the relationship between state transitions and expected rewards in integral form.

[0114] In the above embodiments, by transforming the second Bellman equation into the first Bellman equation, an important theoretical basis and mathematical model can be provided for the subsequent design and optimization of control strategies.

[0115] To provide a complete understanding of the interception and guidance strategy acquisition method provided in the embodiments of this application, another aspect of the embodiments of this application provides an alternative interception and guidance strategy acquisition method. Figure 2 This illustration shows a flowchart of another interception guidance strategy acquisition method provided in an embodiment of this application, as shown below. Figure 2 The interception guidance strategy acquisition method may include the following steps.

[0116] S21, Initialize parameters, using two permissive control strategies as the initial control strategies.

[0117] S22, the initial control strategy is input into the system after adding exploratory noise, and the system's control input data and status output data are collected.

[0118] Here, the system can refer to the nonlinear interception dynamics equations constructed as described above.

[0119] S23. The first Bellman equation is solved iteratively using the collected control input data and state output data, and the game value function and the optimal input control strategy of both sides are updated at the same time.

[0120] S24, determine if the convergence condition is met. If it is met, output the optimal input control strategy for the kth iteration. If it is not met, let k = k + 1, use the optimal input control strategy obtained in the kth iteration as the initial control strategy, and return to S22.

[0121] In the above embodiments, an online off-policy integral reinforcement learning algorithm is used to iteratively solve the first Bellman equation of model-free off-policy integral reinforcement learning, thereby obtaining the optimal input control policy that can achieve the Nash equilibrium condition. Thus, compared to traditional guidance control methods, the process of obtaining the interception guidance policy in this application has the advantage of being model-free, and the obtained interception guidance policy is optimal.

[0122] Another aspect of this application provides a simulation experiment to verify the effectiveness of the optimal interception guidance control strategy obtained in the foregoing embodiments of this application.

[0123] The initial parameters for the simulation include: the missile's initial position. initial velocity Initial track inclination Initial track deflection Initial position of the target initial velocity Initial track inclination Initial track deflection And dynamic parameters, missile: mass parameter m M =150kg, lift coefficient c LM =35 rad, drag coefficient c DM =0.74, reference cross-sectional area s M =0.0324m 2 Thrust T M =5880N; Target: mass parameter m T =7500kg, lift coefficient c LT = 4.01 rad, drag coefficient c DT =0.0169, reference cross-sectional area s T =26m 2 Thrust T T = 65000N, local gravitational acceleration constant g = 9.8m / s² 2 Atmospheric density ρ = 0.9096 kg / m³ 3 The performance index function parameters are Q = 100I, R M =5I,R T =I, where I is the identity matrix.

[0124] The output matrix is ​​defined as follows:

[0125] Explore the noise as a complex trigonometric function signal summation. Define the neural network basis functions as follows: φ(x)=x, The initial value of the constant weight matrix of the neural network is

[0126]

[0127] Simulation results are as follows Figures 3 to 6 As shown. Among them, Figure 3 This application illustrates an evaluation method for neural network weight parameters. Convergence plot, Figure 4 This application illustrates a tracking neural network weight parameter provided by an implementation of the present application. Convergence plot. Figure 5 This application provides an embodiment of an escape-side neural network weight parameter. Convergence plot. The horizontal axis represents the number of iterations, and the vertical axis represents the convergence value of the weight parameters. From... Figures 3 to 5 It can be seen that all three neural networks converged in 15 iterations, thus the optimal interception control strategy can be output in the 15th iteration. In this way, the optimal interception control strategy can be output quickly through the policy integral reinforcement learning algorithm.

[0128] Figure 6 This diagram illustrates a three-dimensional schematic of a missile interception target implemented according to an optimal interception control strategy, as provided in an embodiment of this application. Figure 6 As shown, the missile and the target fly based on the optimal interception control strategy, which can make the final position of the missile and the position of the target relatively close, and the distance between the two is within the allowable error range, thus verifying the accuracy of the optimal interception control strategy.

[0129] In this embodiment, firstly, a nonlinear interception dynamics equation is obtained by modeling the relative motion model of the interception scenario. Simultaneously, a differential game model is performed based on this relative motion model, defining a zero-sum game under the interception problem. The theoretically optimal control strategy is then solved using the minimization principle. Secondly, to solve the optimal interception strategy under the condition of unknown system dynamics information, a reinforcement learning framework is introduced. The model-free, policy-independent integral reinforcement learning Bellman equation is derived theoretically, and a corresponding online policy-independent integral reinforcement learning algorithm is proposed. Finally, the system output state information is collected by adding stable control input and exploration noise. Three neural networks are used to fit the game value function, the pursuer's strategy, and the escaper's strategy. The integral reinforcement learning Bellman equation is iteratively solved until the neural network converges, thus obtaining the optimal interception guidance control strategy. Simulation experiments verify the effectiveness of the neural network parameters obtained through the reinforcement learning algorithm and the effectiveness of the optimal interception guidance control strategy.

[0130] Another aspect of this application provides an interception guidance strategy acquisition device. Figure 7 This illustration shows a structural schematic diagram of an interception guidance strategy acquisition device provided in an embodiment of this application, as shown below. Figure 7 As shown, the interception guidance strategy acquisition device 70 may include the following modules.

[0131] The first equation acquisition module 71 is used to acquire the first Bellman equation, which is a policy integral reinforcement learning Bellman equation obtained based on the nonlinear interception dynamics equation, the game value function, and the iterative control strategy.

[0132] The input and output data acquisition module 72 is used to acquire control input data and status output data;

[0133] The control strategy generation module 73 is used to iteratively solve the first Bellman equation based on the control input data and the state output data to obtain the optimal interception guidance control strategy.

[0134] In the above embodiments, a first Bellman equation is obtained. This first Bellman equation is a Bellman equation derived from off-policy integral reinforcement learning based on nonlinear interception dynamics equations, game value functions, and iterative control strategies. Therefore, by using this first Bellman equation, the interception problem can be described and solved more accurately, forming a framework suitable for off-policy integral reinforcement learning. Control input data and state output data are obtained, and based on these data, the first Bellman equation is iteratively solved to obtain the optimal interception guidance control strategy. Thus, using the off-policy integral reinforcement learning method, the optimal control strategy can be found through a data-driven approach without relying on a specific system model, resulting in optimal control input. Simultaneously, the optimal control strategy can better achieve the guided interception task against the maneuvering escapee, and the interception accuracy is high.

[0135] In some embodiments, the control policy generation module 73 may include the following sub-modules.

[0136] The data matrix transformation submodule is used to transform control input data and status output data into a first data matrix and a second data matrix, respectively.

[0137] The fitting submodule is used to fit the game value function to be updated, the pursuer control strategy, and the escaper control strategy in the first Bellman equation based on the evaluation neural network, the pursuer neural network, and the escaper neural network, respectively, to obtain the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy.

[0138] The fitting equation transformation submodule is used to transform the first Bellman equation into a fitting equation based on the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy. The terms in the fitting equation are represented by the Kronen product.

[0139] The linear matrix transformation submodule is used to transform the fitted equation into a linear matrix equation based on the first data matrix, the second data matrix, the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy.

[0140] The linear equation solving submodule is used to solve linear matrix equations to obtain the optimal interception guidance control strategy.

[0141] In some embodiments, the first equation acquisition module 71 may include several sub-modules.

[0142] The second equation yields a submodule used to derive the second Bellman equation based on the nonlinear interception dynamics equation and the game value function.

[0143] The equation transformation submodule is used to transform the second Bellman equation into the first Bellman equation based on an iterative control strategy.

[0144] In some embodiments, the equation transformation submodule may include the following units.

[0145] The intermediate equation transformation unit is used to transform the second Bellman equation into an intermediate equation based on an iterative control strategy.

[0146] The first equation yields a unit, which is used to multiply both sides of the intermediate equation by an exponential function and then integrate to obtain the first Bellman equation.

[0147] In some embodiments, the interception guidance strategy acquisition device 70 may further include the following modules.

[0148] The nonlinear equation construction module is used to construct nonlinear interception dynamic equations based on the motion models of the pursuer and the escaper.

[0149] The function retrieval module is used to retrieve the game value function in zero-sum game theory.

[0150] In some embodiments, the nonlinear equation construction module may include the following sub-modules.

[0151] The initial equation construction submodule is used to construct the initial nonlinear interception dynamic equations based on the motion models of the pursuer and the escaper.

[0152] The nonlinear equation transformation module is used to transform the initial nonlinear interception dynamics equation into a nonlinear interception dynamics equation based on the time derivative of the initial nonlinear interception dynamics equation.

[0153] In some embodiments, the first Bellman equation is:

[0154]

[0155] Among them, u M and u T Indicates the initial input control strategy. and The optimal input control strategy for the k-th iteration is... and Let x be the optimal input control strategy for the (k+1)th iteration, and let x represent the output state information.

[0156] It should be understood that the specific features, operations, and details described herein with respect to the methods of this application can also be similarly applied to the apparatus and system of this application, or vice versa. Furthermore, each step of the methods of this application described above can be performed by a corresponding component or unit of the apparatus or system of this application.

[0157] It should be understood that the various modules / units of the device of this application can be implemented wholly or partially through software, hardware, firmware, or a combination thereof. Each module / unit can be embedded in the processor of the electronic device in hardware or firmware form or independent of the processor, or it can be stored in the memory of the electronic device in software form for the processor to call to execute the operation of each module / unit. Each module / unit can be implemented as an independent component or module, or two or more modules / units can be implemented as a single component or module.

[0158] In another aspect, this application provides an electronic device. Figure 8 This illustration shows a structural diagram of an electronic device provided in an embodiment of this application, such as... Figure 8 As shown, the electronic device 80 includes a processor 81 and a memory 82 storing computer program instructions. The processor 81 executes the computer program instructions to implement the steps of the aforementioned method for implementing the changing scene. This electronic device 80 can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities.

[0159] In one embodiment, the electronic device 80 may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the electronic device 80 can be used to provide necessary computing, processing, and / or control capabilities. The memory of the electronic device 80 may include non-volatile storage media and internal memory. The non-volatile storage media may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface and communication interface of the electronic device 80 can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the interception guidance strategy acquisition method of this application.

[0160] This application provides a computer-readable storage medium storing an interception and guidance strategy program. When the vehicle-mounted conference management program is executed by a processor, it implements the steps of the interception and guidance strategy acquisition method provided in any of the above embodiments.

[0161] Those skilled in the art will understand that the method steps of this application can be performed by a computer program instructing related hardware, such as electronic device 80 or a processor. The computer program can be stored in a non-transitory computer-readable storage medium, and its execution causes the steps of this application to be performed. Depending on the context, any reference herein to memory, storage, or other media may include non-volatile or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0162] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for obtaining interception guidance strategies, characterized in that, include: The first Bellman equation is obtained, which is a policy-integral reinforcement learning Bellman equation based on the nonlinear interception dynamics equation, the game value function, and the iterative control strategy; the first Bellman equation is: in, and Indicates the initial input control strategy. and The optimal input control strategy for the k-th iteration is... and Let x represent the (k+1)th optimal input control strategy, and Q represent the output state information. c This represents the augmented state weight coefficient. This represents the first control weight coefficient. This represents the second control weight coefficient. V represents the discount factor, T represents the time interval, and V represents the discount factor. k (x(t)) represents the value function of the k-th game, and τ represents the integration time; Acquire control input data and status output data; Based on the control input data and the state output data, the first Bellman equation is solved iteratively to obtain the optimal interception guidance control strategy. The step of iteratively solving the first Bellman equation based on the control input data and the state output data to obtain the optimal interception guidance control strategy includes: The control input data and the status output data are respectively converted into a first data matrix and a second data matrix; Based on the evaluation neural network, the pursuer neural network, and the escaper neural network, the game value function to be updated, the pursuer control strategy, and the escaper control strategy in the first Bellman equation are fitted respectively, so as to obtain the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy. Based on the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy, the first Bellman equation is converted into a fitted equation, and each term in the fitted equation is represented by the Krone product. Based on the first data matrix, the second data matrix, the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy, the fitted equation is transformed into a linear matrix equation. Solving the linear matrix equation yields the optimal interception guidance control strategy.

2. The interception guidance strategy acquisition method according to claim 1, characterized in that, The process of obtaining the first Bellman equation includes: Based on the nonlinear interception dynamics equation and the game value function, the second Bellman equation is obtained; Based on an iterative control strategy, the second Bellman equation is transformed into the first Bellman equation.

3. The interception guidance strategy acquisition method according to claim 2, characterized in that, The method of transforming the second Bellman equation into the first Bellman equation based on the iterative control strategy includes: Based on the iterative control strategy, the second Bellman equation is transformed into an intermediate equation; Multiplying both sides of the intermediate equation by an exponential function and integrating, we obtain the first Bellman equation.

4. The interception guidance strategy acquisition method according to claim 1, characterized in that, Prior to obtaining the first Bellman equation, the method further includes: Based on the motion models of the pursuer and the escaper, a nonlinear interception dynamic equation is constructed. Obtain the game value function in zero-sum game theory.

5. The interception guidance strategy acquisition method according to claim 4, characterized in that, The nonlinear interception dynamic equations, constructed based on the motion models of the pursuer and the escaper, include: Based on the motion models of the pursuer and the escaper, an initial nonlinear interception dynamic equation is constructed. Based on the time derivative of the initial nonlinear interception dynamics equation, the initial nonlinear interception dynamics equation is transformed into a nonlinear interception dynamics equation.

6. An interception guidance strategy acquisition device, characterized in that, include: The first equation acquisition module is used to acquire the first Bellman equation, which is a policy-integral reinforcement learning Bellman equation obtained based on the nonlinear interception dynamics equation, the game value function, and the iterative control strategy; the first Bellman equation is: in, and Indicates the initial input control strategy. and The optimal input control strategy for the k-th iteration is... and Let x represent the (k+1)th optimal input control strategy, and Q represent the output state information. c This represents the augmented state weight coefficient. This represents the first control weight coefficient. This represents the second control weight coefficient. V represents the discount factor, T represents the time interval, and V represents the discount factor. k (x(t)) represents the value function of the k-th game, and τ represents the integration time; The input and output data acquisition module is used to acquire control input data and status output data; The control strategy generation module is used to iteratively solve the first Bellman equation based on the control input data and the state output data to obtain the optimal interception guidance control strategy. The control strategy generation module includes: a data matrix conversion submodule, used to convert the control input data and the state output data into a first data matrix and a second data matrix, respectively; The fitting submodule is used to fit the game value function to be updated, the pursuer control strategy, and the escaper control strategy in the first Bellman equation based on the evaluation neural network, the pursuer neural network, and the escaper neural network, respectively, to obtain the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy. The fitting equation transformation submodule is used to transform the first Bellman equation into a fitting equation based on the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy. The terms in the fitting equation are represented by the Kronen product. The linear matrix transformation submodule is used to convert the fitted equation into a linear matrix equation based on the first data matrix, the second data matrix, the fitted game value function to be updated, the fitted pursuer control strategy, and the fitted escaper control strategy. The linear equation solving submodule is used to solve the linear matrix equation to obtain the optimal interception guidance control strategy.

7. An electronic device, characterized in that, The electronic device includes a processor and a memory storing computer program instructions, wherein when the computer program instructions are executed by the processor, they implement the steps of the interception guidance strategy acquisition method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer-executable program, wherein when the computer-executable program is executed by a processor, it implements the steps of the interception guidance strategy acquisition method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Soft landing adaptive proportional guidance method based on reinforcement learning

    CN117826585A

  • Method for determining control strategy of fixed-wing unmanned aerial vehicle based on reinforcement learning

    CN117826860A