Incomplete information constraint uncertain game strategy optimization method and system

By constructing a nonlinear uncertain game process model and introducing constrained robust optimization and event-triggered secure robust reinforcement learning techniques, the uncertain game problem with incomplete information constraints in a highly adversarial scenario is solved, achieving optimization of game strategy and resource conservation in the process of missile interception of highly maneuverable targets.

CN121526360APending Publication Date: 2026-02-13UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511381534.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In high-intensity adversarial scenarios, existing technologies struggle to effectively address uncertain game problems constrained by incomplete information. This is particularly true in missile interception of highly maneuverable targets, where the lack of transparency between the two sides and the complex nonlinear dynamic processes make it difficult to establish accurate mathematical models. Consequently, existing reinforcement learning algorithms cannot directly solve the Hamilton-Jacobi-Isax equations and thus cannot obtain game strategies that meet the needs of practical applications.

Method used

A nonlinear uncertain differential game process model is constructed, and the ideas of constraint robust optimization and event-triggered safe robust reinforcement learning are introduced. Through a non-quadratic energy function and a control barrier function, it is transformed into a nonlinear nominal system constraint robust optimization problem. The optimal performance index function is approximated online through a neural network, and the Hamilton-Jacobi-Bellman equation is solved approximately to obtain the player's optimal game strategy.

Benefits of technology

It effectively overcomes the challenges of unknown opponent strategies and unknown local dynamics in game strategy optimization, improves the optimal performance of game strategies, balances robustness and economy, and reduces computational resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121526360A_ABST
    Figure CN121526360A_ABST
Patent Text Reader

Abstract

The invention provides an incomplete information constraint uncertain game strategy optimization method and system, and the method comprises the steps: building a nonlinear uncertain differential game process model according to the incomplete information characteristics of a nonlinear uncertain game dynamic process in a strong confrontation scene, and enabling the strategy amplitudes and game states of two parties to be constrained; taking an opponent game strategy and an unknown uncertain item as a disturbance item, introducing a non-quadratic form energy function, a control barrier function and a disturbance function upper bound on the basis of describing a nonlinear disturbance system of a nonlinear uncertain game dynamic process, and constructing a performance index function; converting a safety robust game problem of a nonlinear uncertain dynamic game process into a nonlinear nominal system constraint robust optimization problem; and introducing an event-triggered security robust reinforcement learning technology, and completing approximate solution of an HJB equation by evaluating an online approximation optimal performance index function of a neural network to obtain an own optimal game strategy. According to the method, the uncertain game strategy can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of game strategy optimization, in particular to a method and system for strategy optimization of constrained uncertain game with incomplete information. BACKGROUND

[0002] Game theory (also known as countermeasure theory and competition theory) mainly studies how the interest subjects with antagonistic and competitive relationship make interactive decisions, how to be driven by their own interests, and how to be restricted by each other to achieve system "equilibrium". In competition, whether there is the most reasonable action plan and how to find the plan is the main research content of game theory. Differential game problem is to use differential equation to deal with continuous antagonistic dynamic game problem. As a typical application scenario, pursuit-evasion differential game is a pursuit-escape scenario composed of game parties or multiple parties, and the research object involves missiles, aircraft, robots, ships, etc. Due to its strong application background, especially in the problem of missile interception guidance, pursuit-evasion differential game has attracted widespread attention from scholars at home and abroad in recent years.

[0003] Considering that qualitative differential game requires to judge the result after the game ends, quantitative differential game becomes the research focus. Generally, the quantitative differential game problem is evaluated by designing a performance index to judge the differential game result, that is, by introducing a corresponding performance index function, according to the maximum value principle and multi-agent system theory, the cooperative pursuit-evasion differential game guidance solving problem is converted into a Hamilton-Jacobi-Isaac (HJI) equation solving problem, and then the solution of the equation is the optimal strategy of the game participants. However, the HJI equation is a second-order nonlinear partial differential equation, which is very difficult to solve analytically, and depends on the dynamic game process model information. Therefore, the numerical solution method of differential game problem has gradually become the mainstream. Among all the numerical solution methods, the strategy iteration and value function approximation method are studied most. The current mainstream research methods are mostly focused on reinforcement learning technology, and require that the information of both parties of the game is open and transparent.

[0004] However, considering the actual strong confrontation scene (such as missile interception), the information of the attack and defense confrontation parties is not transparent (that is, both parties do not know the strategy and intention of the other party), and the complex nonlinear game dynamic process is difficult to establish an accurate mathematical model, the differential game problem in the strong confrontation scene is essentially a kind of incomplete information game problem. Because the game information (such as strategy and intention) of the opponent cannot be directly obtained and the calculation resources of the missile are limited, the game strategy that meets the actual application requirements cannot be obtained by directly solving the HJI equation by applying the existing mature reinforcement learning algorithm. In addition, in some actual scenes, the strategy selection of the game parties is limited and the game dynamics must also evolve within a given range. For example, in the process of intercepting a high-maneuvering target by a missile, due to the physical constraints of the driving mechanism, the interception strategy of the interceptor and the maneuvering defense strategy of the target missile are both amplitude-limited, and the interception guidance process variables (which can be composed of missile-target line-of-sight distance, approach rate, line-of-sight angle, interception angle, etc.) must be within a given range to ensure that the maneuvering target is not lost. Therefore, the differential game process in the strong confrontation scene is essentially a kind of incomplete information constrained uncertain game problem. According to the literature research, the incomplete information constrained uncertain game problem remains open.

[0005] Although the event-triggered safe robust reinforcement learning technology has been deeply studied and successfully applied to solve the game strategy optimization in the strong confrontation scene, how to introduce the event-triggered safe robust reinforcement learning technology and develop a method suitable for the incomplete information constrained uncertain game strategy optimization in the strong confrontation scene has important application value. SUMMARY

[0006] In order to solve the technical problems existing in the prior art, the present application provides an incomplete information constrained uncertain game strategy optimization method and system, and the technical scheme is as follows:

[0007] On the one hand, an incomplete information constrained uncertain game strategy optimization method is provided, which comprises the following steps:

[0008] S1, according to the incomplete information characteristics of the nonlinear uncertain game dynamic process in the strong confrontation scene, a nonlinear uncertain differential game process model is constructed, and the strategy amplitude and game state of the two parties of the nonlinear uncertain game dynamic process are constrained;

[0009] S2, regarding the opponent's game strategy and unknown uncertain term as a disturbance term, by introducing the idea of constraint robust optimization, on the basis of describing the nonlinear disturbance system of the nonlinear uncertain game dynamic process, introducing a non-quadratic energy function, a control barrier function and an upper bound of the disturbance function, a suitable performance index function is constructed, and the safe robust game problem of the nonlinear uncertain dynamic game process is converted into a corresponding nonlinear nominal system constraint robust optimization problem;

[0010] S3, for the nonlinear nominal system constraint robust optimization problem, an event-triggered safety robust reinforcement learning technology is introduced, an optimal performance index function is approximated online by evaluating a neural network, and approximate solution of a corresponding Hamilton-Jacobi-Bellman (HJB) equation is completed, so that an optimal game strategy of the self is obtained.

[0011] In another aspect, an incomplete information constraint uncertain game strategy optimization system is provided, and the system comprises:

[0012] A first construction module is configured to construct a nonlinear uncertain differential game process model according to the incomplete information characteristics of a nonlinear uncertain game dynamic process in a strong confrontation scenario, wherein the strategy amplitude of both parties and the game state of the nonlinear uncertain game dynamic process are constrained.

[0013] A second construction module is configured to regard the game strategy of an opponent and an unknown uncertain term as a disturbance term, introduce a constraint robust optimization idea, introduce a non-quadratic energy function, a control barrier function and an upper bound of a disturbance function on the basis of describing a nonlinear disturbance system of the nonlinear uncertain game dynamic process, construct a suitable performance index function, and convert a safety robust game problem of the nonlinear uncertain dynamic game process into a corresponding nonlinear nominal system constraint robust optimization problem.

[0014] A solution module is configured to, for the nonlinear nominal system constraint robust optimization problem, introduce an event-triggered safety robust reinforcement learning technology, approximate an optimal performance index function online by evaluating a neural network, complete approximate solution of a corresponding Hamilton-Jacobi-Bellman (HJB) equation, and obtain an optimal game strategy of the self.

[0015] In another aspect, an electronic device is provided, and the electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, the at least one instruction is loaded and executed by the processor to implement the above-described incomplete information constraint uncertain game strategy optimization method.

[0016] In another aspect, a computer readable storage medium is provided, and the storage medium stores at least one instruction, the at least one instruction is loaded and executed by a processor to implement the above-described incomplete information constraint uncertain game strategy optimization method.

[0017] The technical scheme provided by the present application has at least the following beneficial effects:

[0018] 1) The incomplete information constraint uncertain game strategy optimization method provided by the present application fully considers the incomplete information characteristics of a strong confrontation game, takes into account the unknown situation of a local game dynamic, and effectively overcomes the ideal conditions that a game strategy optimization method based on a differential game requires an optimal strategy of an opponent and an accurate known game dynamic model.

[0019] 2) The application provides an incomplete information constraint uncertain game strategy optimization method, which converts a safety robust game problem of a nonlinear uncertain dynamic game process into a corresponding nonlinear nominal system robust constraint optimization problem by introducing a constraint robust optimization thought, effectively processes local uncertain game dynamics, strategy amplitude constraints, state constraints and influences of unknown opponent strategies on game strategy optimization, and finally realizes incomplete information constraint uncertain game strategy optimization in a strong confrontation scene.

[0020] 3) The application provides an incomplete information constraint uncertain game strategy optimization method, which introduces an event-triggered safety robust reinforcement learning technology, improves the optimal performance of the game strategy, and takes into account the robust performance, safety performance and economy (mainly the consumption of calculation and communication resources of the algorithm). BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 It is a flow chart of an incomplete information constraint uncertain game strategy optimization method provided by the embodiment of the application.

[0023] Figure 2 It is a schematic diagram of relative plane relationship of a one-to-one missile interception provided by the embodiment of the application.

[0024] Figure 3 It is a schematic diagram of a missile-target interception flight trajectory provided by the embodiment of the application.

[0025] Figure 4 It is a schematic diagram of a missile-target line-of-sight distance provided by the embodiment of the application.

[0026] Figure 5 It is an interception missile lateral acceleration evolution curve diagram provided by the embodiment of the application.

[0027] Figure 6 It is a variation curve diagram of an event-triggered time provided by the embodiment of the application.

[0028] Figure 7 It is a system block diagram of an incomplete information constraint uncertain game strategy optimization system provided by the embodiment of the application.

[0029] Figure 8 It is a structural schematic diagram of an electronic device provided by the embodiment of the application. DETAILED DESCRIPTION

[0030] In order to make the technical problems, technical solutions and advantages to be solved by the present application clearer, the following will be described in detail in combination with the drawings and specific embodiments.

[0031] The embodiment of the present application provides an incomplete information constraint uncertain game strategy optimization method based on event-triggered safe robust reinforcement learning in a strong confrontation scene, which is different from the existing game strategy optimization method based on differential game, and is used for a kind of nonlinear uncertain game dynamic process, according to the incomplete information characteristics of nonlinear uncertain game dynamic process in a strong confrontation scene, a nonlinear uncertain differential game process model is constructed, both strategy amplitude and game state of the nonlinear uncertain game dynamic process are constrained;The strategy of the opponent and the unknown uncertain term are regarded as disturbance term, by introducing the idea of constraint robust optimization, on the basis of describing the nonlinear disturbance system of the nonlinear uncertain game dynamic process, a non-quadratic energy function, a control barrier function and an upper bound of the disturbance function are introduced, a suitable performance index function is constructed, and the safe robust game problem of the nonlinear uncertain dynamic game process is converted into the constraint robust optimization problem of the corresponding nonlinear nominal system;Considering the scarcity of computing resources in actual application scenarios, for the constraint robust optimization problem of the nonlinear nominal system, the event-triggered safe robust reinforcement learning technology is introduced, the optimal performance index function is approximated online by evaluating the neural network, and the approximate solution of the corresponding Hamilton-Jacobi-Bellman (HJB) equation is completed, to obtain the optimal game strategy of the self side, compared with the game strategy optimization method based on differential game, the introduction of constraint robust optimization idea and event-triggered safe robust reinforcement learning technology ensures that the constructed game strategy is more flexible and robust, not only overcomes the influence of local uncertain game dynamic, amplitude constraint, state constraint and unknown strategy of the opponent on game strategy optimization, but also improves the optimal performance of the game strategy while considering the robust performance, safety performance and economy (mainly the consumption of computing resources of the algorithm).

[0032] The incomplete information constraint uncertain game strategy optimization method provided by the embodiment of the present application can be implemented by an electronic device, which can be a terminal or a server, Figure 1 As shown in the figure, the method flowchart can include the following steps:

[0033] S1, according to the incomplete information characteristics of nonlinear uncertain game dynamic process in a strong confrontation scene (such as missile interception / penetration guidance), a nonlinear uncertain differential game process model is constructed, both strategy amplitude and game state of the nonlinear uncertain game dynamic process are constrained;

[0034] Optionally, the nonlinear uncertain differential game process model includes a nonlinear uncertain game dynamic equation, strategy amplitude constraints and game state constraints of the nonlinear uncertain game dynamic equation.

[0035] The nonlinear uncertain game dynamic equation is as follows:

[0036]

[0037] wherein, is a game state, and are game strategies of both sides, is a game strategy of the self side, is a game strategy of the opponent side, and a function is continuous and differentiable with respect to x(t), satisfies f(0) = 0, and depicts a local nonlinear dynamic game behavior, and are nonlinear matrix functions with respect to the game state x(t), wherein and Δ(x(t), u(t), v(t)) depicts an unmodeled local game dynamic behavior, satisfies Δ(0, u(t), v(t)) = 0, and is an initial game state, and a balance state (x e , u e , v e ) of the nonlinear uncertain game equation (1) satisfies 0 = f(x e ) + B(x e )u e + H(x e )v e + Δ(x e , u e , v e ). It is assumed that (0, 0, 0) is a balance state of the nonlinear uncertain game process (1).

[0038] The nonlinear uncertain game equation (1) of the embodiment of the application can be used for modeling a guidance process of a missile intercepting a high-maneuvering target, wherein x(t) is a variable of the intercepting guidance process (which can be composed of a line-of-sight distance between the missile and the target, an approaching speed, a line-of-sight angle, an intercepting angle, etc.), u(t) and v(t) are respectively an intercepting strategy of the intercepting missile and a maneuvering and penetration strategy of the target missile, and Δ(x(t), u(t), v(t)) is an unmodeled guidance dynamic behavior. In the game intercepting guidance process, the intercepting missile and the target missile can only exert limited axial accelerations, that is, the intercepting strategy and the penetration strategy are both limited in amplitude.

[0039] The game strategies u(t) and v(t) of the nonlinear uncertain game equation (1) satisfy the following amplitude constraints:

[0040]

[0041] where, and are the amplitude bounds of the game strategy;

[0042] The game state x(t) satisfies the following constraints (to ensure the interception effect, the interception guidance state x(t) must evolve within a given range, that is, the state x(t) of the game equation (1) is limited):

[0043]

[0044] where, and are the upper and lower bounds of the game state variable x i (t).

[0045] S2, the opponent game strategy and unknown uncertainty term are regarded as disturbance terms, by introducing the constraint robust optimization idea, on the basis of describing the nonlinear disturbance system of the nonlinear uncertain game dynamic process, the non-quadratic energy function, the control barrier function and the upper bound of the disturbance function are introduced, the appropriate performance index function is constructed, and the safety robust game problem of the nonlinear uncertain dynamic game process is converted into the corresponding nonlinear nominal system constraint robust optimization problem;

[0046] Optionally, the S2 specifically comprises:

[0047] Due to the existence of the game strategy amplitude constraints (2) and (3), the game state constraint (4) and the unmodeled local game dynamic behavior Δ(x(t),u(t),v(t)), the design of the own game strategy u(t) of the nonlinear dynamic game equation (1) not only needs to resist the influence of the opponent game strategy, but also needs to overcome the adverse factors of the unmodeled game dynamic behavior, while ensuring that the above various safety constraints are not violated, therefore, the design of the own game strategy u(t) of the nonlinear dynamic game equation (1) is a kind of safety robust game problem, which is essentially a kind of constrained uncertain optimization problem, by using the amplitude constraints (2) and (3) of the game strategies u(t) and v(t), regarding H(x(t))v(t) and Δ(x(t),u(t),v(t)) as bounded disturbances, the nonlinear dynamic game equation (1) is modeled as a nonlinear disturbance system:

[0048]

[0049] where, is the nonlinear disturbance function of the game process and satisfies ||d(x(t),u(t),v(t))||≤d M (x(t)),

[0050] For the nonlinear perturbed system (5), the performance index function of the optimal strategy u(t) is designed under the game strategy amplitude constraints (2), (3) and the game state constraints (4):

[0051]

[0052] where, is the weight matrix of the game state, d M (x(t)) is the upper bound of the perturbation function d(x(t),u(t),v(t)), and the purpose of introducing d M (x(t)) is to offset the influence of the nonlinear perturbation function d(x(t),u(t),v(t)) on the game strategy performance, and δ>0 is a design parameter that will affect the convergence of the guidance system, and U(u(t)) is a non-quadratic energy function about the game strategy u(t):

[0053]

[0054] Here r j , j∈M is the weight coefficient of the strategy u(t), and CB(x) is the control barrier function, which is expressed as Here s(x) is a continuous differentiable function related to the safety boundary of the game state x(t), κ>0 is the adjustment parameter of the barrier function CB(x) when the game state x(t) approaches its safety boundary, and the function CB(x) is used as a risk penalty term for the safety constraint of the game dynamics in the performance index function;

[0055] Because the upper bound d M (x(t)) of the perturbation function d(x(t),u(t),v(t)) has been introduced in the performance index function (6) to offset its influence on the optimization of the game strategy u(t), the optimization of the game strategy u(t) in the nonlinear uncertain dynamic game equation (1) can ignore the influence of the perturbation function d(x(t),u(t),v(t)) in the nonlinear perturbed system (5) and only consider the nonlinear nominal system In addition, the optimization of the game strategy u(t) must satisfy the constraints (2) and the game state constraints (4), which are essentially the constraints of the nonlinear nominal system The game strategy u(t) optimization problem of the performance index function (6) is a constrained robust optimization problem, therefore, with the help of the concept of constrained robust optimization, according to the nonlinear perturbed system (5) and the performance index function (6), the safety robust game problem of the nonlinear uncertain dynamic game equation (1) is transformed into a constrained robust optimization problem of the following nonlinear nominal system:

[0056]

[0057] S3, for the nonlinear nominal system constraint robust optimization problem, an event-triggered safety robust reinforcement learning technology is introduced, the optimal performance index function is approximated online by evaluating the neural network, and approximate solution of the corresponding Hamilton-Jacobi-Bellman (HJB) equation is completed, so that the optimal game strategy of the self side is obtained.

[0058] Optionally, the S3 specifically comprises:

[0059] For the nonlinear nominal system constraint robust optimization problem (8), the Hamilton function is defined as follows:

[0060]

[0061] Wherein, is the gradient of the performance index function J, and the following optimal value function is introduced:

[0062]

[0063] According to the Pontryagin extremum principle, the partial derivative of the performance index function should satisfy the HJB equation:

[0064]

[0065] Assuming that the solution of the HJB equation (11) exists and is unique, the optimal game strategy of the nonlinear nominal system constraint robust optimization problem (8) is:

[0066] u * (x)=-u M tanh(η(x)) (12)

[0067] Wherein,

[0068]

[0069] and

[0070] The HJB equation (11) is rewritten as:

[0071]

[0072] Wherein, and

[0073] ​The optimal game strategy (12) depends on the analytical solution of the HJB equation (13). However, since the HJB equation (13) is a typical nonlinear partial differential equation, it is difficult to obtain its analytical solution directly. Therefore, an evaluation neural network based on reinforcement learning is designed to approximate the optimal performance index function J online by combining the adaptive dynamic programming algorithm. * (x), and complete the approximate solution of HJB equation (13) to obtain the optimal game strategy of our side.

[0074] Optionally, the specific expression of the evaluation neural network is:

[0075] J * (x)=(w * ) T θ(x)+ρ(x) (14)

[0076] in, These are the ideal weights for evaluating neural networks. It evaluates the basis functions of a neural network, where N is the number of neurons and ρ(x) represents the approximation error;

[0077] Accordingly, J * The gradient of (x) with respect to the game state x is in the form of:

[0078]

[0079] in, Let θ(x) be the partial derivative of the basis function with respect to the game state x. The optimal game strategy (12) is expressed as:

[0080] u * (x)=-u M tanh(η(x,w * (16)

[0081] in,

[0082] To effectively reduce the resource consumption of game strategies, an event-triggered technique is introduced. Therefore, a set of monotonically increasing trigger sequences is defined. in, Indicates the time when the event is triggered and the event dwell time. Let x(t) q ) represents the trigger time t q Given the dynamic nature of the game process, the corresponding triggering error is:

[0083]

[0084] Using the trigger time t q The game state x(t) q The event triggering form of game strategy (16) is as follows:

[0085] u * (x(t q ))=-u M tanh(η(x(t q ),w * (18)

[0086] in, And the event triggering condition is:

[0087]

[0088] Where ω>0 is the trigger coefficient;

[0089] The corresponding event-triggered Hamiltonian function is:

[0090]

[0091] Because the ideal weights w for evaluating neural networks * Since it is unknown, in order to realize the event-triggered game strategy (18), the estimated weights are used instead of the ideal weights, and the following approximate expression is constructed:

[0092]

[0093] in,

[0094] The approximate event-triggered Hamiltonian function is:

[0095]

[0096] Combining the gradient descent algorithm, the weight updates of the neural network are evaluated as follows:

[0097]

[0098] in, μ > 0 is the given learning rate.

[0099] An example of an embodiment of the present invention is given below with reference to the accompanying drawings:

[0100] Consider as Figure 2 The scenario shown is a one-to-one missile interception game. Figure 2 In the diagram, M and T represent the interceptor missile and the maneuvering target, respectively; V M and V T Let α and β represent the velocities of the interceptor missile and the maneuvering target, respectively, and θ be the line-of-sight angle. The corresponding line-of-sight angular rate is: Intercepting missiles and maneuvering targets' maneuverability is achieved by applying control u in the normal direction perpendicular to their respective velocity directions. Mand v T Achieved. The relative distance and relative speed between the interceptor missile and the maneuvering target are r and r, respectively.

[0101] Choose the system state variable x = [θ σ] T The state equation of the interception guidance system is:

[0102]

[0103] in, According to the parallel guidance principle, as long as the line-of-sight angular velocity of all missiles approaches zero and the relative velocity between the missile and the target is less than zero, that is... This ensures that the missile can hit the target.

[0104] In the terminal intercept guidance phase, assuming the missile and target have constant velocities, let V... M =600m / s, V T =400m / s, the initial position coordinates of the missile and the target are respectively (x M ,y M ) = (0,0) and (x T ,y T ) = (2500, 0). The target is v T The missile maneuvers with a lateral velocity of 50*sin(t). Considering that in actual interception and guidance processes, missiles are often subject to physical constraints such as overload saturation, we assume the missile's input constraint is |u... M ≤500m / s 2 .

[0105] Assume the missile's initial trajectory angles are α0 = 30°, β0 = 60°, and the initial line-of-sight angle θ(0) = 0°. In the evaluation neural network, the basis function is... The initial weights w(0) are set to [50, 50, 50]. T ;choose R = 0.04 and a = 10, learning rate α = 0.07, where Q = 200I². And δ = 1. Let ω = 0.05. The event-triggered robust game-theoretic interception guidance method is applied to the guidance system, and the simulation results are as follows: Figures 3-5 As shown:

[0106] Figure 3 The trajectory curve of the missile intercepting the maneuvering target was depicted, which shows that the missile successfully intercepted the maneuvering target within an acceptable miss range. Figure 4 The curve showing the change in the line-of-sight distance between the missile and the target is presented. It can be seen that the relative distance gradually decreases from 2500m to within the allowable miss range. This also shows that the missile is always approaching the target during the interception guidance process until it finally hits the target successfully.Figure 5 The trajectory of the lateral acceleration a M of the missile is shown.

[0107] In order to intuitively show the advantages of the embodiments of the present application, the simulation experimental results of the period sampling safe robust reinforcement learning game guidance method are also shown in Figures 3-5 .

[0108] In addition, Figure 6 The change curve of the event triggering moment is shown, and it can be seen that the minimum triggering interval is 0.01s, so the Zeno behavior will not occur.

[0109] Table 1 shows the guidance performance of the missile intercepting a maneuvering target under the action of the event-triggered and period-sampling safe robust reinforcement learning guidance method, and it can be seen that the guidance performance based on the event-triggered mechanism is equivalent to that based on the period-sampling method. However, the event-triggered sampling scheme only needs to transmit 260 times of data, while the period-sampling guidance method needs to transmit 800 times of data. Through calculation, it can be known that the number of data transmission is saved by 67.5% under the event-triggered sampling scheme. Therefore, compared with the period-sampling guidance method, the event-triggered optimal interception guidance method can save the communication cost of the system while ensuring the guidance performance.

[0110] Table 1.

[0111] Event trigger Periodic sampling Off-target amount r(m) 1.000 0.428 Interception time t(s) 8.00 8.00 Data transmission times t r (times) 260 800

[0112] As Figure 7 shown, the embodiments of the present application further provide an incomplete information constrained uncertain game strategy optimization system, which comprises:

[0113] The first construction module 710 is configured to construct a nonlinear uncertain differential game process model according to the incomplete information characteristics of a nonlinear uncertain game dynamic process in a strong confrontation scene, and the amplitude of the strategy of both parties and the game state of the nonlinear uncertain game dynamic process are constrained.

[0114] The second construction module 720 is configured to regard the opponent game strategy and the unknown uncertain term as a disturbance term, introduce a constraint robust optimization idea, introduce a non-quadratic energy function, a control barrier function and an upper bound of a disturbance function on the basis of describing a nonlinear disturbance system of the nonlinear uncertain game dynamic process, construct a suitable performance index function, and convert a safe robust game problem of the nonlinear uncertain dynamic game process into a corresponding nonlinear nominal system constraint robust optimization problem.

[0115] The solving module 730 is configured to introduce an event-triggered safe robust reinforcement learning technique for the nonlinear nominal system constraint robust optimization problem, to evaluate a neural network to online approximate an optimal performance index function, and to complete approximate solving of a corresponding Hamilton-Jacobi-Bellman (HJB) equation to obtain an optimal game strategy of the ego.

[0116] Optionally, the nonlinear uncertain differential game process model comprises a nonlinear uncertain game dynamic equation, a strategy amplitude constraint of the nonlinear uncertain game dynamic equation, and a game state constraint.

[0117] The nonlinear uncertain game dynamic equation is as follows:

[0118]

[0119] wherein, is a game state, and are game strategies of two parties, is an ego game strategy, is an opponent game strategy, and a function is continuous and differentiable with respect to x(t), satisfies f(0) = 0, and describes a local nonlinear dynamic game behavior, and are nonlinear matrix functions with respect to the game state x(t), where and Δ(x(t), u(t), v(t)) describes an unmodeled local game dynamic behavior and satisfies Δ(0, u(t), v(t)) = 0 and is an initial game state, and a balance state (x e , u e , v e ) of the nonlinear uncertain game equation (1) satisfies 0 = f(x e ) + B(x e )u e + H(x e )v e + Δ(x e , u e , v e ), and it is assumed that (0, 0, 0) is the balance state of the nonlinear uncertain game process (1).

[0120] The game strategies u(t) and v(t) of the nonlinear uncertain game equation (1) satisfy the following amplitude constraints:

[0121]

[0122] wherein, and is the amplitude bound of the game strategy;

[0123] The game state x(t) satisfies the following constraints:

[0124]

[0125] wherein, and are the upper and lower bounds of the game state variable x i (t).

[0126] Optionally, the S2 specifically comprises:

[0127] Due to the game strategy amplitude constraints (2) and (3), the game state constraint (4), and the existence of the unmodeled local game dynamic behavior Δ(x(t),u(t),v(t)), the design of the own game strategy u(t) of the nonlinear dynamic game equation (1) not only needs to counteract the influence of the opponent game strategy, but also needs to overcome the adverse factors of the unmodeled game dynamic behavior, while ensuring that the above various safety constraints are not violated. Therefore, the design of the own game strategy u(t) of the nonlinear dynamic game equation (1) is a kind of safe robust game problem, and its essence is a kind of constrained uncertain optimization problem. By using the amplitude constraints (2) and (3) of the game strategies u(t) and v(t), the H(x(t))v(t) and Δ(x(t),u(t),v(t)) are regarded as bounded disturbances, and the nonlinear dynamic game equation (1) is modeled as a nonlinear disturbance system:

[0128]

[0129] wherein, is a nonlinear disturbance function of the game process and satisfies ||d(x(t),u(t),v(t))||≤d M (x(t)),

[0130] For the nonlinear disturbance system (5), under the game strategy amplitude constraints (2), (3) and the game state constraint (4), the performance index function of the optimization strategy u(t) is designed as:

[0131]

[0132] wherein, is the weight matrix of the game state, d M (x(t)) is the upper bound of the disturbance function d(x(t),u(t),v(t)), and d MThe purpose of the term (x(t)) is to offset the influence of the nonlinear disturbance function d(x(t), u(t), v(t)) on the performance of the game strategy, and δ > 0 is a design parameter that will affect the convergence of the guidance system, and U(u(t)) is a non-quadratic energy function related to the game strategy u(t):

[0133]

[0134] Here r j , j∈M is the weight coefficient of the strategy u(t), and CB(x) is a control barrier function, and its expression is Here s(x) is a continuous and differentiable function related to the safety boundary of the game state x(t), κ > 0 is an adjustment parameter of the barrier function CB(x) when the game state x(t) approaches its safety boundary, and the function CB(x) is a risk penalty term of the performance index function as a safety constraint of the game dynamics;

[0135] Because the upper bound d M (x(t)) of the disturbance function d(x(t), u(t), v(t)) has been introduced into the performance index function (6) to offset its influence on the optimization of the game strategy u(t), the optimization of the game strategy u(t) of the nonlinear uncertain dynamic game equation (1) can ignore the influence of the disturbance function d(x(t), u(t), v(t)) in the nonlinear disturbance system (5) and only consider the nonlinear nominal system In addition, the optimization of the game strategy u(t) needs to satisfy the constraint (2) and the game state constraint (4), and essentially, the optimization of the game strategy u(t) for the nonlinear nominal system The game strategy u(t) optimization problem of the performance index function (6) is a constrained robust optimization problem, therefore, by means of the concept of constrained robust optimization, according to the nonlinear disturbance system (5) and the performance index function (6), the safety robust game problem of the nonlinear uncertain dynamic game equation (1) is converted into a constrained robust optimization problem of the following nonlinear nominal system:

[0136]

[0137] Optionally, the S3 specifically comprises:

[0138] For the constrained robust optimization problem (8) of the nonlinear nominal system, the following Hamilton function is defined:

[0139]

[0140] Where is the gradient of the performance index function J, and the following optimal value function is introduced:

[0141]

[0142] According to Pontryagin's maximum principle, the partial derivative of the performance index function The HJB equation should be satisfied:

[0143]

[0144] Assuming that the solution of the HJB equation (11) exists and is unique, the optimal game strategy of the nonlinear nominal system constraint robust optimization problem (8) is:

[0145] u * (x)=-u M tanh(η(x)) (12)

[0146] where,

[0147]

[0148] and

[0149] The HJB equation (11) is rewritten as:

[0150]

[0151] where, and

[0152] The optimal game strategy (12) depends on the analytical solution of the HJB equation (13), however, since the HJB equation (13) is a typical nonlinear partial differential equation, it is difficult to directly obtain its analytical solution, and an evaluation neural network based on reinforcement learning is designed to online approximate the optimal performance index function J * (x) and complete the approximate solution of the HJB equation (13), and obtain the optimal game strategy of the own side.

[0153] Alternatively, the specific expression of the evaluation neural network is:

[0154] J * (x)=(w * ) T θ(x)+ρ(x) (14)

[0155] where, is the ideal weight of the evaluation neural network, is the evaluation neural network base function, N is the number of neurons and ρ(x) represents the approximation error;

[0156] Correspondingly, the gradient form of J * (x) with respect to the game state x is:

[0157]

[0158] where, is the partial derivative of the base function θ(x) with respect to the game state x, the optimal game policy (12) can be expressed as:

[0159] u * (x) = -u M tanh(η(x, w * )) (16)

[0160] where,

[0161] To effectively reduce the resource consumption of the game policy, the event-triggered technique is introduced, for which a set of monotonically increasing trigger sequences where, denotes the event-triggered time, and the event-triggered dwell time Let x(t q ) be the game process dynamics at the trigger time t q , then the corresponding trigger error is:

[0162]

[0163] Using the game state x(t q ) at the trigger time t q , the event-triggered form of the game policy (16) is:

[0164] u * (x(t q )) = -u M tanh(η(x(t q ), w * )) (18)

[0165] where, and the event-triggered condition is:

[0166]

[0167] where ω > 0 is the trigger coefficient;

[0168] The corresponding event-triggered Hamiltonian function is:

[0169]

[0170] Because the ideal weight w * of the evaluation neural network is unknown, in order to realize the event-triggered game policy (18), the estimated value of the weight is used to replace its ideal weight, and the following approximate expression is constructed:

[0171]

[0172] wherein,

[0173] The approximate event-triggered Hamilton function is:

[0174]

[0175] In combination with the gradient descent algorithm, the weight update of the neural network is evaluated as:

[0176]

[0177] wherein, μ>0 is a given learning rate.

[0178] The function structure of the incomplete information constraint uncertain game strategy optimization system provided by the embodiment of the present application corresponds to the incomplete information constraint uncertain game strategy optimization method provided by the embodiment of the present application, and will not be repeated here.

[0179] Figure 8 Fig. 8 is a structural schematic diagram of an electronic device 800 provided by the embodiment of the present application. The electronic device 800 can have a large difference due to different configurations or performances, and can include one or more processors (central processing units, CPUs) 801 and one or more memories 802. The memory 802 stores at least one instruction, which is loaded and executed by the processor 801 to realize the steps of the incomplete information constraint uncertain game strategy optimization method described above.

[0180] In the exemplary embodiment, a computer readable storage medium, such as a memory including instructions, is also provided. The instructions can be executed by a processor in a terminal to complete the incomplete information constraint uncertain game strategy optimization method described above. For example, the computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0181] Those of ordinary skill in the art can understand that all or part of the steps of the above-described embodiments can be completed by hardware, or by programs instructing relevant hardware to complete, and the programs can be stored in a computer readable storage medium, such as a read-only memory, a magnetic disk or an optical disk.

[0182] The above description is only the preferred embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A strategy optimization method for uncertain games with incomplete information constraints, characterized in that, The method includes: S1. Based on the incomplete information characteristics of the dynamic process of nonlinear uncertain game under strong adversarial scenarios, a nonlinear uncertain differential game process model is constructed, wherein the strategy magnitudes of both parties and the game state in the dynamic process of the nonlinear uncertain game are constrained. S2. Treating the opponent's game strategy and unknown uncertainties as perturbation terms, and introducing the concept of constrained robust optimization, based on the nonlinear perturbation system that characterizes the dynamic process of nonlinear uncertain games, we introduce a non-quadratic energy function, a control barrier function, and an upper bound of the perturbation function to construct a suitable performance index function, thereby transforming the safe and robust game problem of the nonlinear uncertain dynamic game process into a corresponding nonlinear nominal system constrained robust optimization problem. S3. For the nonlinear nominal system constrained robust optimization problem, an event-triggered secure robust reinforcement learning technique is introduced. The optimal performance index function is approximated online through an evaluation neural network, and the corresponding Hamilton-Jacobi-Bellman HJB equation is approximated to obtain the player's optimal game strategy.

2. The method according to claim 1, characterized in that, The nonlinear uncertain differential game process model includes a nonlinear uncertain game dynamic equation, a two-player strategy magnitude constraint, and a game state constraint. The dynamic equations of the nonlinear uncertain game are as follows: in, It is a game state. and It is the strategy of both sides in the game. For one's own game strategy, For the opponent's game strategy, function It describes the local nonlinear dynamic game behavior of x(t) that is continuously differentiable and satisfies f(0) = 0. and It is a nonlinear matrix function of the game state x(t), here and Δ(x(t),u(t),v(t)) characterizes the dynamic behavior of the unmodeled local game and satisfies Δ(0,u(t),v(t))=0 and This is the initial game state, the equilibrium state (x) of the nonlinear uncertain game equation (1). e ,u e ,v e ) satisfies 0 = f(x) e )+B(x e )u e +H(x e )v e +Δ(x e ,u e ,v e Assume that (0,0,0) is the equilibrium state of the nonlinear uncertain game process (1); The game strategies u(t) and v(t) of the nonlinear uncertain game equation (1) satisfy the following magnitude constraints: in, and It is the amplitude limit of the game strategy; The game state x(t) satisfies the following constraints: in, and The game state variable x i The upper and lower bounds of (t).

3. The method according to claim 2, characterized in that, S2 specifically includes: Due to the existence of game strategy magnitude constraints (2) and (3), game state constraints (4), and unmodeled local game dynamic behavior Δ(x(t),u(t),v(t)), the design of the player's game strategy u(t) in the nonlinear dynamic game equation (1) must not only counteract the influence of the opponent's game strategy, but also overcome the adverse factors of unmodeled game dynamic behavior, while ensuring that it does not violate the above-mentioned safety constraints. Therefore, the design of the player's game strategy u(t) in the nonlinear dynamic game equation (1) is a class of safe and robust game problems, which is essentially a class of constrained uncertainty optimization problems. Using the magnitude constraints (2) and (3) of game strategies u(t) and v(t), H(x(t))v(t) and Δ(x(t),u(t),v(t)) are regarded as bounded perturbations, and the nonlinear dynamic game equation (1) is modeled as a nonlinear perturbation system: in, Let be the nonlinear perturbation function of the game process and satisfy ||d(x(t),u(t),v(t))||≤d M (x(t)), For the nonlinear perturbation system (5), under the game strategy magnitude constraints (2), (3) and game state constraints (4), design the performance index function of the optimization strategy u(t): in, It is the weight matrix of the game state, d M (x(t)) is an upper bound of the perturbation function d(x(t),u(t),v(t)), and d is introduced. M The purpose of (x(t)) is to counteract the influence of the nonlinear perturbation function d(x(t),u(t),v(t)) on the performance of the game strategy. δ>0 is a design parameter that will affect the convergence of the guidance system. U(u(t)) is a non-quadratic energy function with respect to the game strategy u(t). Here r j , Here are the weight coefficients of policy u(t), and CB(x) is the control barrier function, expressed as follows: Here, s(x) is a continuously differentiable function related to the safety bound of the game state x(t), κ>0, and is the adjustment parameter of the barrier function CB(x) when the game state x(t) approaches its safety bound. The function CB(x) serves as the risk penalty term of the game dynamic safety constraint in the performance index function. Because an upper bound d of the perturbation function d(x(t),u(t),v(t)) has already been introduced into the performance index function (6). M (x(t)) to offset its influence on the optimization of the game strategy u(t), so the optimization of the game strategy u(t) of the nonlinear uncertain dynamic game equation (1) can ignore the influence of the perturbation function d(x(t),u(t),v(t)) in the nonlinear perturbation system (5), and only consider the nonlinear nominal system. Furthermore, the optimization of the game strategy u(t) must satisfy constraints (2) and game state constraints (4). Essentially, this applies to nonlinear nominal systems. The optimization problem of the game strategy u(t) of the performance index function (6) is a class of constrained robust optimization problems. Therefore, with the help of the concept of constrained robust optimization, based on the nonlinear perturbation system (5) and the performance index function (6), the safe and robust game problem of the nonlinear uncertain dynamic game equation (1) is transformed into the following class of nonlinear nominal system constrained robust optimization problems:

4. The method according to claim 3, characterized in that, S3 specifically includes: For the constrained robust optimization problem (8) of a nonlinear nominal system, the Hamiltonian function is defined as follows: in This is the gradient of the performance index function J, and the following optimal value function is also introduced: According to Pontryagin's principle of extrema, the partial derivative of the performance index function... The HJB equation should be satisfied: Assuming that the solution to the HJB equation (11) exists and is unique, the optimal game strategy for the nonlinear nominal system constrained robust optimization problem (8) is: u * (x)=-u M tanh(η(x)) (12) in, and Combining expressions (9) and (12), HJB equation (11) is rewritten as: in, and The optimal game strategy (12) depends on the analytical solution of the HJB equation (13). However, since the HJB equation (13) is a typical nonlinear partial differential equation, it is difficult to obtain its analytical solution directly. Therefore, an evaluation neural network based on reinforcement learning is designed to approximate the optimal performance index function J online by combining the adaptive dynamic programming algorithm. * (x), and complete the approximate solution of HJB equation (13) to obtain the optimal game strategy of our side.

5. The method according to claim 4, characterized in that, The specific expression for the evaluation neural network is: J * (x)=(w * ) T θ(x)+ρ(x) (14) in, These are the ideal weights for evaluating neural networks. It evaluates the basis functions of a neural network, where N is the number of neurons and ρ(x) represents the approximation error; Accordingly, J * The gradient of (x) with respect to the game state x is in the form of: in, Let θ(x) be the partial derivative of the basis function with respect to the game state x. The optimal game strategy (12) is expressed as: u * (x)=-u M tanh(η(x,w * )) (16) in, To effectively reduce the resource consumption of game strategies, an event-triggered technique is introduced. Therefore, a set of monotonically increasing trigger sequences is defined. in, Indicates the time when the event is triggered and the duration of the event's persistence. Let x(t) q ) represents the trigger time t q Given the dynamic nature of the game process, the corresponding triggering error is: Using the trigger time t q The game state x(t) q The event triggering form of game strategy (16) is as follows: u * (x(t q ))=-u M tanh(η(x(t q ),w * )) (18) in, And the event triggering condition is: Where ω > 0 is the trigger coefficient; The corresponding event-triggered Hamiltonian function is: Because the ideal weights w for evaluating neural networks * Since it is unknown, in order to realize the event-triggered game strategy (18), the estimated weights are used instead of the ideal weights, and the following approximate expression is constructed: in, The approximate event-triggered Hamiltonian function is: Combining the gradient descent algorithm, the weight updates of the neural network are evaluated as follows: in, μ > 0 is the given learning rate.

6. A strategy optimization system for uncertain games with incomplete information constraints, characterized in that, The system includes: The first construction module is used to construct a nonlinear uncertain differential game process model based on the incomplete information characteristics of the nonlinear uncertain game dynamic process under strong adversarial scenarios. The strategy magnitudes and game states of both parties in the nonlinear uncertain game dynamic process are constrained. The second construction module is used to treat the opponent's game strategy and unknown uncertainties as perturbation terms. By introducing the concept of constraint robust optimization, based on the nonlinear perturbation system that characterizes the dynamic process of nonlinear uncertain games, a non-quadratic energy function, a control barrier function, and an upper bound of the perturbation function are introduced to construct a suitable performance index function. This transforms the safe and robust game problem of the nonlinear uncertain dynamic game process into a corresponding nonlinear nominal system constraint robust optimization problem. The solution module is used to solve the constraint-robust optimization problem of the nonlinear nominal system by introducing event-triggered secure robust reinforcement learning technology. It approximates the optimal performance index function online through an evaluation neural network and completes the approximate solution of the corresponding Hamilton-Jacobi-Bellman HJB equation to obtain the player's optimal game strategy.

7. The system according to claim 6, characterized in that, The nonlinear uncertain differential game process model includes a nonlinear uncertain game dynamic equation, a two-player strategy magnitude constraint, and a game state constraint. The dynamic equations of the nonlinear uncertain game are as follows: in, It is a game state. and It is the strategy of both sides in the game. For one's own game strategy, For the opponent's game strategy, function It describes the local nonlinear dynamic game behavior of x(t) that is continuously differentiable and satisfies f(0) = 0. and It is a nonlinear matrix function of the game state x(t), here and Δ(x(t),u(t),v(t)) characterizes the dynamic behavior of the unmodeled local game and satisfies Δ(0,u(t),v(t))=0 and This is the initial game state, the equilibrium state (x) of the nonlinear uncertain game equation (1). e ,u e ,v e ) satisfies 0 = f(x) e )+B(x e )u e +H(x e )v e +Δ(x e ,u e ,v e Assume that (0,0,0) is the equilibrium state of the nonlinear uncertain game process (1); The game strategies u(t) and v(t) of the nonlinear uncertain game equation (1) satisfy the following magnitude constraints: in, and It is the amplitude limit of the game strategy; The game state x(t) satisfies the following constraints: in, and The game state variable x i The upper and lower bounds of (t).

8. The system according to claim 7, characterized in that, S2 specifically includes: Due to the existence of game strategy magnitude constraints (2) and (3), game state constraints (4), and unmodeled local game dynamic behavior Δ(x(t),u(t),v(t)), the design of the player's game strategy u(t) in the nonlinear dynamic game equation (1) must not only counteract the influence of the opponent's game strategy, but also overcome the adverse factors of unmodeled game dynamic behavior, while ensuring that it does not violate the above-mentioned safety constraints. Therefore, the design of the player's game strategy u(t) in the nonlinear dynamic game equation (1) is a class of safe and robust game problems, which is essentially a class of constrained uncertainty optimization problems. Using the magnitude constraints (2) and (3) of game strategies u(t) and v(t), H(x(t))v(t) and Δ(x(t),u(t),v(t)) are regarded as bounded perturbations, and the nonlinear dynamic game equation (1) is modeled as a nonlinear perturbation system: in, Let be the nonlinear perturbation function of the game process and satisfy ||d(x(t),u(t),v(t))||≤d M (x(t)), For the nonlinear perturbation system (5), under the game strategy magnitude constraints (2), (3) and game state constraints (4), design the performance index function of the optimization strategy u(t): in, It is the weight matrix of the game state, d M (x(t)) is an upper bound of the perturbation function d(x(t),u(t),v(t)), and d is introduced. M The purpose of (x(t)) is to counteract the influence of the nonlinear perturbation function d(x(t),u(t),v(t)) on the performance of the game strategy. δ>0 is a design parameter that will affect the convergence of the guidance system. U(u(t)) is a non-quadratic energy function with respect to the game strategy u(t). Here r j , Here are the weight coefficients of policy u(t), and CB(x) is the control barrier function, expressed as follows: Here, s(x) is a continuously differentiable function related to the safety bound of the game state x(t), κ>0, and is the adjustment parameter of the barrier function CB(x) when the game state x(t) approaches its safety bound. The function CB(x) serves as the risk penalty term of the game dynamic safety constraint in the performance index function. Because an upper bound d of the perturbation function d(x(t),u(t),v(t)) has already been introduced into the performance index function (6). M (x(t)) to offset its influence on the optimization of the game strategy u(t), so the optimization of the game strategy u(t) of the nonlinear uncertain dynamic game equation (1) can ignore the influence of the perturbation function d(x(t),u(t),v(t)) in the nonlinear perturbation system (5), and only consider the nonlinear nominal system. Furthermore, the optimization of the game strategy u(t) must satisfy constraints (2) and game state constraints (4). Essentially, this applies to nonlinear nominal systems. The optimization problem of the game strategy u(t) of the performance index function (6) is a class of constrained robust optimization problems. Therefore, with the help of the concept of constrained robust optimization, based on the nonlinear perturbation system (5) and the performance index function (6), the safe and robust game problem of the nonlinear uncertain dynamic game equation (1) is transformed into the following class of nonlinear nominal system constrained robust optimization problems:

9. The system according to claim 8, characterized in that, S3 specifically includes: For the constrained robust optimization problem (8) of a nonlinear nominal system, the Hamiltonian function is defined as follows: in This is the gradient of the performance index function J, and the following optimal value function is also introduced: According to Pontryagin's principle of extrema, the partial derivative of the performance index function... The HJB equation should be satisfied: Assuming that the solution to the HJB equation (11) exists and is unique, the optimal game strategy for the nonlinear nominal system constrained robust optimization problem (8) is: u * (x)=-u M tanh(η(x)) (12) in, and Combining expressions (9) and (12), HJB equation (11) is rewritten as: in, and The optimal game strategy (12) depends on the analytical solution of the HJB equation (13). However, since the HJB equation (13) is a typical nonlinear partial differential equation, it is difficult to obtain its analytical solution directly. Therefore, an evaluation neural network based on reinforcement learning is designed to approximate the optimal performance index function J online by combining the adaptive dynamic programming algorithm. * (x) and complete the approximate solution of HJB equation (13) to obtain the optimal game strategy of our side.

10. The system according to claim 9, characterized in that, The specific expression for the evaluation neural network is: J * (x)=(w * ) T θ(x)+ρ(x) (14) in, These are the ideal weights for evaluating neural networks. It evaluates the basis functions of a neural network, where N is the number of neurons and ρ(x) represents the approximation error; Accordingly, J * The gradient of (x) with respect to the game state x is in the form of: in, Let θ(x) be the partial derivative of the basis function with respect to the game state x. The optimal game strategy (12) is expressed as: u * (x)=-u M tanh(η(x,w * )) (16) in, To effectively reduce the resource consumption of game strategies, an event-triggered technique is introduced. Therefore, a set of monotonically increasing trigger sequences is defined. in, Indicates the time when the event is triggered and the duration of the event's persistence. Let x(t) q ) represents the trigger time t q Given the dynamic nature of the game process, the corresponding triggering error is: Using the trigger time t q The game state x(t) q The event triggering form of game strategy (16) is as follows: u * (x(t q ))=-u M tanh(η(x(t q ),w * )) (18) in, And the event triggering condition is: Where ω > 0 is the trigger coefficient; The corresponding event-triggered Hamiltonian function is: Because the ideal weights w for evaluating neural networks * Since it is unknown, in order to realize the event-triggered game strategy (18), the estimated weights are used instead of the ideal weights, and the following approximate expression is constructed: in, The approximate event-triggered Hamiltonian function is: Combining the gradient descent algorithm, the weight updates of the neural network are evaluated as follows: in, μ > 0 is the given learning rate.