One-to-one target defense game decision-making method under perception limitation, electronic equipment and medium

By constructing target defense game scenarios and game models under perceptually restricted conditions, combining microelement method and Apollonian circles to solve the maximum and critical pursuit angles, the problem of one-to-one target defense game under perceptual restricted conditions is solved, the optimal strategic decision of the agent is achieved, and the interception effect is improved.

CN120373462APending Publication Date: 2025-07-25BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510492238.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology cannot effectively solve the one-to-one target defense game problem in multi-agent systems under the conditions of perception restriction. The traditional Apollonian circle analysis is no longer applicable, resulting in the defender being unable to effectively intercept attackers within the perception range.

Method used

Build a target defense game scenario and game model, combine the microelement method and the Apollonian circle, and determine the optimal strategy of the agent by solving the maximum pursuit angle and critical pursuit angle.

Benefits of technology

It provides an optimal control strategy in perceptually restricted environments, improves the effect of defenders intercepting attackers, and is suitable for multi-agent decision-making control in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373462A_ABST
    Figure CN120373462A_ABST
Patent Text Reader

Abstract

The invention discloses a one-to-one target defense game decision-making method under perception limitation, electronic equipment and a medium. The method comprises the following steps: establishing a target defense game scene, and constructing a target defense game model; solving a maximum pursuit angle through an Apollonius circle; solving a critical pursuit angle through an infinitesimal method; and according to the maximum pursuit angle and the critical pursuit angle, solving an intelligent agent optimal strategy. On the basis of a differential game, a target defense game scene and a game model are constructed, an infinitesimal method is combined with an Apollonius circle, and classification discussion is performed on the pursuit angle of the game so as to solve the optimal decision of the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-agent decision-making control, and more specifically, to a one-on-one target defense game decision-making method, an electronic device, and a medium under limited perception. Background Art

[0002] Multi-agent decision-making control in an adversarial environment has been a hot topic in recent years. With the development of artificial intelligence technology, multi-agent systems have been widely used in many fields, including unmanned aerial vehicle formations, autonomous vehicle convoys, smart grid management, logistics and supply chain optimization, and complex strategic game simulations. In these scenarios, especially in adversarial environments, the behavior of agents not only needs to consider cooperation but also must deal with potential or explicit competition and conflicts. And the macroscopic large-scale cluster confrontation scenario can be regarded as composed of several sub-confrontation scenarios each containing only a few agents. Therefore, analyzing the strategy tendencies of agents in sub-confrontation scenarios is of great significance for studying the multi-agent decision-making control problem in macroscopic confrontation scenarios.

[0003] The target defense problem is a typical scenario in the field of multi-agent decision-making control. In a sub-confrontation scenario, the target defense problem usually involves one or two attacking agents responsible for approaching and destroying the target, and such agents are called attackers; another type of agent is responsible for intercepting these attackers approaching the target to protect the target, and is called a defender. The attacking and defending sides take approaching and protecting the target as the basis of the game. The attacker tries to take the optimal action strategy as much as possible to break through the defender's interception; while the defender also takes the optimal strategy to intercept the attacker as early as possible. Eventually, the result of this sub-confrontation game ends with the attacker successfully seizing the target or the defender successfully intercepting the attacker. Therefore, the key to the target defense problem is to find the optimal strategies of the attacker and the defender based on their respective goals. At the practical application level, the results of the target defense problem can be applied in scenarios such as hostage rescue and flag capture on the battlefield. Unmanned aerial vehicles or unmanned vehicles take optimal strategy actions according to the conclusions of the target defense problem to complete the mission objectives. Therefore, solving the target defense problem has guiding significance for the research in the field of multi-agent decision-making control.

[0004] The target defense problem usually assumes that each agent has global information, that is, it knows the state information of all other agents and the strategies they adopt. However, in the actual environment, due to factors such as obstacle occlusion, noise interference, and area shielding, it is often difficult for each agent to obtain the state information of the enemy agents. Therefore, studying the target defense game under the constraint of the perception range is more in line with the actual scenario. Under the constraint of the perception radius, some agents can only obtain the state information of the enemy agents within the perception range. When the defender has a perception range constraint, it should keep the attacker within the perception range as much as possible without losing the target during the pursuit; while the attacker needs to escape from the defender's perception range as much as possible to escape the pursuit. Therefore, the one-on-one target defense problem under the limited perception range has great research value.

[0005] Usually, the differential game method is used to model and solve such target defense problems, and the Apollonius circle is used to analyze the optimal strategies and decisive regions of the attacking and defending agents, so as to solve the target defense problem in the sub-confrontation scenario. However, due to the existence of perception constraints, the defender needs to keep the attacker within the perception radius during the pursuit. Therefore, using the traditional Apollonius circle to analyze the optimal strategies of agents is no longer applicable to solve the target defense problem under limited perception.

[0006] Therefore, it is necessary to develop a one-on-one target defense game decision-making method, electronic device and medium under limited perception.

[0007] The information disclosed in the background art section of the present invention is only intended to deepen the understanding of the general background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0008] The present invention proposes a one-on-one target defense game decision-making method, electronic device and medium under limited perception. On the basis of differential game, a target defense game scenario and game model are constructed, and the microelement method is combined with the Apollonius circle to classify and discuss the pursuit angle of the game to solve the optimal decision of the agent.

[0009] In the first aspect, an embodiment of the present disclosure provides a one-on-one target defense game decision-making method under limited perception, including:

[0010] Establish a target defense game scenario and construct a target defense game model;

[0011] Solve the maximum pursuit angle through the Apollonius circle;

[0012] Solve the critical pursuit angle through the microelement method;

[0013] Solve the optimal strategy of the agent according to the maximum pursuit angle and the critical pursuit angle.

[0014] Preferably, establishing the target defense game scenario includes:

[0015] Set that there are two agents in the target defense game scenario, namely the attacker and the defender, and conduct the target defense game on the plane where the target defense game is located;

[0016] Determine that the initial distance between the attacker and the defender on the plane is r0. If the real-time distance r between the attacker and the defender on the plane is less than or equal to r0 and decreasing, then set the sensing range of the defender as a circular area with the defender as the center and a radius of r, and the attacker has no sensing range limit;

[0017] The goal of the attacker is to break through the interception of the defender, capture the target, and reach the coordinates where the target is located. The goal of the defender is to intercept the attacker and prevent it from capturing the target.

[0018] Preferably, constructing the target defense game model includes:

[0019] Taking the defender as the origin, establish a plane rectangular coordinate system on the plane where the target defense game is located;

[0020] Determine the termination condition of the target defense game as: when the attacker captures the target or the attacker is outside the sensing range of the defender, the attacker wins; when the defender pursues and catches the attacker, the defender wins;

[0021] Establish the objective function of the target defense game:

[0022]

[0023] Among them, x0 represents the initial state of the system, t f represents the game termination time;

[0024] Furthermore, determine the value function V(x) of the target defense game as:

[0025]

[0026] Preferably, when the pursuit direction of the defender is tangent to the current Apollonius circle, the maximum pursuit angle is obtained.

[0027] Preferably, the critical pursuit angle is:

[0028]

[0029] Among them, θ cis the critical pursuit angle, and λ is the maximum speed ratio of the defender to the attacker.

[0030] Preferably, solving the optimal strategy of the agent according to the maximum pursuit angle and the critical pursuit angle includes:

[0031] Determine the magnitude relationship between the maximum pursuit angle and the critical pursuit angle, and then determine the optimal strategies of the defender and the attacker.

[0032] Preferably, when the maximum speeds of the defender and the attacker When, the maximum pursuit angle is less than or equal to the critical pursuit angle, then the optimal strategies of the defender and the attacker are:

[0033]

[0034] Preferably, when the maximum speeds of the defender and the attacker When, the maximum pursuit angle is greater than the critical pursuit angle, then the optimal strategies of the defender and the attacker are:

[0035]

[0036] In a second aspect, an embodiment of the present disclosure further provides an electronic device, which includes:

[0037] A memory storing executable instructions;

[0038] A processor that runs the executable instructions in the memory to implement the one-on-one target defense game decision-making method under limited perception.

[0039] In a third aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the one-on-one target defense game decision-making method under limited perception.

[0040] Its beneficial effects are as follows:

[0041] 1. The control decision-making method of the present invention based on differential game theory has strict mathematical derivations, providing a mathematical expression of the control strategy for agents to perform target defense tasks in a limited perception environment;

[0042] 2. The present invention takes into account the limited perception situation in the actual complex environment. Compared with similar tracking and interception algorithms, this algorithm has obvious interception advantages and is the optimal decision under this target defense game.

[0043] The methods and apparatuses of the present invention have other characteristics and advantages that will be apparent from, or will be described in detail in, the accompanying drawings and the subsequent detailed description incorporated herein. The accompanying drawings and the detailed description together are used to explain the specific principles of the present invention. Description of the Drawings

[0044] The above and other objects, features, and advantages of the present invention will become more apparent by describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, in which, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.

[0045] Figure 1 A flowchart showing the steps of a one-on-one target defense game decision-making method under perception constraints according to an embodiment of the present invention.

[0046] Figure 2 A schematic diagram showing a one-on-one target defense game scenario under perception constraints according to an embodiment of the present invention.

[0047] Figure 3 A schematic diagram showing an interception scenario where the target is outside the Apollonius circle according to an embodiment of the present invention.

[0048] Figure 4 A schematic diagram showing an interception scenario where the target is inside the Apollonius circle according to an embodiment of the present invention.

[0049] Figure 5 A schematic diagram showing the Apollonius circle and the maximum pursuit angle according to an embodiment of the present invention.

[0050] Figure 6 A schematic diagram showing the analysis of the movement tendency of an agent by the infinitesimal element method according to an embodiment of the present invention.

[0051] Figure 7 A schematic diagram showing that the maximum pursuit angle is less than or equal to the critical pursuit angle according to an embodiment of the present invention.

[0052] Figure 8 A schematic diagram showing that the maximum pursuit angle is greater than the critical pursuit angle according to an embodiment of the present invention.

[0053] Figure 9 A schematic diagram showing a target defense game simulation under perception constraints with a speed ratio λ = 1.5 according to an embodiment of the present invention.

[0054] Figure 10 A schematic diagram showing a target defense game simulation under perception constraints with a speed ratio λ = 1.45 according to an embodiment of the present invention.

[0055] Figure 11 The schematic diagram shows the optimal interception strategy with a speed ratio λ = 1.35 according to an embodiment of the present invention.

[0056] Figure 12 The schematic diagram shows the tracking interception strategy with a speed ratio λ = 1.35 according to an embodiment of the present invention. Detailed implementation manners

[0057] The preferred embodiments of the present invention will be described in more detail below. Although the preferred embodiments of the present invention are described below, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0058] Figure 1 The flowchart shows the steps of the one-to-one target defense game decision-making method under limited perception according to an embodiment of the present invention.

[0059] As Figure 1 shown, the one-to-one target defense game decision-making method under limited perception includes:

[0060] Step 101, establish a target defense game scenario and construct a target defense game model;

[0061] Step 102, solve the maximum pursuit angle through the Apollonius circle;

[0062] Step 103, solve the critical pursuit angle through the infinitesimal method;

[0063] Step 104, solve the optimal strategy of the agent according to the maximum pursuit angle and the critical pursuit angle.

[0064] In one example, establishing a target defense game scenario includes:

[0065] Set that there are two agents in the target defense game scenario, namely the attacker and the defender, and conduct a target defense game on the plane where the target defense game is located;

[0066] Determine that the initial distance between the attacker and the defender on the plane is r0. If the real-time distance r between the attacker and the defender on the plane is less than or equal to r0 and decreasing, set the perception range of the defender as a circular area with the defender as the center and a radius of r, and the attacker has no perception range limit;

[0067] The goal of the attacker is to break through the interception of the defender, capture the target, and reach the coordinates where the target is located. The goal of the defender is to intercept the attacker and prevent it from capturing the target.

[0068] In one example, constructing a target defense game model includes:

[0069] Taking the defender as the origin, establish a plane rectangular coordinate system on the plane where the target defense game is located;

[0070] Determine the termination condition of the target defense game as: when the attacker captures the target or the attacker is outside the defender's perception range, the attacker wins; when the defender pursues and catches the attacker, the defender wins;

[0071] Establish the objective function of the target defense game:

[0072]

[0073] where \(x_0\) represents the initial state of the system, and \(t\) f represents the termination time of the game;

[0074] Furthermore, determine the value function \(V(x)\) of the target defense game as:

[0075]

[0076] In an example, when the pursuit direction of the defender is tangent to the current Apollonius circle, the maximum pursuit angle is obtained.

[0077] In an example, the critical pursuit angle is:

[0078]

[0079] where \(\theta\) c is the critical pursuit angle, and \(\lambda\) is the maximum speed ratio of the defender to the attacker.

[0080] In an example, according to the maximum pursuit angle and the critical pursuit angle, solving the optimal strategy of the agent includes:

[0081] Determine the size relationship between the maximum pursuit angle and the critical pursuit angle, and then determine the optimal strategies of the defender and the attacker.

[0082] In an example, when the maximum speeds of the defender and the attacker satisfy a certain condition, if the maximum pursuit angle is less than or equal to the critical pursuit angle, the optimal strategies of the defender and the attacker are:

[0083]

[0084] In an example, when the maximum speeds of the defender and the attacker satisfy another condition, if the maximum pursuit angle is greater than the critical pursuit angle, the optimal strategies of the defender and the attacker are:

[0085]

[0086] Specifically, the present invention is an optimal decision-making control algorithm for target defense tasks in a perception-limited environment. Based on differential games, a target defense game scenario and a game model are constructed. By combining the infinitesimal method with the Apollonius circle, the pursuit angle of the game is classified and discussed to solve the optimal decision of the agent. The specific technical solutions are as follows:

[0087] Figure 2 FIG. shows a schematic diagram of a one-on-one target defense game scenario under perception limitation according to an embodiment of the present invention.

[0088] Construct a target defense game scenario. Assume that the plane Ω is the plane where the target defense game is located, and the target is located at a certain point on the plane. The two agents are the attacker and the defender, and they conduct a target defense game on the plane. The attacker and the defender are located at two points on the plane, with a distance of r between them. The goal of the attacker is to break through the interception of the defender and capture the target (i.e., reach the coordinates where the target is located), while the goal of the defender is to intercept the attacker and prevent it from capturing the target. When the coordinates of the defender and the attacker coincide, it is regarded as a successful interception. The defender has limited perception. It is determined that the initial distance between the attacker and the defender on the plane is r0. If the real-time distance r between the attacker and the defender on the plane is less than or equal to r0 and decreasing, then the perception range of the defender is set as a circular area with the defender as the center and a radius of r; while the attacker has no perception range limitation. Since the perception radius r cannot be increased, and the defender can only obtain the state information of the attacker within the perception range, the defender needs to ensure that the distance between it and the attacker will not increase at any moment during the pursuit process. In addition, different from the general target defense game problem, the attacker has a decision-making advantage, that is, the defender needs to make a decision first, and the attacker makes a decision based on the decision made by the defender. The specific target defense game scenario is as Figure 2 shown.

[0089] When the attacker captures the target or the attacker is outside the perception range of the defender, the game ends and the attacker wins; while when the defender catches up with the attacker, the defender wins. If the defender catches up with the attacker exactly when the attacker captures the target, neither wins. For the target defense game problem under perception limitation, the strategies with the greatest possibility of winning for the attacker and the defender need to be found respectively, that is, the optimal penetration strategy and the optimal interception strategy.

[0090] Construct a target defense game model. Taking the defender as the origin, a plane rectangular coordinate system is established on the plane Ω. Let the initial coordinates of the attacker be (x a0 , y a0 ), and the coordinates at time t be (x a (t), y a (t)) (subsequently abbreviated as (x a , y a )); the initial coordinates of the defender are (xd0 , y d0 )(Since the defender is initially located at the origin, it can be known that x d0 = y d0 = 0). At time t, the coordinates are (x d (t), y d (t)) (subsequently abbreviated as (x d , y d )). Let the target coordinates be (x f , y f ). According to the given conditions, the initial distance between the attacker and the defender is exactly the perception radius r of the defender. It can be obtained that:

[0091]

[0092] Let the maximum speed of the attacker be v a , and the maximum speed of the defender be v d . The speed ratio of the defender to the attacker is λ. Then there are:

[0093]

[0094] Establish a first-order motion model for the agent. Let the control inputs of the attacker and the defender be respectively Satisfy:

[0095]

[0096] Respectively represent the motion directions of the attacker and the defender at time t. The agent can instantaneously change the motion direction (first-order motion model). According to differential game theory, both the attacker and the defender adopt optimal strategies during the game. Therefore, the attacker and the defender always move at their maximum speeds v a and v d respectively. Therefore, the motion equations of the attacker and the defender during the entire target defense game can be obtained:

[0097]

[0098] Regarding the attacker and the defender as a system, and using to represent the state quantity of the system, the differential equation of the system evolution can be obtained:

[0099]

[0100] where f represents a function with the system state, the attacker control input, and the defender control input as variables.

[0101] According to the game victory conditions of the attacker and the defender: when the attacker captures the target, the game ends and the attacker wins; when the defender captures the attacker, the defender wins. Based on this, the termination conditions of the target defense game are given.

[0102]

[0103] According to the optimization objectives of the game between the attacker and the defender, the following objective function of the target defense game is established:

[0104]

[0105] In the above equation, x0 represents the initial state of the system, and t f represents the game termination time. It can be seen from the equation that the objective function of the game is the distance from the attacker to the target at the end of the game. The attacker hopes that the objective function is as small as possible, and when it is reduced to zero, it means successful capture; the defender hopes that the objective function is as large as possible, so as to intercept the attacker at a farther distance. Based on this, the value function V(x) of the target defense game is defined as:

[0106]

[0107] The value function of the target defense game can accurately reflect the interest trends of both the attacker and the defender. It can not only reflect whether the game is won or not, but also reflect how much is won or lost. The value function transforms the entire target defense game from a qualitative game to a quantitative game, and is an important indicator to verify the advantages and disadvantages of the strategies of both the attacker and the defender, so it has important significance.

[0108] Solve the maximum pursuit angle through the Apollonius circle. The Apollonius circle, also known as the Apollonian circle or A-circle, is defined as the set of points whose distances to two points A and B on the plane are in a fixed ratio (not equal to 1). In recent years, the Apollonius circle has been widely used in differential game problems to solve the optimal strategies of both sides of the differential game. In the target defense game, the Apollonius circle can be redefined as the set of points that the attacker and the defender can reach simultaneously.

[0109] Figure 3 Shows a schematic diagram of an interception scenario where the target is outside the Apollonius circle according to an embodiment of the present invention.

[0110] Figure 4 Shows a schematic diagram of an interception scenario where the target is inside the Apollonius circle according to an embodiment of the present invention.

[0111] Through the Apollonius circle, the entire plane is divided into three parts: inside the Apollonius circle, outside the Apollonius circle, and on the Apollonius circle. According to the definition of the Apollonius circle, the points on the circle are the points that both the attacker and the defender can reach simultaneously, so they can be regarded as the boundary fence of the general target defense game. If the target is on the Apollonius circle, the attacker and the defender cannot distinguish a winner, because when both the attacker and the defender adopt optimal decisions, they will necessarily reach the location of the target at the same time. For the points outside the Apollonius circle, the defender can reach them earlier than the attacker. Therefore, if the target is outside the Apollonius circle, the defender will surely be able to intercept the attacker before the attacker seizes the target. On the contrary, for the points inside the Apollonius circle, the attacker can reach them earlier than the defender. Therefore, if the target is inside the Apollonius circle, the attacker will seize it first and the defender cannot intercept. Two typical target defense scenarios are as Figure 3 , Figure 4 shown.

[0112] In the process of the target defense game, according to the algebraic and geometric relationships of the Apollonius circle, there is the following theorem: If the current state variables of the attacker are (x a , y a ), and the state variables of the defender are (x d , y d ), the distance between the two sides is r, and the speed ratio of the defender to the attacker is λ. Then the Apollonius circle formed by the attacker and the defender is a circle with O c (x c , y c ) as the center and r c as the radius, and satisfies the following conditions:

[0113]

[0114] This theorem shows that the current Apollonius circle can be solved through the positions of the current attacker and defender. This is of great significance for studying the target defense game problem, that is, the result of the target defense game (both the attacker and the defender adopt optimal strategies) can be predicted in advance by determining the relationship between the target and the current Apollonius circle.

[0115] Define the pursuit angle as the angle between the pursuit direction of the defender and the line connecting the attacker and the defender. Then the pursuit angle has a maximum value, that is: Let the pursuit angle of the defender be θ, then the pursuit angle has a maximum value θ max , satisfying:

[0116]

[0117] Figure 5 shows a schematic diagram of the Apollonius circle and the maximum pursuit angle according to an embodiment of the present invention.

[0118] When θ reaches its maximum value, and only when the pursuer's chasing direction is tangent to the current Apollonius circle. As Figure 5 shown.

[0119] Figure 6 shows a schematic diagram of analyzing the motion tendency of an agent by the infinitesimal element method according to an embodiment of the present invention.

[0120] Solve the critical pursuit angle by the infinitesimal element method. In an environment with limited sensing range, the defender needs to always keep the attacker within the sensing range and make a decision prior to the attacker. Therefore, the traditional Apollonius circle is no longer applicable to solve such target defense game problems, so the Apollonius circle needs to be improved. The infinitesimal element method is applicable to analyzing the algebraic and geometric relationships of the motion of an object within a tiny moment, and has a wide range of application scenarios, which is an important entry point for analyzing the current problem. Analyze the motion tendency of the agent within a tiny time interval by the infinitesimal element method as Figure 6 shown. Within a tiny time Δt, the defender tends to intercept the attacker at the target point F. Let the distance advanced by the attacker within the infinitesimal time be Δx, then the distance advanced by the defender is λΔx, and the pursuit angle at this time is θ, and the defender advances to the position D′ accordingly. After the defender determines the pursuit strategy, the attacker makes a decision later, and the decision-making process can be regarded as finding a point on the circle with the attacker as the center and Δx as the radius to determine its velocity direction. Since the defender's sensing range is limited, the defender needs to keep the attacker within the field of view during the pursuit process; while the attacker has the tendency to escape from the defender's sensing range r. Therefore, the attacker's optimal decision is to find a point on the circle such that the point A * , which is the intersection of the extension line of D′A and the circle, is the farthest from the defender's next position D′.

[0121] The attacker moving in the AA * direction is the strategy by which it can escape from the constraint of the defender's sensing radius to the greatest extent. Therefore, the critical situation needs to be considered, that is, if the attacker adopts this strategy, it can just maintain the distance from the defender unchanged (or just escape from the defender's sensing range), that is:

[0122] ||A * D′|| = r

[0123] Combining the above equations, in ΔAD′D, according to the cosine theorem, we can get:

[0124] (r - Δx) 2 = r 2 +(λΔx) 2 - 2rλΔx cosθ

[0125] It is solved that the pursuit angle in the critical situation should satisfy the following relational expression:

[0126]

[0127] That is, the critical pursuit angle θ c should satisfy:

[0128]

[0129] When the pursuit angle of the defender is less than or equal to the critical pursuit angle θ c no matter what strategy the attacker takes, it cannot escape from the perception range of the defender; however, if the pursuit angle of the defender is greater than the critical pursuit angle θ c there exists a certain strategy for the attacker to escape from the perception range of the defender, resulting in the failure of the defender's interception. Thus, it can be obtained that in the one-on-one target defense game with limited perception, the defender needs to always ensure that the current pursuit angle θ is less than or equal to the critical pursuit angle θ c , that is, satisfy:

[0130] θ ≤ θ c

[0131] Otherwise, there exists a certain strategy for the attacker to escape from the perception range of the defender. In the formula, θ c satisfies

[0132] Solve the optimal strategy of the agent. In the process of the one-on-one target defense game with limited perception, there is a maximum value for the pursuit angle, and it is restricted by the critical pursuit angle θ c Therefore, the strategy of the defender depends on the size relationship between the two, and thus a case-by-case discussion is carried out.

[0133] Figure 7 shows a schematic diagram of the maximum pursuit angle being less than or equal to the critical pursuit angle according to an embodiment of the present invention.

[0134] Figure 8 shows a schematic diagram of the maximum pursuit angle being greater than the critical pursuit angle according to an embodiment of the present invention.

[0135] As Figure 7 shown, at this time the maximum pursuit angle is less than or equal to the critical pursuit angle. At this time, no matter what strategy the attacker takes, it cannot escape from the perception range of the defender. Therefore, in this case, the constraint of the critical pursuit angle does not need to be considered, and the traditional Apollonius circle can be used for analysis.

[0136] As Figure 8As shown, the maximum pursuit angle is greater than the critical pursuit angle at this time. At this time, the attacker has a certain strategy to escape from the defender's perception range. Therefore, this situation needs to consider the constraint of the critical pursuit angle and cannot be analyzed using the traditional Apollonius circle. If the defender intercepts according to the original advancing direction, the attacker can take advantage of the later-mover decision-making advantage and adopt a strategy to escape from the defender's perception range. At this time, the defender should advance along the direction of the critical constraint angle, which can not only ensure that the attacker cannot escape from the perception range but also ensure approaching the attacker as quickly as possible.

[0137] By comparing the magnitude relationship between the maximum pursuit angle and the critical pursuit angle, the following inequality group can be solved to obtain different cases of classification discussion:

[0138] θ≤θ c

[0139] When the speed ratio is such that the maximum pursuit angle is less than or equal to the critical pursuit angle, the defender does not need to consider the constraints brought by the limited perception range; when the speed ratio is such that the maximum pursuit angle is greater than the critical pursuit angle, the defender needs to adopt a strategy to ensure that the attacker cannot escape from its perception range. And the defender's strategy in this case is to intercept the attacker along the direction of the critical pursuit angle. After the defender makes a decision and considering the attacker's decision, it can be inferred that the attacker's optimal strategy is to go straight to the direction where the target is located to capture the target as quickly as possible. To sum up, it can be obtained that in the one-on-one target defense game under perception limitation, assuming the expected interception point is T(x t ,y t ). When the speed ratio is such that the optimal strategies of both sides are:

[0140]

[0141] When the speed ratio is such that the optimal strategies of both sides are:

[0142]

[0143] In solving the problem of target defense, the present invention considers the complex environment of perception limitation, and through rigorous mathematical derivation, gives the mathematical expression of the optimal control strategy under this constraint. Compared with the same type of algorithms, it has better practicability. For the scenario where the defender's perception range is limited, a target defense game scenario and a game model are constructed. This game model is based on differential game theory and combines the continuous decision-making game model of Stackelberg game. And by combining the infinitesimal method with the Apollonius circle, the conditions that the pursuit angle should satisfy in the game are analyzed, and finally the problem is solved through classification discussion.

[0144] The present invention also provides an electronic device, which includes: a memory storing executable instructions; and a processor that runs the executable instructions in the memory to implement the one-on-one target defense game decision-making method under limited perception as described above.

[0145] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the one-on-one target defense game decision-making method under limited perception as described above.

[0146] To facilitate understanding of the solutions and effects of the embodiments of the present invention, the following gives three specific application examples. Those skilled in the art should understand that these examples are only for facilitating understanding of the present invention, and no specific details are intended to limit the present invention in any way.

[0147] Example 1

[0148] This embodiment verifies the effectiveness of the given algorithm in the one-on-one target defense game scenario under limited perception.

[0149] Figure 9 The figure shows a schematic diagram of the target defense game simulation under limited perception with a speed ratio λ = 1.5 according to an embodiment of the present invention.

[0150] Figure 10 The figure shows a schematic diagram of the target defense game simulation under limited perception with a speed ratio λ = 1.45 according to an embodiment of the present invention.

[0151] In the first group of simulations, the speed ratio of the defender to the attacker The defender is initially located at (0, 0), the attacker is initially located at (6, 1), and the target is located at (8, 9). Considering the two cases of the speed ratio of the defender to the attacker being 1.5 and 1.45, the simulation results are respectively as Figure 9 、 Figure 10 shown.

[0152] In Figure 9 the target defense game scenario, the speed ratio of the defender to the attacker is 1.5. Since the speed ratio is greater than the maximum pursuit angle is less than the critical pursuit angle, the influence brought by limited perception does not need to be considered, and the strategies of both the attacker and the defender can be analyzed using the traditional Apollonius circle. Since the target is outside the initial Apollonius circle, the defender successfully intercepts the attacker at the position of T(7.86, 8.48). The value function of the target defense game V(x) = 0.29.

[0153] In Figure 10 the target defense game scenario, the speed ratio of the defender to the attacker is 1.45. Since the speed ratio is also greater than The maximum pursuit angle is less than the critical pursuit angle, so the impact of perception limitation does not need to be considered. However, compared with the scenario shown in Figure 9 when the speed of the defender slows down, it affects the range of the initial Apollonius circle, resulting in the target being within the initial Apollonius circle, so the defender cannot intercept the attacker. The value function of this target defense game V(x) = 0, and the attacker successfully captures the target at the position F(8, 9).

[0154] Figure 11 FIG. shows a schematic diagram of the optimal interception strategy with a speed ratio λ = 1.35 according to an embodiment of the present invention.

[0155] Figure 12 FIG. shows a schematic diagram of the tracking interception strategy with a speed ratio λ = 1.35 according to an embodiment of the present invention.

[0156] In the second group of simulations, the speed ratio of the defender to the attacker The defender is initially located at (0, 0), the attacker is initially located at (6, 1), and the target is located at (8, 12). The speed ratio of the defender to the attacker is 1.35. Considering two defender strategies: in the scenario shown in Figure 11 the defender adopts the optimal interception strategy; while in the target defense game scenario shown in Figure 12 the defender adopts the tracking interception strategy, that is, when the pursuit angle is greater than the critical pursuit angle, the defender moves towards the position of the attacker (the pursuit angle is zero) to ensure quickly shortening the distance from the defender.

[0157] In Figure 11 In the target defense game scenario shown, the speed ratio of the defender to the attacker is 1.35. Since the speed ratio is also less than the maximum pursuit angle is greater than the critical pursuit angle, so the impact of perception limitation needs to be considered. In this scenario, the defender adopts the optimal interception strategy for interception. It can be seen from Figure 11 that the trajectory of the defender is a curve, indicating that the expected pursuit angle of the defender is always greater than the critical pursuit angle. Finally, the defender successfully intercepts the attacker at T(7, 72, 10.50). It is worth mentioning that compared with the traditional Apollonius circle pursuit strategy, the defender can intercept the attacker at a farther place (the interception point is slightly outside the Apollonius circle). However, this result is very reasonable. Because the defender has a limitation in the perception range, in order to avoid losing the attacker, a trade-off strategy is adopted, that is, to intercept the attacker as quickly as possible on the premise of ensuring that the attacker is within the field of view. From this perspective, the defender's strategy is the optimal strategy in this scenario. The value function of this target defense game V(x) = 2.33.

[0158] In Figure 12In the target defense game scenario shown, the speed ratio of the defender to the attacker is 1.35. Different from Figure 11 the only difference in the target defense game scenario shown is that the defender adopts a tracking and intercepting strategy when the pursuit angle is greater than the critical angle. It can be seen from the simulation results that the defender's trajectory consists of two obvious segments. The first segment of the trajectory is a curve, representing that the defender adopts a tracking and intercepting strategy; the second segment of the trajectory is a straight line, representing that the defender adopts the traditional Apollonius circle strategy for interception. Finally, the defender successfully intercepts the attacker at T(7.80, 10.95). For Figure 11 the scenario where the defender adopts the optimal interception strategy, Figure 12 the interception point is more biased towards the target, representing a worse interception effect. It can be seen from the value function of the game that: V(x) = 1.14 < 2.33. The smaller the value function of the target defense game, the more advantageous the attacker is (a value function of zero means the attacker successfully captures the target); while the larger the value function, the better the interception effect of the defender. Since the pursuit angle of the tracking and intercepting strategy is zero, it can indeed shortest the distance with the defender fastest and ensure that the attacker cannot escape the perception range of the defender. However, the tracking and intercepting strategy ignores the element of the target. The primary goal of the attacker is to capture the target. Therefore, the defender adopting the tracking and intercepting strategy, due to not having a strategic advantage, can only follow the footsteps of the attacker, without the foresight of the strategy, and cannot predict the subsequent movement trajectory of the attacker. Therefore, its interception effect is naturally not as good as the optimal interception strategy. Through calculation, it can be found that compared with the tracking and intercepting strategy, the interception effect of the optimal interception strategy has been significantly improved, which also confirms the superiority of the proposed defender interception strategy.

[0159] Example 2

[0160] The present disclosure provides an electronic device, which includes: a memory storing executable instructions; a processor that runs the executable instructions in the memory to implement the above one-on-one target defense game decision method under limited perception.

[0161] The electronic device according to an embodiment of the present disclosure includes a memory and a processor.

[0162] The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0163] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory.

[0164] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain a good user experience effect, well-known structures such as communication buses and interfaces may also be included in this embodiment, and these well-known structures should also be included in the protection scope of the present disclosure.

[0165] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0166] Example 3

[0167] An embodiment of the present disclosure provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the one-to-one target defense game decision-making method under limited perception described above.

[0168] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the methods of the foregoing embodiments of the present disclosure are executed.

[0169] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).

[0170] Those skilled in the art should understand that the purpose of the above description of the embodiments of the present invention is only to exemplarily illustrate the beneficial effects of the embodiments of the present invention, and is not intended to limit the embodiments of the present invention to any of the examples given.

[0171] The above has described the embodiments of the present invention. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

Claims

1. A one-to-one target defense game decision-making method under perception constraints, characterized in that Including: Establish a target defense game scenario and construct a target defense game model; Solve for the maximum pursuit angle through the Apollonius circle; Solve for the critical pursuit angle through the infinitesimal method; Solve for the optimal strategy of the agent according to the maximum pursuit angle and the critical pursuit angle.

2. The one-to-one target defense game decision-making method under limited perception according to claim 1, wherein, Establishing a target defense game scenario includes: Set that there are two agents in the target defense game scenario, namely the attacker and the defender, and conduct a target defense game on the plane where the target defense game is located; Determine that the initial distance between the attacker and the defender on the plane is r0. If the real-time distance r between the attacker and the defender on the plane is less than or equal to r0 and decreasing, set the perception range of the defender as a circular area with the defender as the center and radius r, and there is no perception range limit for the attacker; The goal of the attacker is to break through the interception of the defender, capture the target, and reach the coordinates where the target is located. The goal of the defender is to intercept the attacker and prevent it from capturing the target.

3. The one-to-one target defense game decision-making method under limited perception according to claim 2, wherein, Constructing a target defense game model includes: Taking the defender as the origin, establish a plane rectangular coordinate system on the plane where the target defense game is located; Determine the termination condition of the target defense game as: when the attacker captures the target or the attacker is outside the perception range of the defender, the attacker wins; when the defender catches up with the attacker, the defender wins; Establish the objective function of the target defense game: where \(x_0\) represents the initial state of the system, and \(t\) f represents the game termination time; Furthermore, determine the value function V(x) of the target defense game as:

4. The one-to-one target defense game decision-making method under limited perception according to claim 2, wherein, When the pursuit direction of the defender is tangent to the current Apollonius circle, the maximum pursuit angle is obtained.

5. The one-to-one target defense game decision-making method under limited perception according to claim 1, wherein, The critical pursuit angle is: Among them, θ c is the critical pursuit angle, and λ is the maximum speed ratio between the defender and the attacker.

6. The one-to-one target defense game decision-making method under limited perception according to claim 2, wherein, Solving for the optimal strategy of the agent according to the maximum pursuit angle and the critical pursuit angle includes: Determine the size relationship between the maximum pursuit angle and the critical pursuit angle, and then determine the optimal strategies of the defender and the attacker.

7. The one-to-one target defense game decision-making method under limited perception according to claim 6, wherein, When the maximum speed of the defender and the attacker is such that the maximum pursuit angle is less than or equal to the critical pursuit angle, the optimal strategies for the defender and the attacker are:

8. The one-to-one target defense game decision-making method under limited perception according to claim 6, wherein, When the maximum speed of the defender and the attacker is such that the maximum pursuit angle is greater than the critical pursuit angle, the optimal strategies for the defender and the attacker are as follows:

9. An electronic device, characterized in that, The electronic device includes: A memory storing executable instructions; A processor that runs the executable instructions in the memory to implement the one-on-one target defense game decision-making method under perception limitation described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the one-on-one target defense game decision-making method under perception limitation described in any one of claims 1-8.