Pursuit game control method for attack of multi-agent inter-satellite link under incomplete information

By constructing a pursuit-escape game model and a distributed state observer, combined with an adaptive attack compensator, the problems of communication link attacks and incomplete information in multi-agent systems are solved, and stable pursuit-escape control in complex environments is achieved.

CN121485993APending Publication Date: 2026-02-06XIAN AISHENG TECH GRP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511603537.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies cannot effectively handle the problems of communication link attacks and incomplete state information in multi-agent systems, leading to system performance degradation or failure and making it difficult to achieve multi-agent collaborative control.

Method used

A pursuit-escape game model is constructed, and the Nash equilibrium is solved using Hamilton-Jacobi-Bellman partial differential equations. By combining a distributed state observer and an adaptive attack compensator, control strategies for pursuing the leader and the escapee, as well as pursuing the followers, are designed to form a pursuit-escape game control method for multi-agent inter-satellite links under attack.

Benefits of technology

In situations where communication is disrupted and some information is missing, the system achieves stable operation and Nash equilibrium of the multi-agent system, improving the system's feasibility, security, and optimality in complex adversarial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121485993A_ABST
    Figure CN121485993A_ABST
Patent Text Reader

Abstract

The invention particularly relates to a pursuit game control method for attacking a multi-agent inter-satellite link under incomplete information. According to the method, a many-to-one pursuit game problem is converted into two sub-problems of a one-to-one pursuit game between a pursuit leader and an escaper and cooperative following of a pursuit follower to the pursuit leader. In consideration of the situation that an intelligent agent cannot acquire speed information of other intelligent agents, an intelligent agent system is modeled as a second-order kinematic model, and meanwhile, a distributed state observer is constructed to estimate a non-neighbor state, so that global state information acquisition under an incomplete information condition is realized. Furthermore, an adaptive filter is designed to suppress the influence of the inter-satellite link false data injection attack on the state estimation and control performance, so that the stable operation of the intelligent agent system is kept under the norm bounded random attack interference condition. Therefore, according to the method, multi-agent pursuit game solving can be realized under the conditions that communication is disturbed and part of information is lost, so that the system approaches Nash equilibrium.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of agent control, in particular to a pursuit-evasion game control method for multi-agent inter-satellite link under imperfect information. BACKGROUND

[0002] Under the background of rapid development of intelligent and networked systems, multi-agent pursuit-evasion confrontation research shows important value in many fields. This kind of problem simulates the dynamic game relationship between the "pursuer" and the "evader", and provides a theoretical basis and technical support for autonomous cooperative decision-making, path planning and resource allocation of complex systems.

[0003] In the field of high dynamic cooperative operation, the pursuit-evasion game idea has been widely used in autonomous evasion and interception, group maneuver coordination, target approach and interception, etc. For example, the cooperative evasion system of autonomous equipment can evade threats based on the pursuit-evasion game decision mechanism; in the task of surrounding and approaching the target by multiple flight platforms or multiple ground unmanned platforms, the pursuit-evasion game idea is often used to realize the cooperative control and planning of multiple agents; in the task of searching and inspecting ground targets by intelligent mobile platforms, the target positioning and tracking can be modeled as a pursuit-evasion problem of pursuing static or low-maneuvering targets.

[0004] In the field of space autonomous service and on-orbit maintenance, the pursuit-evasion game theory also has important application value. For example, when a fuel-limited spacecraft approaches a non-cooperative space target, it can use the pursuit-evasion game strategy to achieve safe and efficient approach. In addition, in the field of public safety and civil use, the related technology can be used for target search and path planning in emergency rescue, as well as traffic flow regulation and intelligent traffic management systems.

[0005] Although the method based on differential game theory has covered most of the pursuit-evasion application scenarios, it usually assumes that the multi-agent system has complete state information perception ability and reliable communication link. However, in actual deployment, the space or ground communication network relied on by the multi-agent cooperative system may face malicious behaviors such as external signal interference and data tampering, resulting in degradation or even failure of system performance. Attack behavior is hidden and dynamic, and limited by observation equipment and bandwidth conditions, it is difficult for the agent to obtain complete link state and potential attack strategy in real time. In addition, under actual perception conditions, the agent often cannot accurately obtain the dynamic information of all neighbors, and can only obtain the position information of the neighboring agents, and the speed or control input state is incomplete. Therefore, the existing differential game method cannot be directly applied to the multi-agent pursuit-evasion scene with link disturbance and state incompleteness. SUMMARY

[0006] The application provides a pursuit-evasion game control method for multi-agent inter-satellite link under attack under incomplete information, a computer readable storage medium and a computer program product, which can effectively overcome the defects in the prior art.

[0007] Other characteristics and advantages of the application will become apparent from the detailed description that follows, or will be learned by practice of the application.

[0008] According to a first aspect of the application, a pursuit-evasion game control method for multi-agent inter-satellite link under attack under incomplete information is provided, the method comprising: constructing a pursuit-evasion game model of the pursuit leader and the evader; and solving the pursuit-evasion game model to determine the control strategy of the pursuit leader and the evader and the reference state of the pursuit leader by using a Hamilton-Jacobi-Bellman partial differential equation, wherein the pursuit-evasion game model is established based on a relative motion dynamics model of the pursuit leader and the evader, and a zero-sum cost function of the game model is constructed based on minimizing the pursuit cost and maximizing the escape cost; constructing a cooperative following model between the pursuit follower and the pursuit leader; and solving the cooperative following model by using a distributed state observer and an adaptive attack compensator to obtain the control strategy of the pursuit follower, wherein a following cost function of the cooperative following model is constructed based on the reference state of the pursuit leader and the position vector difference between the pursuit follower and the pursuit leader; forming a pursuit-evasion game control strategy for multi-agent inter-satellite link under attack based on the control strategy of the pursuit leader and the evader and the control strategy of the pursuit follower, wherein the pursuit-evasion game control strategy is used to control the multi-pursuit agents to cooperatively pursue the evading agent and the evading agent to evade the multi-pursuit agents.

[0009] In some example embodiments, the constructing a pursuit-evasion game model of the pursuit leader and the evader comprises: establishing a relative motion dynamics model based on a motion model of the pursuer and the evader and combining relative state variables of the pursuer and the evader; constructing a zero-sum cost function of the pursuit-evasion game model based on the relative motion dynamics model, wherein the zero-sum cost function comprises a pursuit cost function and an escape cost function, the pursuit cost function is constructed based on a relative state variable error, a control cost of the pursuit leader and a control cost of the evader, and the escape cost function is a dual expression of the pursuit cost function.

[0010] In some example embodiments, the zero-sum cost function comprises: the pursuit cost function comprises:

[0011] the escape cost function comprises:

[0012] wherein, is the motion start time; is a state variable of the pursuit-evasion game model; is a system state weight of the pursuit-evasion game model; is a control variable of the pursuit leader; is a proportion of energy consumed by the pursuit leader in the pursuit cost function; is a control variable of the evader; is a proportion of energy consumed by the evader in the pursuit cost function.

[0013] In some example embodiments, the constructing a cooperative following model between the pursuit followers and the pursuit leader comprises: establishing a first communication topology directed graph between the pursuit followers, and a second communication topology directed graph between the pursuit leader and the pursuit followers; constructing a following cost function based on a communication weight and a position vector difference between adjacent pursuit followers in the first communication topology directed graph, a reference state of the pursuit leader, and a communication weight between the pursuit followers and the pursuit leader in the second communication topology directed graph.

[0014] In some example embodiments, the following cost function comprises:

[0015] wherein, ; is a position vector difference between the pursuit follower and other pursuit followers to which the pursuit follower has a communication connection; is a communication weight between adjacent pursuit followers; is a position vector difference between the pursuit follower and the pursuit leader ; is a communication weight between the pursuit follower and the pursuit leader; is a gain matrix of the following cost function.

[0016] In some example embodiments, the solving the cooperative following model by using a distributed state observer and an adaptive attack compensator to obtain a control strategy of the pursuit follower comprises: using the distributed state observer to estimate state information of non-adjacent pursuit followers to obtain a global state estimate; The adaptive attack compensator is used to compensate the global initial state estimator to obtain a modified global state estimator; The modified global state estimator and the reference state of the pursuit leader are input into a following cost function, and the control strategy of the pursuit follower is obtained by solving the following cost function with the dynamic equation of the pursuit follower as a constraint.

[0017] In some example embodiments, the adaptive attack compensator comprises:

[0018] wherein, is a positive definite matrix; is an estimated value of an auxiliary state variable; is an estimate of an upper bound of the injected attack signal; is an arbitrary positive definite continuous bounded function.

[0019] According to a second aspect of the present application, a computer readable storage medium is provided, which comprises a stored executable program, wherein the executable program, when executed, controls a device where the storage medium is located to perform the above-mentioned multi-agent inter-satellite link attack pursuit and evasion game control method under incomplete information.

[0020] According to a third aspect of the present application, a computer program product is provided, which comprises a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned multi-agent inter-satellite link attack pursuit and evasion game control method under incomplete information.

[0021] According to a fourth aspect of the present application, an electronic device is provided, which comprises: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to implement the above-mentioned multi-agent inter-satellite link attack pursuit and evasion game control method under incomplete information by executing the executable instructions.

[0022] According to a fifth aspect of the present application, a storage medium is provided, which stores a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned multi-agent inter-satellite link attack pursuit and evasion game control method under incomplete information.

[0023] The method for pursuit-evasion game control of multi-agent inter-satellite link under attack under incomplete information provided by the embodiment of the application converts the multi-to-one pursuit-evasion game problem into one-to-one pursuit-evasion game of the pursuit leader and the evader, and the two sub-problems of the cooperative following of the pursuit follower to the pursuit leader. Considering that the agent cannot obtain the speed information of other agents, the agent system is modeled as a second-order kinematics model, and at the same time, a distributed state observer is constructed to estimate the non-neighbor state, so as to realize the acquisition of global state information under the condition of incomplete information. Further, an adaptive filter is designed to suppress the influence of the inter-satellite link false data injection attack on the state estimation and control performance, so as to keep the agent system stable under the condition of norm-bounded random attack interference. Therefore, the method can realize the solution of the multi-agent pursuit-evasion game under the condition of disturbed communication and partial information loss, so that the system approaches the Nash equilibrium.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the application. BRIEF DESCRIPTION OF DRAWINGS

[0025] The drawings herein are incorporated into the specification and form part of the specification, show embodiments consistent with the application, and together with the specification serve to explain the principles of the application. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained from these drawings without creative labor for those skilled in the art.

[0026] Figure 1 The flow chart of the method for pursuit-evasion game control of multi-agent inter-satellite link under attack under incomplete information is schematically shown in an exemplary embodiment of the application; Figure 2 The schematic diagram of the attack on the inter-satellite communication link of the agent in the method for pursuit-evasion game control of multi-agent inter-satellite link under attack under incomplete information is schematically shown in an exemplary embodiment of the application; Figure 3 The specific implementation schematic diagram of the method for pursuit-evasion game control of multi-agent inter-satellite link under attack under incomplete information is schematically shown in an exemplary embodiment of the application; Figure 4 The schematic diagram of the pursuit-evasion game between the pursuer and the evader in the method for pursuit-evasion game control of multi-agent inter-satellite link under attack under incomplete information is schematically shown in an exemplary embodiment of the application; Figure 5A The distance curve of the follower 1 from the leader in three coordinate directions in the method for pursuit-evasion game control of multi-agent inter-satellite link under attack under incomplete information is schematically shown in an exemplary embodiment of the application; Figure 5BIllustration of the distance curve of the follower 2 from the leader in three coordinate directions in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 5C Illustration of the distance curve of the follower 3 from the leader in three coordinate directions in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 6A Illustration of the control curve of the follower 1 in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 6B Illustration of the control curve of the follower 2 in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 6C Illustration of the control curve of the follower 3 in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 7 Illustration of the distance curve of the leader from the evader in three coordinate directions in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 8 Illustration of the control curve of the leader in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 9 Illustration of the control curve of the evader in a pursuit-evasion game control method for multiple agents under imperfect information when the inter-satellite link is attacked; Figure 10 Illustration of the composition of an electronic device in an example embodiment of the present application. DETAILED DESCRIPTION

[0027] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0028] Furthermore, the accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this specification. The drawings are not intended to be restrictive in any way. Throughout the drawings, like references numerals denote like features. Descriptions of the same or similar elements are omitted from being repeated. Some of the blocks in the drawings are functional entities that do not necessarily have to correspond to physically or logically independent entities. These functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0029] In view of the disadvantages and deficiencies of the prior art, a pursuit-evasion game control method for multi-agent inter-satellite link attack under incomplete information is provided in the example embodiment. Referring to Figure 1 As shown, the method can specifically include: In step S10, a pursuit-evasion game model of the pursuit leader and the evader is constructed, and a Hamilton-Jacobi-Bellman partial differential equation is used to solve the Nash equilibrium of the pursuit-evasion game model to determine the control strategy of the pursuit leader and the evader and the reference state of the pursuit leader. The pursuit-evasion game model is established based on the relative motion dynamics model of the pursuit leader and the evader, and a zero-sum cost function of the game model is constructed based on minimizing the pursuit cost and maximizing the escape cost. In step S10, the pursuit leader refers to a core pursuit unit in a multi-pursuit agent system for directly pursuing and evading the evading agent. The pursuit leader constructs and solves the Nash equilibrium strategy based on the pursuit-evasion confrontation model to obtain the optimal pursuit control law, which is used as the reference trajectory or control direction of the entire pursuit group to drive other pursuit agents to perform cooperative hunting behavior. That is, the pursuit leader is responsible for calculating the optimal pursuit action and is the strategy center of the pursuit task.

[0030] The evader refers to the target agent to be pursued, whose strategy goal is opposite to that of the pursuit leader, and tries to escape from the pursuit group through maneuvering avoidance, path planning and interference means. The evader constructs a game model against the pursuit leader and generates an optimal escape control input based on the confrontation strategy to maximize the safety distance or escape probability. That is, the evader is the opposite party of the pursuit task, and its behavior drives the dynamic evolution of the entire pursuit-evasion process.

[0031] In step S12, a cooperative following model between the pursuit follower and the pursuit leader is constructed, and a distributed state observer and an adaptive attack compensator are used to solve the cooperative following model to obtain the control strategy of the pursuit follower. The following cost function of the cooperative following model is constructed based on the reference state of the pursuit leader and the position vector difference between the pursuit follower. In step S12, the above-mentioned pursuit follower refers to a pursuit agent in a multi-pursuit agent system, which does not directly game with the escape agent, but tracks the position, velocity or attitude of the pursuit leader through a distributed cooperative control strategy according to the reference state of the pursuit leader. That is, the pursuit follower is responsible for tracking the pursuit leader to achieve the pursuit-escape task.

[0032] For example, in the active removal scene of on-orbit debris, the escapee is a high-speed maneuvering space debris or a failed satellite; the pursuit leader is a removal satellite main hull, which plans to approach and capture the attitude; and the pursuit follower is a small satellite group that collaborates to monitor and provide navigation assistance.

[0033] In step S14, a pursuit-escape game control strategy for multi-agent interstellar link under attack is formed based on the control strategies of the pursuit leader and the escapee and the control strategy of the pursuit follower; wherein the pursuit-escape game control strategy is used to control the multi-pursuit agent to cooperatively pursue and capture the escape agent and the escape agent to evade the multi-pursuit agent.

[0034] Specifically, considering a pursuit-escape game problem of one pursuit agent pursuing one escapee, the multi-pursuit agent system is regarded as a leader-follower formation system, and thus the many-to-one pursuit-escape game problem is converted into two sub-problems: a leader-follower formation problem and a one-to-one pursuit-escape game problem. For these two sub-problems, corresponding pursuit-escape game models of the pursuit leader and the escapee and a cooperative following model between the pursuit follower and the pursuit leader are established.

[0035] Further, for the pursuit-escape game problem under the condition that the interstellar communication link of the multi-agent is attacked by a norm-bounded random false data injection attack signal, a Nash search algorithm based on adaptive filtering is proposed to design a corresponding control strategy for each agent. Meanwhile, considering that the follower can only obtain the position information of the neighbors and cannot obtain the speed information of the neighbor followers, the relative motion dynamics equation of the agent is written in the form of a second-order model, and a distributed observer is designed for each follower to estimate the state information of the non-neighbors. Then, an adaptive filter is designed for each follower to reduce the influence of false data injection attacks on the performance of the follower.

[0036] Based on the above steps S10 to S14, the many-to-one pursuit and evasion game problem is converted into a one-to-one pursuit game between the pursuit leader and the evader, and a two-sub-problem of cooperative following of the pursuit follower to the pursuit leader. Considering that the agent cannot obtain the speed information of other agents, the agent system is modeled as a second-order kinematics model, and at the same time, a distributed state observer is constructed to estimate the non-neighbor state, so as to realize the global state information acquisition under the condition of incomplete information. Further, an adaptive filter is designed to suppress the influence of the false data injection attack on the state estimation and control performance of the inter-satellite link, so as to keep the agent system stable under the condition of norm-bounded random attack interference. Therefore, the method can solve the multi-agent pursuit and evasion game under the condition of disturbed communication and partial information loss, so that the system approaches the Nash equilibrium.

[0037] In the following, each step of the multi-agent pursuit and evasion game control method under incomplete information in the present example embodiment will be described in more detail in combination with the accompanying drawings and examples.

[0038] Illustratively, in step S10, the pursuit and evasion game model of the pursuit leader and the evader is constructed, including: Step S101, based on the motion model of the pursuer and the evader, and in combination with the relative state variables of the pursuer and the evader, a relative motion dynamics model is established; Step S102, based on the relative motion dynamics model, a zero-sum cost function of the pursuit and evasion game model is constructed; wherein the zero-sum cost function includes a pursuit cost function and an evasion cost function, the pursuit cost function is composed based on the relative state variable error, the control cost of the pursuit leader and the control cost of the evader; the evasion cost function is a dual expression of the pursuit cost function.

[0039] Illustratively, the zero-sum cost function includes: The pursuit cost function includes:

[0040] The evasion cost function includes:

[0041] wherein, is the motion start time; is the state variable of the pursuit and evasion game model; is the system state weight of the pursuit and evasion game model; is the control variable of the pursuit leader; is the proportion of the energy consumed by the pursuit leader in the pursuit cost function; is the control variable of the evader; is the proportion of the energy consumed by the evader in the pursuit cost function.

[0042] Specifically, first, the pursuit leader and the evader motion dynamics model is constructed. When the state variables of the agent are all known, the first-order linear system can be used. Therefore, the motion dynamics model of the agent is shown in expression (1): (1) Wherein, the state variable , vector is the position of the agent, vector is the speed of the agent, so ; the state matrix ; the control matrix , is the control variable.

[0043] On this basis, the relative motion dynamics model is established in combination with the relative state variables of the pursuer and the evader. Specifically: in the pursuit and evasion differential game, the motion of the pursuer and the evader both need to satisfy , that is, need to satisfy as shown in expression (2): (2) Wherein, is the state variable of the pursuit leader; is the control variable of the pursuit leader; is the state variable of the evader; is the control variable of the evader.

[0044] In order to describe more conveniently, the state variable of the pursuit and evasion game system is defined as the difference between the state variables of the pursuit leader and the evader relative to the virtual reference point, that is, the form of the state variable of the pursuit and evasion game system is shown in expression (3): (3) For the pursuit and evasion game problem of the multi-agent system, the relative motion between the agents is usually used for description. Further, the relative motion dynamics model of the pursuer and the evader, that is, the state space equation of the pursuit and evasion game system is described as shown in expression (4): (4) Further, in the infinite time differential game, both the pursuit leader and the evader decide to carry out the game to the end, and the fuel they need to consume is sufficient. Then, the cost function of the two parties will change, the pursuit cost function of the pursuit leader is shown in expression (5): (5) The evading cost function of the evader is shown in expression (6): (6) Among them, among them, This refers to the start time of the movement; For the state variables in the pursuit-escape game model; The system state weights for the pursuit-escape game model; Control variables for tracking the leader; The proportion of energy consumed in pursuing the leader in the pursuit cost function; For the escapee, the control variable; The proportion of energy consumed by the escapee in the pursuit cost function.

[0045] Furthermore, the one-to-one pursuit-escape game is a zero-sum game. The optimal control strategies for both players are solved using the Hamilton-Jacobi-Bellman partial differential equations for constructing differential games. The obtained optimal control strategies for the pursuer and the escapee are shown in expression (7): (7) Among them, matrix By solving the algebraic Riccati equation as shown in expression (8), we obtain: (8) For example, in step S12, constructing a collaborative following model between the follower and the leader includes: Step S121: Establish a first communication topology directed graph between pursuers and followers, and a second communication topology directed graph between the pursuer leader and pursuers and followers; Specifically, intelligent agents need to exchange information, therefore an information communication topology needs to be constructed: (1) Use to indicate The first communication topology is a directed graph between the pursuers and followers. In the graph... middle, Represents a set of follower nodes, where, Indicates the first A node that chases and follows others. This indicates the total number of followers; This represents a set of edges in a directed graph.

[0046] Use numbers Indicates the first The pursuer and the first A connection of followers. The weighted adjacency matrix is ​​defined as follows: In the weighted adjacency matrix middle, For directed graphs The connectivity weights. If ,but denotes the chaser follower The information that the chaser follower is able to obtain is denoted by If , then denotes no communication connection. In addition, for any host follower , . Meanwhile, since the graph is a directed communication topology graph, it is not necessarily true that equals , so the weighted adjacency matrix is asymmetric.

[0047] Let denote the set of all chaser followers that are able to communicate with the th chaser follower. The Laplacian matrix of the graph is defined as where , .

[0048] (2) The directed graph denotes a second communication topology directed graph between the chaser leaders and the chaser followers. In the graph , a set of agent nodes is denoted by , where denotes the leader nodes, and denotes a set of edges of the directed graph. The graph only contains the communication between the leaders and the followers, but does not contain their internal communication. The connection that the follower is able to obtain the state information of the leader is denoted by , and the weight of the communication is given by . The connection that the leader is able to obtain the state information of the th follower is denoted by , and the weight of the communication is given by .

[0049] In step S122, a following cost function is constructed based on the communication weight between the adjacent chaser followers in the first communication topology directed graph and the position vector difference, the reference state of the chaser leader, and the communication weight between the chaser followers and the chaser leader in the second communication topology directed graph.

[0050] Illustratively, the following cost function includes:

[0051] where ; is the chaser follower The sum of position vector differences between it and other pursuers with which it has a communication connection; The communication weight between adjacent followers; To pursue the followers With the pursuit of the leader The position vector difference between them; The weight of communication between followers and leaders; This is the gain matrix that follows the cost function.

[0052] Specifically, the pursuit-escape game problem described using a first-order linear system has been well-studied. In this case, the state variables of each agent participate in the game. However, in practical applications, it is difficult for agents to obtain the velocity information of other agents, and the first-order linear system described above is not applicable to this situation. Therefore, it is necessary to write the position information and velocity information of the agents as two separate state variables to construct a second-order system model of the relative motion dynamics equations of the agents. The relative motion dynamics equations of the agents are written as a second-order system model, and the specific expression is shown in expression (9): (9) Furthermore, based on the second-order system model, the local neighborhood error variable and the following cost function of the follower are constructed.

[0053] First, based on the movement target of the pursuer, a local domain error variable is established for the pursuer. : (10) in, To pursue the followers The sum of position vector differences between it and other pursuers with which it has a communication connection; The communication weight between adjacent followers; To pursue the followers With the pursuit of the leader The position vector difference between them; To determine the communication weight between followers and leaders, record... .

[0054] Based on local domain error variables No. The following cost function of the chaser system is shown in expression (11): (11) in, It is a positive definite symmetric matrix, and is the gain matrix of the cost function.

[0055] For example, in step S12, the process of solving the cooperative following model using a distributed state observer and an adaptive attack compensator to obtain the control strategy for the pursuer includes: Step S123: Using a distributed state observer, estimate the state information of non-adjacent pursuers to obtain a global state estimate. Specifically, in the case of an incomplete information chase-follower game, each follower can only obtain the state information of its own neighbors. However, since its follower cost function includes the states of all other followers, each follower also needs to estimate the state information of its non-neighbors in order to further optimize its own follower cost function. Therefore, a distributed state observer is designed for the follower. The details are as follows: First, define new variables. As an auxiliary variable used to estimate global state information, its form is shown in expression (12): (12) in, Indicates the first Auxiliary variables for each follower; Indicates the first One pursuer and follower; Indicates location, Indicates speed.

[0056] For ease of description, the definition is... For the first The estimation of auxiliary state information of other followers by a single pursuer is shown in the form of expression (13): (13) in, It is the first The pursuer followed the first The state variables of the pursuers The estimated value.

[0057] definition It is the first Estimation of auxiliary state variables for each pursuer and other followers The combination of these elements is shown in the specific form of expression (14): (14) It needs to be emphasized that the first A follower transmits a combined state to its neighbors, not its own individual state variable.

[0058] Further, based on the setting of the above-mentioned auxiliary state variable, in the absence of disturbance or attack, the pursuer The form of the designed distributed state observer is shown in expression (15): (15) In order to obtain the Nash equilibrium point of the follower cost function The most commonly used method at present is to take the partial derivative vector of the cost function as part of the control input of the follower system. At the same time, considering that the partial derivative vector of the cost function will be affected by the dynamic characteristics of the relative motion system of the intelligent agent, therefore, a compensation term of the original dynamic system characteristics needs to be added to the control input. In addition, as can be seen from the form of the follower distributed state observer The estimation vector of all auxiliary state variables of the follower is not included in the distributed state observer Therefore, compensation is also needed for this part in the control input. From the above, the control input vector of the pursuer is designed as shown in expression (16): (16) Wherein, the forms of matrices and are shown in expression (17): (17) Further, the dynamic system of the pursuer may be shown in expression (18): (18) According to the definitions of the above and , the following results shown in expression (19) can be obtained: ; ; ; (19) Further, the dynamic equation of the combined vector of the auxiliary state variable of the th pursuer and the auxiliary state variable estimation of other pursuers is shown in expression (20): (20) Step S124, based on the adaptive attack compensator, attack compensation is performed on the global initial state estimation to obtain a modified global state estimation; Step S125, input the modified global state estimation and the reference state of the chasing leader into the follower cost function, solve the follower cost function with the chasing follower's dynamic equation as a constraint, and obtain the control strategy of the chasing follower.

[0059] Exemplarily, the adaptive attack compensator comprises:

[0060] wherein, is a positive definite matrix; is an estimated value of the auxiliary state variable; is an estimation of the upper bound of the injected attack signal; is an arbitrary positive definite continuous bounded function.

[0061] Specifically, since the chasing follower transmits the combined state variable to its neighbor, not the state variable of itself, that is, the state vector is transmitted, and the state vector is transmitted through the communication network, which is easily affected by the network attack. As shown in Figure 2 , the chasers can transmit information, but the information transmission link is attacked, that is, the inter-satellite communication link of the intelligent agent is attacked. Figure 2 is the chasing follower When the chasing follower transmits the state information, it is attacked. Therefore, the communication network is considered to be attacked by false data injection, that is, the chasing follower injects the random false data with a bounded norm into the combined state vector when transmitting it to the chasing follower . In the presence of random false data injection attack, in order to reduce the influence of the attack on the system performance of the chasing follower, an adaptive attack compensator is designed for each chasing follower. The information received by the follower from its neighbor is described as: (21) wherein the false data injection attack signal satisfies , and the upper bound of the attack signal is an unknown positive constant.

[0062] Specifically, for malicious attackers to bypass attack detectors undetected, the attack signal needs certain constraints; on the other hand, considering the finite energy of the injected attack signal, a bounded attack signal is reasonable. Therefore, the upper bound of the attack signal can be known. However, attack signals are often invisible and difficult to detect, making it difficult to obtain their magnitude. Therefore, the assumption that the attack signal is bounded and its amplitude is unknown is more realistic. Thus, in the presence of spoofed data injection attacks, targeting follower attacks... The form of the distributed state observer designed for it is shown in expression (22):

[0063] (twenty two) in, , It is a novel adaptive attack compensator designed to adaptively counteract the effects of attacks.

[0064] The adaptive attack compensator is designed as shown in expression (23): (twenty three) in, It is an estimate of the upper bound of the false data injection attack signal, as shown in expression (24): (twenty four) in, , It is a positive definite matrix; It is any positive definite continuous bounded function that satisfies the inequality shown in expression (25): (25) Among them, the function It can be an exponential decay function ,at this time .

[0065] Furthermore, in the presence of fake data injection attacks, in order to track down followers... Design a control input vector as shown in expression (26): (26) Furthermore, the pursuit of followers The dynamic system in the presence of a fake data injection attack can be rewritten as shown in expression (27): (27) From the above, we can obtain the first... Estimation of auxiliary state variables of individual followers and other followers in the presence of a fake data injection attack. The dynamics equation of the combination vector is shown in the following formula: (28) Each pursuit follower generates its own control input by using the control update law shown in expressions (27)-(28) based on the reference state of the leader, the distributed state observer and the adaptive attack compensator, so as to achieve cooperative approximation to the pursuit leader, thereby maintaining group cooperative pursuit and game optimality under the conditions of incomplete information and false data injection attack.

[0066] Specifically, the method provided by the embodiment of the application is shown in the following formula: Figure 3 The application provides a pursuit-evasion game control method for multiple intelligent agents under attack of inter-satellite links in incomplete information, which comprises the following steps: Step 1, constructing a pursuit-evasion game model of the intelligent agent system: for the pursuit-evasion game problem of the multiple intelligent agent system, the relative motion between the intelligent agents is generally described, and when the state variables of the intelligent agents are all known, the problem is generally described as a first-order linear system; Step 2, solving a one-to-one pursuit-evasion game problem: the one-to-one pursuit-evasion game problem is a zero-sum game, and the Hamilton-Jacobi-Bellman partial differential equation for constructing a differential game is used to solve the optimal control strategy of both parties in the game; Step 3, respectively constructing a communication topology graph model of the leader-follower and the followers: the followers and the leader have information transmission, so it is necessary to construct a communication topology model for them, and the communication connection between the followers and the leader is described by a directed graph; Step 4, constructing a second-order system model of the relative motion dynamics equation of the intelligent agent: the pursuit-evasion game problem described by the first-order linear system has been studied relatively maturely. In this case, the state variables of each intelligent agent are involved in the game. However, generally, the intelligent agent is difficult to obtain the speed information of other intelligent agents, and the first-order linear system is not suitable for this case. Therefore, the position information and the speed information of the intelligent agent are written as two state variables respectively; Step 5, constructing a local neighborhood error variable and a cost function of the follower: according to the motion target of the follower, a local neighborhood error variable is established for the follower, and a cost function is designed for the follower; Step 6, designing a distributed state observer for the follower: in the case of incomplete information pursuit-evasion game, each follower can only obtain the relevant information of its neighbors, but the state of all other followers is included in its cost function, so each follower needs to estimate the state information of its non-neighbors, and further optimize its cost function; Step 7, designing an adaptive filter: the follower inter-satellite communication link is subjected to a norm-bounded random false data injection attack, in order to reduce the influence of the attack on the system performance, an adaptive attack compensator is designed for each follower.

[0067] Specifically, the method provided by the embodiment of the application takes a spacecraft pursuit-evasion game problem as an implementation case to verify the effectiveness of the method. Considering that a distributed multi-pursuit spacecraft system is subjected to a norm-bounded random frequency-diverse array (FDA) signal, the information communication link between the pursuit spacecraft clusters is destroyed. Meanwhile, considering that the speed information measured by the spacecraft is not accurate in an actual scene, the spacecraft can only obtain the position information of neighbor spacecrafts, taking four pursuit spacecrafts (three follower spacecrafts and one leader spacecraft) as an example, the weight matrix parameters in the cost function are respectively:

[0068] The adjacency matrix of the pursuit spacecrafts is respectively: ;

[0069] The initial state of each spacecraft is: ; ; ; ; .

[0070] Taking the positive definite continuous bounded function in the adaptive attack compensator , the attack signal received by the communication link of the first follower spacecraft and its neighbor follower spacecrafts is: .

[0071] The simulation result shown in Fig. 23 is obtained through simulation verification. Figures 4-9 It can be seen from the simulation result graph that, at about 23s, the follower spacecraft catches up with the leader spacecraft, and then the four pursuit spacecrafts form a certain formation to pursue the escape spacecraft. At about 30s, the relative position curve of the leader spacecraft and the escape spacecraft converges to the error allowable range, that is, the pursuit spacecraft successfully catches up with the escape spacecraft, which indicates that the algorithm proposed in the application is effective.

[0072] The application realizes robust distributed state estimation, adaptive anti-attack, and multi-agent optimal pursuit-evasion game control under the conditions of incomplete information and network attacks, and significantly improves the feasibility, security and optimality of the system in a complex confrontation environment. Specific beneficial effects are as follows: (1) Realize multi-agent cooperative pursuit-evasion control under incomplete information By constructing a pursuit leader-pursuit follower cooperative following framework and designing a distributed state observer, the pursuit agent can still estimate the global state under the condition of only being able to obtain local information of neighbors, thereby realizing cooperative pursuit control of the evading target. This mechanism effectively solves the problem of relying on complete global information in traditional methods, and improves the feasibility and robustness of multi-agent cooperation in information-limited scenarios.

[0073] (2) Resist inter-satellite link false data injection attacks and improve system security and reliability By designing an adaptive attack compensator for the communication link between agents, real-time estimation and compensation of false data injection attacks are performed, so that the distributed state observation error remains bounded, thereby avoiding system divergence or failure caused by attacks. This method enhances the communication robustness and network security capability of the inter-satellite link when subjected to malicious interference.

[0074] (3) Realize optimal control and Nash equilibrium of multi-agent pursuit-evasion game The multi-pursuit-single-evade game is divided into leader-follower cooperation and one-on-one differential game sub-problems, and the follower cost function is minimized under the pursuit follower dynamics constraint to obtain the optimal control strategy, realizing the Nash equilibrium of the pursuit-evasion game system. Compared with traditional centralized solution methods, the present invention has the advantages of distributed calculation, strong real-time performance and provable convergence.

[0075] It should be noted that the above-described figures are only schematic representations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to limit the purpose. It is easy to understand that the processes shown in the above-described figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0076] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units embodied.

[0077] Figure 10 A schematic diagram of an electronic device suitable for implementing embodiments of the present application is shown.

[0078] It should be noted that, Figure 10 The electronic device 1000 shown is only an example and should not limit the function and scope of use of the embodiments of the present application.

[0079] like Figure 10 As shown, the electronic device 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from storage section 1008 into Random Access Memory (RAM) 1003. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004. Furthermore, the electronic device 1000 also includes an FPGA device and a System-on-a-Chip (SoC) device.

[0080] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1010 as needed so that computer programs read from them can be installed into storage section 1008 as needed.

[0081] In particular, according to embodiments of the present invention, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.

[0082] Specifically, the aforementioned electronic devices can be airborne intelligent electronic devices.

[0083] It should be noted that the storage medium shown in the embodiments of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the above two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any storage medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination of the above.

[0084] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a part of code containing one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0085] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described may

[0086] It should be noted that, as another aspect, the present application also provides a storage medium, which can be included in an electronic device, or can exist independently without being assembled into the electronic device. The storage medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the method described in the embodiments. For example, the electronic device can implement each step of the method as shown in Figure 1

[0087] In one embodiment, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the above method embodiments.

[0088] In addition, the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not for limiting purposes. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of the processes. In addition, it is also easy to understand that the processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0089] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the present application cover any and all variations of the application that come within the scope of the claims and a concept of the application. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the application being indicated by the following claims.

[0090] It should be understood that the present application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present application. The scope of the present application is limited only by the appended claims.​

Claims

1. A pursuit-evasion game control method for multi-agent inter-satellite link attack under incomplete information, characterized in that, The method includes: A pursuit-escape game model is constructed, involving the pursuit of the leader and the escapee. Furthermore, the Hamilton-Jacobi-Bellman partial differential equation is used to solve the Nash equilibrium of the pursuit-escape game model, determining the control strategies of the pursuer and the escapee, as well as the reference state of the pursuer. The pursuit-escape game model is based on the relative motion dynamics model of the pursuer and the escapee, and a zero-sum cost function is constructed based on minimizing the pursuit cost and maximizing the escape cost. A cooperative following model between a pursuer and a pursuer leader is constructed; and a control strategy for the pursuer is obtained by solving the cooperative following model using a distributed state observer and an adaptive attack compensator; wherein, the following cost function of the cooperative following model is constructed based on the reference state of the pursuer leader and the position vector difference between the pursuers and the pursuers. Based on the control strategies for pursuing the leader and the escapee, and the control strategies for pursuing the followers, a multi-agent inter-satellite link attack control strategy is formed. The pursuit and escape game control strategy is used to control the multi-pursuing agents to cooperate in pursuing the escaped agent and the escaped agent to evade the multi-pursuing agents.

2. The method according to claim 1, characterized in that, The construction of the pursuit-escape game model for the leader and the escapee includes: Based on the motion model of the pursuer and the escapee, and combined with the relative state variables of the pursuer and the escapee, a relative motion dynamics model is established; Based on the relative motion dynamics model, a zero-sum cost function for the pursuit-escape game model is constructed. The zero-sum cost function includes the pursuit cost function and the escape cost function. The pursuit cost function is constructed based on the relative state variable error, the control cost of pursuing the leader, and the control cost of the escapee. The escape cost function is the dual expression of the pursuit cost function.

3. The method according to claim 2, characterized in that, The zero-sum cost function includes: The pursuit cost function includes: The escape cost function includes: in, This refers to the start time of the movement; For the state variables in the pursuit-escape game model; The system state weights for the pursuit-escape game model; Control variables for tracking the leader; The proportion of energy consumed in pursuing the leader in the pursuit cost function; For the escapee, the control variable; The proportion of energy consumed by the escapee in the pursuit cost function.

4. The method according to claim 2, characterized in that, The construction of a collaborative following model between followers and leaders includes: Establish a first directed graph of communication topology between pursuers and followers, and a second directed graph of communication topology between the pursuer leader and pursuers and followers; Based on the communication weights and position vector differences between adjacent followers in the first communication topology directed graph, the reference state of the leader, and the communication weights between followers and the leader in the second communication topology directed graph, a following cost function is constructed.

5. The method according to claim 4, characterized in that, The following cost function includes: in, ; To pursue the followers The sum of position vector differences between it and other pursuers with which it has a communication connection; The communication weight between adjacent followers; To pursue the followers With the pursuit of the leader The position vector difference between them; The weight of communication between followers and leaders; This is the gain matrix that follows the cost function.

6. The method according to claim 4, characterized in that, The method of solving the cooperative following model using a distributed state observer and an adaptive attack compensator to obtain the control strategy for the pursuer includes: By using a distributed state observer, the state information of non-adjacent pursuers is estimated to obtain a global state estimate. Based on the adaptive attack compensator, the global initial state estimate is compensated for the attack to obtain the corrected global state estimate. The corrected global state estimate and the reference state of the leader are input into the follower cost function. The follower cost function is solved by using the dynamic equation of the follower as a constraint, and the control strategy of the follower is obtained.

7. The method according to claim 1, characterized in that, The adaptive attack compensator includes: in, It is a positive definite matrix; These are estimates of the auxiliary state variables; This is for estimating the upper bound of the injection attack signal; Let be any positive definite continuous bounded function.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the method according to any one of claims 1 to 7.

9. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Optimal state consistency control method for multi-agent system

    CN112445132A

  • Analytic solution method and system for spacecraft finite time pursuit game control

    CN114911167A

  • Pursuit game control method and system based on multi-spacecraft inter-satellite attack

    CN116800467A

  • Optimal capturing method under multi-spacecraft pursuit game based on reinforcement learning

    CN117332684A

  • Multi-agent system control method, device and equipment

    CN119045377A