Adaptive dynamic programming-based spacecraft game control method and device under incomplete information

By combining adaptive dynamic programming and extended state observer with neural networks and superquadratic surface models, the time constraint and safety issues in spacecraft game control under incomplete information conditions are solved, and safe and efficient spacecraft game control in complex environments is achieved.

CN120756676APending Publication Date: 2025-10-10BEIHANG UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510951180.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Under conditions of incomplete information, there are problems of time constraints, safety and strategy unknowns in spacecraft game control, which are difficult to effectively solve with existing technologies. Especially in complex space environments, the dynamic behavior of spacecraft is affected by multiple factors and cannot effectively avoid collision risks.

Method used

An adaptive dynamic programming method is used to construct a fixed-time dilated state observer and the Ada-delta algorithm, combined with a neural network and a superquadratic surface model, to achieve real-time estimation of the unknown strategies and disturbances of the escaping spacecraft, and to design a safety-constrained game strategy to ensure near-optimal game control within a fixed time.

Benefits of technology

It improves the robustness and safety of spacecraft game control, ensures the achievement of mission objectives in complex environments, avoids collision risks, meets time constraints, and improves strategy convergence efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120756676A_ABST
    Figure CN120756676A_ABST
Patent Text Reader

Abstract

The invention discloses a spacecraft game control method and device under incomplete information based on adaptive dynamic programming, and belongs to the technical field of spaceflight. The method comprises the following steps: establishing a spacecraft pursuit game relative motion model, and estimating an unknown strategy and spatial disturbance of a non-cooperative target through a fixed time expansion state observer; constructing an incomplete information pursuit game model based on the estimation state, designing an adaptive dynamic programming game strategy, and approaching an optimal game strategy by using a neural network; an Ada-delta method is introduced to adaptively adjust the learning rate of the neural network, so that the convergence efficiency is improved; an escape spacecraft shape model is constructed based on a super-quadric surface, and collision avoidance is ensured in combination with a potential function; a time constraint problem is solved through a fixed time compensation item, and an approximate optimal game strategy in fixed time is realized. According to the method, the problems of time constraint and strategy optimization of spacecraft game control under incomplete information are effectively solved while the safety is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of aerospace technology, and in particular relates to a spacecraft game control method and device under incomplete information based on adaptive dynamic programming. Background Art

[0002] Non-cooperative targets present numerous challenges for controller design due to their uncertainties, including unknowable maneuvering information, uncoordinated maneuvering behavior, and lack of state communication. Abstracting tasks such as on-orbit approach and interception of non-cooperative targets into an on-orbit pursuit-and-escape game problem, a feasible solution is to use game theory to determine the strategy between the pursuing spacecraft and the escaping spacecraft under incomplete information. In this spacecraft pursuit-and-escape game, an extreme scenario during the approach of a non-cooperative target is considered, in which the pursuing and escaping spacecraft have opposite objectives. The pursuing spacecraft aims to quickly approach the escaping spacecraft to capture or rendezvous with it, while the escaping spacecraft's primary objective is to avoid the pursuing spacecraft.

[0003] In existing research, a widely studied method for solving optimal game strategies is to establish quadratic differential game strategies and use the algebraic Riccati equation to solve the optimal game strategy. The zero-sum game assumption lacks engineering feasibility, relying on the knowledge of the opponent's behavior strategy and real-time status, which is the so-called "complete information assumption." All participants can access the status and strategies of other participants. However, in real-world scenarios, information is often partially available, meaning it is incomplete. Participants may not be able to accurately understand the intentions and decisions of others. In this case, game strategies based on the zero-sum game assumption cannot achieve the optimal decision based on partial information and may generate inappropriate strategies.

[0004] During the game, close proximity encounters were not considered, posing safety risks associated with collisions. For maneuverable spacecraft, this poses a more serious safety risk than cooperative rendezvous and docking, or the progressive rendezvous of a disabled spacecraft. In complex space environments, spacecraft dynamic behavior can be affected by multiple factors, including orbital perturbations, system failures, and external interference. Therefore, safety must be a primary consideration when designing and implementing close-proximity game strategies between spacecraft.

[0005] Few algorithms for spacecraft game problems consider mission time constraints. In actual game tasks, the total duration of a tracking spacecraft's approach cannot be infinite. Due to the complexity of the space environment and the urgency of the mission, the tracking spacecraft must complete rendezvous, docking, or intervention operations on a disabled spacecraft within a limited time. This time constraint not only affects strategy selection but also the probability of mission success. The spacecraft should be able to dynamically adjust its approach strategy based on time requirements and mission objectives, adaptively adjusting its strategy to the ever-changing game environment. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a spacecraft game control method under incomplete information based on adaptive dynamic programming. For the time constraints and safety constraints in the game problem, a fixed-time adaptive dynamic programming game strategy is constructed. For the problems of unknown strategies and unknown spatial disturbances of non-cooperative targets, estimation and step size are achieved by establishing a fixed-time extended state observer, and the learning rate of the neural network weights is adaptively adjusted through the Ada-delta algorithm.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A spacecraft game control method under incomplete information based on adaptive dynamic programming, the method comprising:

[0009] Step 1: Establish a relative motion model for the spacecraft pursuit-escape game to describe the relative motion state between the pursuit spacecraft and the escape spacecraft;

[0010] Step 2: Design a fixed-time dilated state observer to estimate the unknown behavior strategy and spatial disturbance of the escaping spacecraft, and construct incomplete information pursuit-escape game models for both the pursuing and escaping spacecraft based on the estimated state and the actual state.

[0011] Step 3: Based on the estimator of the fixed-time dilated state observer, the objective functions of the pursuit and escape parties are constructed. By constructing the Hamilton-Jacobi-Bellman equation and fitting the optimal solution of the HJB equation using a neural network, the spacecraft game strategy based on adaptive dynamic programming is obtained.

[0012] Step 4: Use the Ada-delta method to adjust the learning rate of the neural network;

[0013] Step 5: Construct an outer shape model of the escape spacecraft based on the superquadratic surface, and construct an artificial potential field function based on this;

[0014] Step 6, the artificial potential function, the fixed time compensation term and the spacecraft game strategy of adaptive dynamic programming are combined, a spacecraft game strategy based on fixed time adaptive dynamic programming is obtained based on an incomplete information pursuit game model.

[0015] In another aspect, the application provides a spacecraft game control device based on adaptive dynamic programming under incomplete information, comprising:

[0016] A motion model construction unit is configured to establish a relative motion model of the spacecraft pursuit game, and describe the relative motion state between the tracking spacecraft and the escaping spacecraft.

[0017] An observer construction unit is configured to design a fixed time extended state observer, estimate the unknown behavior strategy of the escaping spacecraft and the space disturbance, and construct incomplete information pursuit game models of the tracking spacecraft and the escaping spacecraft based on the estimated state and the actual state, respectively.

[0018] A preliminary strategy acquisition unit is configured to construct target functions of the tracking spacecraft and the escaping spacecraft based on the estimated value of the fixed time extended state observer, construct a Hamilton-Jacobi-Bellman equation, and obtain a spacecraft game strategy based on adaptive dynamic programming by fitting the optimal solution of the HJB equation using a neural network.

[0019] An adjustment unit is configured to adjust the learning rate of the neural network using an Ada-delta method.

[0020] A function construction unit is configured to construct an external shape model of the escaping spacecraft based on a hyperquadric surface, and construct an artificial potential function based on the external shape model.

[0021] A final strategy acquisition unit is configured to combine the artificial potential function, the fixed time compensation term and the spacecraft game strategy of adaptive dynamic programming, and obtain a spacecraft game strategy based on fixed time adaptive dynamic programming based on an incomplete information pursuit game model.

[0022] In a third aspect, the application provides an electronic device, comprising: one or more processors; a memory configured to store one or more programs; and wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned spacecraft game control method based on adaptive dynamic programming under incomplete information.

[0023] In a fourth aspect, the application provides a computer readable storage medium having stored executable instructions, which when executed by a processor, enable the processor to implement the aforementioned spacecraft game control method based on adaptive dynamic programming under incomplete information.

[0024] The application has the following beneficial effects:

[0025] The present invention first establishes a relative motion model for the spacecraft pursuit-escape game under the condition of complete information availability, describing the relative motion state between the pursuing and escaping spacecraft. By establishing a fixed-time dilated state observer, real-time estimation of the unknown strategies and spatial perturbations of non-cooperative targets is achieved, improving the robustness of the system in uncertain environments. Based on the estimated state, an incomplete information pursuit-escape game model is constructed, respectively, to construct relative motion and dynamic models for the pursuing and escaping spacecraft based only on the available information. An adaptive dynamic programming game strategy is designed, and an adaptive planning method based on neural networks is designed for both parties in the pursuit-escape game; the Ada-delta method is introduced to adaptively update the learning rate of the gradient descent method, and the adaptive learning rate update of different parameters is realized by dynamically calculating the exponential decay average of historical gradient information, thereby improving the strategy convergence efficiency. It is combined with the adaptive dynamic programming algorithm to adjust its adaptive rate and avoid the oscillation problem in weight learning; the shape of the escaping spacecraft is modeled based on the superquadratic surface, and the shape of the target can be fitted by adjusting the superquadratic parameters. A potential function is introduced to ensure that the tracking spacecraft will not enter the interior of the superquadratic surface, thereby ensuring safety requirements during the game mission; the fixed-time compensation term is combined with the potential function, and combined with the adaptive dynamic programming game strategy, a fixed-time adaptive dynamic programming game strategy is developed that has the performance of guaranteeing the convergence of relative position and velocity states within a fixed time. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a logic diagram of a spacecraft game control method under incomplete information based on adaptive dynamic programming of the present invention;

[0027] Figure 2 Schematic diagram of simulating spacecraft shape based on superquadratic surfaces. DETAILED DESCRIPTION

[0028] The present invention will be further described below with reference to the accompanying drawings and examples.

[0029] This paper addresses incomplete information game scenarios. First, a relative motion model of a complete information spacecraft pursuit-escape game is constructed. An extended state observer (ESO) is used to estimate the unknown perturbations and the escaping spacecraft's strategy within the model. The ESO infers the system's internal state and external disturbances through real-time monitoring of the system's inputs and outputs. Integrating this with fixed-time control theory, a fixed-time ESO is constructed. Relying solely on the inputs and outputs of the relative position states, the ESO estimates relative velocity information, the escaping spacecraft's control variables and perturbations, and the tracking spacecraft's own perturbation information. The algorithm's estimation error is guaranteed to converge to near zero within a fixed time. Based on this, incomplete information pursuit-escape game models are constructed for the tracking and escaping spacecraft, respectively, based on the estimated and actual states. Subsequently, an adaptive dynamic programming method is used to construct the payoff functions of the pursuing and escaping spacecraft. The Hamilton-Jacobi-Bellman (HJB) equation is established, and an evaluative neural network is introduced for online learning to approximate the optimal game strategy of the HJB equation. To address the problem of approaching an escaping spacecraft within a restricted zone, a superquadratic surface-based restricted zone is used to describe the non-cooperative escaping spacecraft and the equipment attached to it. A potential function is then established to prevent collisions between the tracking spacecraft and the escaping spacecraft. By incorporating the potential function into the tracking spacecraft game strategy, a near-optimal game strategy is achieved while meeting safety requirements during the rendezvous process. Furthermore, to ensure the convergence and stability of the algorithm, the Ada-delta method is introduced to adaptively update the learning rate of the gradient descent method. This method uses the mean squared gradient of past iterations to update the current gradient, thus maintaining stability even with large gradient fluctuations. For the tracking spacecraft, a fixed-time compensation term is introduced into the adaptive dynamic programming method, allowing the game strategy to achieve a near-optimal non-cooperative target approach strategy within a fixed time while avoiding collisions during the mission.

[0030] Step 1: A relative motion model of spacecraft pursuit-escape game based on two-body dynamics was established. By transforming the coordinate system, the relative motion between spacecraft was described in the tracking spacecraft system, and a nonlinear relative motion dynamics model was constructed. The parameters set included:

[0031] ,

[0032] Where, and The tracking and escaping spacecraft are respectively in the geocentric inertial coordinate system The position vector under is the Earth's gravitational constant; , The tracking and escaping spacecraft are respectively in the geocentric inertial coordinate system Under the control, and The tracking and escaping spacecraft are respectively in the geocentric inertial coordinate system external disturbances under and are the masses of the pursuit and escape spacecraft, respectively. The superscript ·· denotes the second-order time derivative. Represents the magnitude of a vector.

[0033] The relative motion model in the complete information scenario is rotated to the tracking spacecraft system, which can be described as:

[0034] ,

[0035] in, and They are the desired position and velocity of the tracking spacecraft; , are the expected position and velocity of the escaping spacecraft, It is the tracking spacecraft in the Earth-centered inertial coordinate system Velocity vector under The coordinate system of the tracking spacecraft is To the escape spacecraft's coordinate system The rotation matrix of , They are the pursuit and escape spacecraft in their respective systems and Under the control, The geocentric inertial coordinate system To the tracking spacecraft body coordinate system The rotation matrix of The geocentric inertial coordinate system To the escape spacecraft's coordinate system The rotation matrix of To track the angular velocity of the spacecraft; for any variable Defining common symbols ; Define auxiliary variables and To simplify the description of the formula, the superscript · represents the first-order time derivative.

[0036] Step 2: Establish a relative motion model between the two parties in the spacecraft pursuit game under an incomplete information scenario, construct a fixed-time dilated state observer, and estimate the unknown disturbances of the escaping spacecraft information and orbital state in the game problem; among them, only using relative position information can achieve the estimation value convergence to the true value within a fixed time.

[0037] The relative motion model of the tracking spacecraft in the complete information scenario of this system is simplified and can be described as:

[0038] ,

[0039] The following auxiliary variables are defined to simplify the description of the formula, including , , , .

[0040] The designed fixed-time dilated state observer is as follows:

[0041] ,

[0042] in, is the observer estimation error, Is for The estimated value of ; To simplify the representation, define auxiliary variables , ,in, , , is the error exponential weight parameter that affects the convergence of observations, ; For intermediate parameter variables ,have . is a linear weight parameter that affects the convergence of the observer.

[0043] Based on the estimated state and the actual state, the incomplete information pursuit-escape game model of the two spacecraft is constructed respectively. The relative motion model of the escaping spacecraft is based on the actual partial state information and can be written as:

[0044] ,

[0045] The incomplete information pursuit-escape game model based on observation information for tracking spacecraft can be written as:

[0046] ,

[0047] Among them, auxiliary variables are defined to simplify the formula description, including: , represents a zero matrix of dimension 3×3, Represents the identity matrix of dimension 3×3;

[0048] .

[0049] Step 3. Based on the estimator of the fixed-time dilated state observer, the objective function of the pursuit-escape game is constructed. By solving the Hamilton-Jacobi-Bellman (HJB) equation, an adaptive dynamic programming game strategy based on the evaluation neural network is obtained to obtain the approximate optimal game strategy for the pursuit-escape game problem.

[0050] Approximately optimal game strategy for tracking spacecraft It can be expressed as:

[0051] ,

[0052] in, is a symmetric positive definite weight matrix, is the ideal weight vector of the neural network, is the selected neural network basis function, Representation basis function Estimated value The partial derivative of .

[0053] Step 4: The Ada-delta method is used to address the difficulty of selecting the learning rate during the neural network learning process. By introducing historical gradient information, rapid adaptive adjustment of the learning rate can be achieved to avoid falling into local optimality or oscillation at the equilibrium point.

[0054] ,

[0055] ,

[0056] ,

[0057] in, is the adaptive learning rate, Indicates the current time The gradient shift parameter, Indicates the last step time The gradient shift parameter, whose initial value is , Indicates the current time The historical gradient information parameter, Indicates the last calculation The historical gradient information parameter, whose initial value is ; It is a minimum value, which ensures the stability of the algorithm; is the decay factor that controls the exponential moving average, is the gradient of the adaptation rate.

[0058] Step 5: Construct a shape model of the escaping spacecraft based on the superquadric surface, and construct a potential function to ensure collision avoidance between the tracking spacecraft and the escaping spacecraft;

[0059] like Figure 2 As shown in the figure, for the target shape, a cuboid is constructed based on the super quadratic surface. and cylinders The no-fly zone is formed to enclose the escaping spacecraft:

[0060] ,

[0061] Then construct the artificial potential function In the following form:

[0062] ,

[0063] in, represents the relative distance between the tracking spacecraft and the center of mass of the escaping spacecraft, and the shape parameters of the superquadric surface: and It is a positive constant and is related to the size of the target.

[0064] Step 6. Establish a fixed-time compensation term to solve the time constraint problem of the game strategy. Combine the artificial potential field with the fixed-time compensation term and the adaptive dynamic programming method to obtain the spacecraft feedback game strategy. Adaptive update of the control law is achieved through value function approximation and strategy iterative optimization, while ensuring that the closed-loop system converges within a fixed time.

[0065] The approximate optimal game strategy with fixed-time output feedback is expressed as:

[0066] ,

[0067] Among them, the fixed time compensation term , is an auxiliary variable The generalized inverse variable of , , is the weight coefficient; time weight parameters include: , , , ; Auxiliary variables , switching function , auxiliary variables , auxiliary variables , auxiliary variables , auxiliary variables , symbolic function .

[0068] On the other hand, the present invention provides a spacecraft game control device under incomplete information based on adaptive dynamic programming, which includes various units capable of executing various steps of the aforementioned method, specifically including:

[0069] The motion model building unit is used to establish the relative motion model of the spacecraft pursuit-escape game, describing the relative motion state between the pursuit spacecraft and the escaping spacecraft;

[0070] The observer construction unit is used to design a fixed-time dilated state observer, estimate the unknown behavior strategy and spatial disturbance of the escaping spacecraft, and construct the incomplete information pursuit game model of the pursuing and escaping spacecraft based on the estimated state and the actual state;

[0071] The preliminary strategy acquisition unit is used to construct the objective functions of both the pursuer and the fugitive based on the estimated value of the fixed-time dilated state observer. By constructing the Hamilton-Jacobi-Bellman equation and fitting the optimal solution of the HJB equation using a neural network, the spacecraft game strategy based on adaptive dynamic programming is obtained.

[0072] An adjustment unit for adjusting the learning rate of the neural network using the Ada-delta method;

[0073] A function construction unit is used to construct an outer shape model of the escape spacecraft based on a superquadric surface, and to construct an artificial potential field function based on this;

[0074] The final strategy acquisition unit is used to combine the artificial potential field function, the fixed-time compensation term and the spacecraft game strategy of adaptive dynamic programming, and obtain the spacecraft game strategy based on fixed-time adaptive dynamic programming based on the incomplete information pursuit and escape game model.

[0075] In a third aspect, the present invention provides an electronic device comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned spacecraft game control method under incomplete information based on adaptive dynamic programming.

[0076] In a fourth aspect, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the aforementioned spacecraft game control method under incomplete information based on adaptive dynamic programming.

[0077] The above-described specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application, and it should be understood that the above-described is only a specific embodiment of the present application and is not intended to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A spacecraft game control method under incomplete information based on adaptive dynamic programming, characterized in that: The method comprises: Step 1: Establish a relative motion model for the spacecraft pursuit-escape game to describe the relative motion state between the pursuit spacecraft and the escape spacecraft; Step 2: Design a fixed-time dilated state observer to estimate the unknown behavior strategy and spatial disturbance of the escaping spacecraft, and construct incomplete information pursuit-escape game models for both the pursuing and escaping spacecraft based on the estimated state and the actual state. Step 3: Based on the estimator of the fixed-time dilated state observer, the objective functions of the pursuit and escape parties are constructed. By constructing the Hamilton-Jacobi-Bellman equation and fitting the optimal solution of the HJB equation using a neural network, the spacecraft game strategy based on adaptive dynamic programming is obtained. Step 4: Use the Ada-delta method to adjust the learning rate of the neural network; Step 5: Construct an outer shape model of the escape spacecraft based on the superquadratic surface, and construct an artificial potential field function based on this; Step 6: Combine the artificial potential field function, the fixed-time compensation term, and the spacecraft game strategy of adaptive dynamic programming, and obtain the spacecraft game strategy based on fixed-time adaptive dynamic programming based on the incomplete information pursuit-escape game model.

2. The spacecraft game control method under incomplete information based on adaptive dynamic programming according to claim 1 is characterized in that: The relative motion model of the spacecraft pursuit-escape game in step 1 is: , Where, and The tracking and escaping spacecraft are respectively in the geocentric inertial coordinate system The position vector under is the Earth's gravitational constant; , The tracking and escaping spacecraft are respectively in the geocentric inertial coordinate system Under the control, and The tracking and escaping spacecraft are respectively in the geocentric inertial coordinate system external disturbances under and are the masses of the pursuing and escaping spacecraft, respectively. The superscript ·· denotes the second-order time derivative. represents the magnitude of a vector; The relative motion model is rotated to the local system of the tracking spacecraft.

3. The spacecraft game control method under incomplete information based on adaptive dynamic programming according to claim 2 is characterized in that: In step 2, the fixed-time extended state observer is designed as follows: , in, is the observer estimation error, For variables The estimated value of , , are the expected position and velocity of the escaping spacecraft respectively; define auxiliary variables , ,in, , , is the error exponential weight parameter that affects the convergence of observations, ; For intermediate parameter variables ,have , is the linear weight parameter that affects the convergence of the observer; The construction of the incomplete information pursuit and escape game model for both parties of the pursuing and escaping spacecraft includes establishing an incomplete information pursuit and escape game model based on actual partial state information for the escaping spacecraft, and establishing an incomplete information pursuit and escape game model based on observation information for the tracking spacecraft.

4. The spacecraft game control method under incomplete information based on adaptive dynamic programming according to claim 3 is characterized in that: Step 3 is based on the spacecraft game strategy of adaptive dynamic programming for: , in, is a symmetric positive definite weight matrix, is the ideal weight vector of the neural network, is the selected neural network basis function, Representation basis function Estimated value The partial derivative of , represents a zero matrix of dimension 3×3, Represents the identity matrix of dimension 3×3.

5. The method for controlling a spacecraft under incomplete information based on adaptive dynamic programming according to claim 4, characterized in that: The step 4 includes implementing rapid adaptive adjustment of the learning rate by introducing historical gradient information.

6. The method for controlling a spacecraft under incomplete information based on adaptive dynamic programming according to claim 5, characterized in that: The step 5 includes constructing a rectangular parallelepiped based on the super quadratic surface for the target shape. and cylinders The no-fly zone is formed to enclose the escaping spacecraft: , Then construct the artificial potential function In the following form: , in, , represents the relative distance between the tracking spacecraft and the center of mass of the escaping spacecraft, is the desired position of the tracking spacecraft, The coordinate system of the tracking spacecraft is To the escape spacecraft's coordinate system The rotation matrix and superquadric shape parameters are: and It is a positive constant related to the target's dimensions.

7. The method for controlling a spacecraft under incomplete information based on adaptive dynamic programming according to claim 6, characterized in that: In step 6, the spacecraft game strategy based on fixed-time adaptive dynamic programming is: , in, is a fixed time compensation item.

8. A spacecraft game control device under incomplete information based on adaptive dynamic programming, characterized in that: include: The motion model building unit is used to establish the relative motion model of the spacecraft pursuit-escape game, describing the relative motion state between the pursuit spacecraft and the escaping spacecraft; The observer construction unit is used to design a fixed-time dilated state observer, estimate the unknown behavior strategy and spatial disturbance of the escaping spacecraft, and construct the incomplete information pursuit game model of the pursuing and escaping spacecraft based on the estimated state and the actual state; The preliminary strategy acquisition unit is used to construct the objective functions of both the pursuer and the fugitive based on the estimated value of the fixed-time dilated state observer. By constructing the Hamilton-Jacobi-Bellman equation and fitting the optimal solution of the HJB equation using a neural network, the spacecraft game strategy based on adaptive dynamic programming is obtained. An adjustment unit for adjusting the learning rate of the neural network using the Ada-delta method; A function construction unit is used to construct an outer shape model of the escape spacecraft based on a superquadric surface, and to construct an artificial potential field function based on this; The final strategy acquisition unit is used to combine the artificial potential field function, the fixed-time compensation term and the spacecraft game strategy of adaptive dynamic programming, and obtain the spacecraft game strategy based on fixed-time adaptive dynamic programming based on the incomplete information pursuit and escape game model.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; Wherein, when one or more programs are executed by the one or more processors, the one or more processors implement the spacecraft game control method under incomplete information based on adaptive dynamic programming as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that Executable instructions are stored thereon, which, when executed by a processor, enable the processor to implement the spacecraft game control method under incomplete information based on adaptive dynamic programming as described in any one of claims 1-7.

Citation Information

Cited By

  • A spacecraft rendezvous control method for a tumbling non-cooperative target

    CN122501549A