A spacecraft rendezvous control method for a tumbling non-cooperative target

CN122501549APending Publication Date: 2026-08-04HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-07-07
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0012]本发明的目的是为了克服现有技术的不足,提出一种面向翻滚非合作目标的航天器交会控制方法,重点解决翻滚目标场景下参考轨迹、策略观测、奖励评价和椭球禁飞区安全评价可能因姿态估计偏差而采用不同几何基准的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122501549A_ABST
    Figure CN122501549A_ABST
Patent Text Reader

Abstract

This invention discloses a spacecraft rendezvous control method for non-cooperative targets undergoing tumbling. Using the same target spacecraft's intrinsic reference sequence as a common intermediate benchmark, it constructs both estimated dynamic reference states and actual dynamic reference states, achieving functional separation between the estimated and actual target spacecraft state channels in the training closed loop. The estimated dynamic reference state is used for policy observation vector construction and deployment phase control, while the actual dynamic reference state is used for reward calculation, success determination, and safety evaluation during the training phase. This mechanism avoids training and evaluation distortion caused by different target attitude benchmarks for reference trajectory generation, policy observation vectors, and safety evaluations. Furthermore, it improves the consistency of safe rendezvous control strategy training and deployment under conditions where target tumbling, state estimation bias, and ellipsoidal no-fly zone constraints coexist.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spacecraft relative motion control technology, specifically a spacecraft rendezvous control method for non-cooperative targets that tumble. Background Technology

[0002] The rapid development of missions such as on-orbit servicing, defunct spacecraft cleanup, and active debris removal has created an urgent need for the ability to track spacecraft for safe rendezvous with non-cooperative targets. Unlike traditional rendezvous with cooperative targets, the attitude of tumbling non-cooperative targets (such as runaway satellites or rocket debris) is usually unknown or only partially observable. The catchable point, the target itself, and the no-fly zone are all fixed in the target's coordinate system and move as the target tumbles. Therefore, the core challenge of such missions lies in ensuring that, given deviations in target state estimation, the tracking spacecraft can always safely approach the target by consistently focusing on the same set of real geometric regions on the target (i.e., the approach point and no-fly zone in the target's own coordinate system), rather than simply reaching a fixed point in the inertial frame.

[0003] To address the aforementioned challenges, existing technologies have explored various approaches, which can be broadly categorized as follows:

[0004] Guidance methods based on robust and adaptive control: For example, CN108287476B discloses an autonomous rendezvous guidance method based on high-order sliding mode control and a disturbance observer. This type of method suppresses model uncertainties and external disturbances by designing robust control laws, ensuring the convergence of relative motion. However, its focus is on compensating for deviations at the dynamic level, without establishing a mechanism to ensure that the reference trajectory, policy observation, and safety assessment all revolve around the same geometric object on the target's own system. When the target rolls and attitude estimation errors exist, there may be a misalignment between the reference trajectory generation benchmark and the true geometric boundary of the safety assessment.

[0005] Obstacle avoidance methods based on potential functions or safe corridors: For example, CN116443275A achieves safe obstacle avoidance rendezvous by constructing potential functions for the target and obstacles. Such methods work well when dealing with static or well-defined obstacles. However, in scenarios involving rolling, non-cooperative targets, the no-fly zone is not a static region fixed in the LVLH (Local Vertical and Local Horizontal) coordinate system, but rather a dynamic geometric boundary defined within the target's own coordinate system. When there is a deviation in the target attitude estimation, the safe boundary constructed based on the estimated attitude may be inconsistent with the no-fly zone based on the actual attitude, leading to distorted safety assessments.

[0006] Constrained optimization methods based on Model Predictive Control (MPC): For example, the Tube-based Robust Output Feedback MPC method proposed by Dong et al. can systematically handle constraints such as thrust saturation, approach corridors, and collision avoidance. However, these methods rely on solving complex constrained optimization problems online within each control cycle. When the sampling period is short, the no-fly zone changes rapidly with the target, and there are state estimation errors and parameter perturbations, the computational burden of the optimization problem increases dramatically, which may lead to timeouts or non-convergence, making it difficult to meet the real-time requirements of onboard computers.

[0007] Parameter tuning or direct control methods based on reinforcement learning (RL):

[0008] CN110850719A employs reinforcement learning to self-tune the parameters of an existing control law. This method does not alter the structure of the original control law. However, when the target approach point and no-fly zone dynamically change with the target's roll, a fixed-structure control law struggles to effectively incorporate dynamic geometric constraint information into the control loop.

[0009] In their work on asymmetric reinforcement learning (ADSAM), Shao et al. improved the robustness of the policy under partially observable conditions by using real states from simulations to assist the value network learning during the training phase. However, the information asymmetry of this method is mainly reflected at the network input level, and it does not enforce and structurally adopt the same target ontology geometric reference in the three key stages of reference signal generation, observation construction, and evaluation feedback. Therefore, the "good" behavior used for evaluation during training may correspond to different target pose references than the state observed based on the estimated state during the deployment phase, thus affecting the actual deployment effect of the trained policy.

[0010] In summary, existing technologies have studied the rendezvous problem of non-cooperative targets from the perspectives of robust control, obstacle avoidance, constraint optimization, and reinforcement learning. However, none of them have clearly proposed a solution that can systematically address the fundamental problem that "due to attitude estimation bias, reference trajectory, policy observation, reward evaluation, and safety constraints are deployed around different geometric benchmarks." For close-range rendezvous tasks of tumbling non-cooperative targets, the lack of a common intermediate benchmark defined on the target's own framework significantly increases the complexity of the control logic and the risk of mission failure.

[0011] Therefore, how to provide a control method that enables reference generation, observation construction, reward calculation, and security evaluation to all revolve around the same set of dynamic geometric relationships on the target ontology, even when there are biases in the target state estimation, thereby improving the consistency between training evaluation and deployment execution and reducing the online computational burden, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0012] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a spacecraft rendezvous control method for non-cooperative rolling targets. It focuses on addressing the problem that reference trajectories, policy observations, reward evaluations, and ellipsoidal no-fly zone safety evaluations in rolling target scenarios may employ different geometric benchmarks due to attitude estimation deviations. Unlike approach control methods that only configure asymmetric information at the Actor / Critic input level, this invention uses the same target spacecraft's intrinsic reference sequence as a common intermediate benchmark, constructing both estimated dynamic reference states and actual dynamic reference states. This achieves functional separation between the target spacecraft's estimated state channel and actual state channel in the training closed loop. Specifically, the estimated dynamic reference state is used for policy observation vector construction and deployment phase control, while the actual dynamic reference state is used for reward calculation, success determination, and safety evaluation during the training phase. This mechanism avoids training evaluation distortion caused by different target attitude benchmarks for reference trajectory generation, policy observation vectors, and safety evaluations, and improves the consistency of safe rendezvous control strategy training and deployment under the simultaneous presence of target rolling, state estimation deviations, and ellipsoidal no-fly zone constraints.

[0013] To achieve the above objectives, the technical solution specifically adopted by the present invention is as follows:

[0014] A spacecraft rendezvous control method for non-cooperative targets undergoing tumbling includes the following steps:

[0015] Step 1: Establish a training environment: Establish a training environment that includes the relative translational dynamics of the tracking spacecraft, the rolling attitude dynamics of the target spacecraft, the upper limit of the three-axis thrust, and the control sampling period, and define the target spacecraft's approach point and ellipsoidal no-fly zone within the target spacecraft's own system.

[0016] Step 2, Dual-channel state maintenance: Maintain two channels in parallel in the training environment: the actual state of the target spacecraft and the estimated state of the target spacecraft.

[0017] Step 3, Reference Trajectory Generation and Binding: Based on the estimated state of the target spacecraft, the approach point of the target's intrinsic system is transformed into the LVLH coordinate system, generating a reference trajectory in the LVLH coordinate system. Then, this reference trajectory in the LVLH coordinate system is transformed by the inverse rotation of the estimated attitude of the target spacecraft and saved as a reference position sequence and a reference velocity sequence in the target spacecraft's intrinsic system.

[0018] Step 4, Dual-channel dynamic reconstruction: At each control moment, the estimated state and the actual state of the target spacecraft are used to reconstruct the reference position sequence and reference velocity sequence of the same target spacecraft in the same system into the estimated dynamic reference state and the actual dynamic reference state in the LVLH coordinate system.

[0019] Step 5, Policy Observation and Action Output: Construct a policy observation vector using the estimated dynamic reference state, input the policy observation vector into a continuous action reinforcement learning policy network, output a three-dimensional continuous action from the policy network, and map the three-dimensional continuous action as a three-axis thrust applied to the tracking spacecraft.

[0020] Step 6, Dual-channel evaluation and training: Calculate the real tracking error using the real dynamic reference state, calculate the degree of violation of the ellipsoidal no-fly zone using the position of the tracked spacecraft in the target spacecraft's own system corresponding to the real target attitude, and calculate the training reward based at least on the real tracking error and the degree of violation of the no-fly zone to train the reinforcement learning policy network.

[0021] Step 7, Deployment Control: During the deployment phase, the policy observation vector is constructed using only the real-time estimated target spacecraft state. The trained policy network then forward-infers the three-axis thrust to execute control actions.

[0022] Preferably, in the training environment, the relative translational dynamics of the tracking spacecraft are described by the relative motion equations in the near-circular orbit LVLH coordinate system; the tumbling attitude dynamics of the target spacecraft are described by the Euler dynamics equations or quaternion attitude kinematics equations in the target spacecraft's own system.

[0023] Preferably, the estimated state of the target spacecraft is generated by a nominal estimation mode, a fully observed mode, or a noisy estimation mode, and is propagated according to the target tumbling attitude dynamics within the same control cycle as the actual state of the target spacecraft.

[0024] Preferably, in step 3, when generating the reference trajectory in the LVLH coordinate system, its terminal velocity is determined based on the estimated angular velocity of the target spacecraft and the rotational induced velocity generated by the target's approach point as the target rolls.

[0025] Preferably, the reference velocity sequence in the target spacecraft's intrinsic system is generated by subtracting the velocity component in the LVLH coordinate system, which is the cross product of the estimated angular velocity of the target spacecraft and the reference position, from the reference velocity in the LVLH coordinate system. Then, the result is transformed into the target spacecraft's intrinsic system by the transpose of the rotation matrix corresponding to the estimated attitude of the target spacecraft.

[0026] Preferably, in step 4, the reconstruction method for estimating the dynamic reference state is as follows: the reference position sequence under the target spacecraft's own system is transformed to the LVLH coordinate system using the rotation matrix corresponding to the estimated attitude of the target spacecraft to obtain the estimated dynamic reference position; the reference velocity sequence under the target spacecraft's own system is transformed to the LVLH coordinate system using the same rotation matrix, and then the velocity component generated by the cross product of the estimated angular velocity of the target spacecraft and the estimated dynamic reference position in the LVLH coordinate system is added to obtain the estimated dynamic reference velocity; the reconstruction method for the true dynamic reference state is as follows: the reference position sequence under the same target spacecraft's own system is transformed to the LVLH coordinate system using the rotation matrix corresponding to the true attitude of the target spacecraft to obtain the true dynamic reference position; the reference velocity sequence under the same target spacecraft's own system is transformed to the LVLH coordinate system using the same true attitude rotation matrix, and then the velocity component generated by the cross product of the true angular velocity of the target spacecraft and the true dynamic reference position in the LVLH coordinate system is added to obtain the true dynamic reference velocity.

[0027] Preferably, the strategy observation vector includes: the position error of the tracking spacecraft relative to the estimated dynamic reference state, the velocity error of the tracking spacecraft relative to the estimated dynamic reference state, the estimated attitude quaternion of the target spacecraft, the estimated angular velocity and mission progress of the target spacecraft, the position increment between adjacent reference points in the target spacecraft's own system, and the velocity increment between adjacent reference points in the target spacecraft's own system.

[0028] Preferably, the training reward includes at least: a negative value of the true dynamic reference position error, a negative value of the true dynamic reference velocity error, and a penalty for the degree of violation of the ellipsoidal no-fly zone.

[0029] Preferably, the training reward also includes one or more of the following: a position error propulsion reward consisting of the difference between the norm of the true dynamic reference position error between the previous control time and the current control time; a thrust consumption penalty consisting of the norm of the three-axis thrust; a tracking hold reward to encourage the strategy to maintain a close state after successful tracking; and a task progress reward given according to the task completion progress.

[0030] Preferably, the continuous action reinforcement learning policy network is trained using the SoftActor-Critic algorithm; the state input of the SoftActor-Critic algorithm is the policy observation vector, the action output is a three-dimensional continuous action, and the loss function during training includes a policy entropy term, which is used to balance the maximization of cumulative reward and the degree of policy exploration.

[0031] Preferably, when training the reinforcement learning policy network, a policy network, two action value networks, two target action value networks corresponding to the two action value networks, and an experience replay pool are set up. In each control step, the policy observation vector, action, reward, next policy observation vector, and termination flag are stored in the experience replay pool, and a small batch of samples are randomly drawn from the experience replay pool to update the network parameters. The two target action value networks adopt a soft update method, that is, in each training step, the parameters of the current action value network and the parameters of the target action value networks are weighted and averaged.

[0032] Preferably, in the training environment, at the beginning of each training round, at least one of the following parameters is resampled in a preset random domain: the target spacecraft's true attitude or true quaternion, the target spacecraft's true angular velocity, the target spacecraft's state estimation bias, the tracking spacecraft's initial relative position, the tracking spacecraft's initial relative velocity, and the target spacecraft's moment of inertia; within the same training round, the estimated dynamic reference state is constructed from the sampled estimated state, and the true dynamic reference state and the degree of no-fly zone violation are calculated from the sampled true state.

[0033] In addition, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0034] This invention has the following characteristics and beneficial effects:

[0035] 1. A unified geometric benchmark was established, which fundamentally solved the problem of inconsistency between reference, observation and evaluation caused by attitude estimation deviation.

[0036] To address the issue in existing robust guidance or potential function methods where reference trajectory generation, strategy observation construction, and safety assessment may correspond to different target attitude benchmarks, this invention proposes for the first time to use a "target spacecraft intrinsic system reference sequence" as a common intermediate benchmark. The reference trajectory is first bound to the target intrinsic system, and then dynamically expanded from the estimated attitude and the actual attitude into the estimated dynamic reference state and the actual dynamic reference state, respectively. Thus, the estimated benchmark upon which the strategy network observation is based, the actual benchmark upon which the reward calculation depends, and the geometric boundary upon which the ellipsoidal no-fly zone safety assessment is based all originate from the same set of target intrinsic geometric relationships. This mechanism effectively reduces the risk of approach points deviating from the target intrinsic system or no-fly zone assessment misalignment due to attitude estimation bias, improving the consistency between training evaluation and deployment observations.

[0037] 2. It improves the authenticity and reliability of safety assessments and avoids distortion of no-fly zone assessments due to estimation bias.

[0038] To address the shortcomings of potential function or safety corridor methods that treat no-fly zones as static or construct them solely based on estimated attitudes, this invention directly uses the target's own system corresponding to the actual target attitude to calculate the degree of violation of the ellipsoidal no-fly zone during the training phase. This actual violation degree is then used for rewards and safety evaluation. This allows the reinforcement learning strategy to accurately perceive the real geometric boundaries during the learning process, avoiding the problem of inconsistencies between safety evaluations and real no-fly zones caused by attitude estimation biases. This results in training a rendezvous control strategy with high safety even in real physical environments.

[0039] 3. Significantly reduced the online computational burden and improved the real-time performance and engineering feasibility of closed-loop control.

[0040] Compared to methods like TubeMPC, which require online rolling solutions to constrained optimization problems in each control cycle, this invention outputs three-axis thrust through a single forward inference operation performed by the trained policy network during the deployment phase, without requiring any online optimization solutions. Simulation results under the same conditions (see Table A3) show that the average computation time per step for this method is only 1.01 ms, while the TubeMPC method requires 23.65 ms. In close-range rendezvous tasks with limited sampling periods (e.g., 0.5 seconds), this invention can stably and promptly output control commands, avoiding the impact of optimization timeouts or non-convergence on closed-loop real-time performance, making it more suitable for environments with limited onboard computing resources.

[0041] 4. It achieves end-to-end continuous control directly facing the dynamic geometry of the rolling target, exceeding the upper limit of parameter self-tuning capability.

[0042] To address the limitations of parameter self-tuning reinforcement learning methods (such as CN110850719A), which still rely on fixed control law structures and struggle to express the dynamic geometric relationships of the target, this invention enables the policy network to directly output three-dimensional continuous actions and map them to three-axis thrust based on observation vectors such as estimated dynamic reference state, estimated target attitude angular velocity, reference increment, and mission progress. This end-to-end structure allows the policy to directly learn the optimal thrust feedback law under dynamic conditions such as target roll and no-fly zone changes, eliminating the need for manual design of intermediate control laws and providing stronger mission adaptability.

[0043] 5. It improves the consistency between reinforcement learning training and deployment at the structural level, unlike general asymmetric information configuration.

[0044] Unlike Shao et al.'s ADSAM methods, which only configure asymmetric information at the Actor / Critic input level, this invention mandates in its system architecture that the estimated channel and the real channel share the same target system reference sequence, and are dynamically reconstructed into an estimated dynamic reference state (for observation and deployment) and a real dynamic reference state (for training evaluation) at each control time. This structural design ensures that the reward objective in the training phase and the observation objective in the deployment phase originate from the same geometric benchmark, improving the transfer effect from simulation training to actual deployment at the algorithm architecture level and reducing the "simulation-reality" difference.

[0045] 6. It has good generalization ability and can adapt to various tumbling states, estimation errors and model uncertainties.

[0046] This invention introduces random domain perturbations, including target attitude, angular velocity, initial relative state of the tracking spacecraft, state estimation bias, and moment of inertia, into the training rounds. This causes the policy network to learn not the optimal action sequence under a single condition during offline training, but rather a set of feedback control laws under uncertain conditions. Closed-loop simulation results under strong perturbation scenarios show that, with an initial relative distance of approximately 50m and multiple superimposed perturbations, the position error norm remains within approximately 1.2m throughout the entire process and within 0.4m at the end. The velocity error is mainly distributed within the range of approximately 0.05m / s, and no violations of no-fly zones occur, verifying the good generalization adaptability of this invention to complex uncertainties. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the dynamic reference dual-channel training environment structure of the target system in an embodiment of the present invention.

[0048] Figure 2 This is a schematic diagram of the dynamic reference generation process of the target system in an embodiment of the present invention.

[0049] Figure 3 This is a schematic diagram of the closed-loop control process during the deployment phase in an embodiment of the present invention.

[0050] Figure 4 This is a schematic diagram of the three-dimensional closed-loop intersection trajectory and the ellipsoidal no-fly zone in an embodiment of the present invention.

[0051] Figure 5 This is a schematic diagram of the position tracking error curve implemented in a strongly disturbed scenario according to an embodiment of the present invention.

[0052] Figure 6 This is a schematic diagram of the position error norm curve implemented under a strong disturbance scenario in an embodiment of the present invention.

[0053] Figure 7This is a schematic diagram of the speed tracking error curve implemented under a strong disturbance scenario in an embodiment of the present invention. Detailed Implementation

[0054] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0055] A spacecraft rendezvous control method for non-cooperative targets undergoing tumbling, such as... Figure 1 As shown, it includes the following steps:

[0056] Step 1: Establish the training environment. Establish a training environment that includes the relative translational dynamics of the tracking spacecraft, the rolling attitude dynamics of the target spacecraft, the upper limit of the three-axis thrust, and the control sampling period. Define the target spacecraft's approach point and the ellipsoidal no-fly zone within the target spacecraft's own system.

[0057] In this embodiment, an environment for reinforcement learning training is first established. This environment describes the relative translational state of the tracked spacecraft in a near-circular orbit using the LVLH coordinate system, and describes the target spacecraft's tumbling attitude and no-fly zone geometry using the target spacecraft's intrinsic system. It includes a training environment for the tracked spacecraft's relative translational dynamics, the target spacecraft's tumbling attitude dynamics, the three-axis thrust limit, and controlling the sampling period, and defines the target spacecraft's intrinsic approach point and ellipsoidal no-fly zone within the target spacecraft's intrinsic system.

[0058] Specifically, the reference orbital angular velocity is taken as Ω = 1.0454 × 10⁻⁶. -3 rad / s corresponds to a near-circular orbit with a semi-major axis of approximately 7144.8 km. The control sampling period T... s =0.5s, meaning the reinforcement learning policy outputs a three-dimensional action every 0.5s. The physical integration step size is... This allows for the execution of 5 integral substeps within a single control cycle.

[0059] The mass of the tracking spacecraft is taken as 850 kg, used to calculate the relative translational acceleration from the thrust. The moment of inertia of the target spacecraft is represented by the following matrix (unit: kg·m²):

[0060]

[0061] The upper limit of the three-axis thrust is taken as (Newton). The normalized action of the strategy output is limited and multiplied element-wise by the upper limit of the thrust to obtain the actual control input.

[0062] Define an ellipsoidal no-fly zone within the target spacecraft's own system, such as Figure 4 As shown, let the center of the ellipsoid be... (meters), the semi-axial length of the ellipsoid is [17,8,6]m. Define a symmetric positive definite matrix P such that the dimensionless quantity... in To track the spacecraft's position within the target's system. When g < 1, the spacecraft is located within a no-fly zone.

[0063] Furthermore, the relative translational dynamics of the tracking spacecraft are described by the relative motion equations in the near-circular orbit LVLH coordinate system; the tumbling attitude dynamics of the target spacecraft are described by the Euler dynamics equations or quaternion attitude kinematics equations in the target spacecraft's own system.

[0064] set up The LVLH coordinate system represents the near-circular reference orbit. Indicates the target spacecraft's system; tracks the spacecraft in The relative positions and relative velocities in the system are respectively and The reference orbital angular velocity is The mass of the tracked spacecraft is The three-axis thrust is Maintain within one control cycle When the position remains constant, the relative translational dynamics of the tracking spacecraft, expressed using the relative motion equations in the near-circular orbit LVLH coordinate system, are as follows:

[0065]

[0066] In the formula, These represent the three-axis position components of the tracking spacecraft relative to the target reference point in the LVLH coordinate system. These are the corresponding velocity components.

[0067] Equivalently, equation (1) can be written in state-space form:

[0068]

[0069] in

[0070]

[0071] Assume the target spacecraft is relative to the inertial frame. The angular velocity of the target spacecraft system The following is represented as The rotational inertia matrix of the target spacecraft is The target external torque is The target spacecraft's tumbling attitude dynamics are based on the Euler dynamic equations of the target spacecraft's own system:

[0072]

[0073] In a free tumbling implementation with no external torque or negligible external torque, it can be made Then we have:

[0074]

[0075] The target spacecraft attitude quaternions follow the format in the instruction manual. , in For the vector part, The scalar part is used. The quaternion attitude kinematic equations can be expressed as:

[0076]

[0077] In the formula, For the reason Construct the antisymmetric cross product matrix. After each integration update, the cross product matrix is... Normalization is performed to meet the requirements. .

[0078] Step 2: Dual-channel state maintenance and initialization. In the training environment, two channels, the actual state of the target spacecraft and the estimated state of the target spacecraft, are maintained in parallel.

[0079] The training environment maintains two state channels in parallel: the target spacecraft's actual state channel and the target spacecraft's estimated state channel.

[0080] The actual state of the target spacecraft is propagated according to the dynamics of the target's tumbling attitude, such as... Figure 2 As shown, this includes the target's true pose quaternion. and true angular velocity (Represented within the target's own framework). The estimated state of the target spacecraft includes the estimated attitude quaternions. and estimated angular velocity The true state is used to advance the target's true attitude, generate a true dynamic reference state, and conduct a true safety assessment; the estimated state is used to generate a reference trajectory, generate an estimated dynamic reference state, construct strategy observation vectors, and control during the deployment phase. The estimated state of the target spacecraft can be propagated using nominal estimation, full observation, or noisy estimation modes. The nominal estimation mode is used to simulate situations with modeling errors or initial estimation biases, the full observation mode is used for comparative training, and the noisy estimation mode is used to simulate the impact of measurement noise.

[0081] Specifically, the target spacecraft's true state channel is represented as follows:

[0082]

[0083] in, The true attitude quaternion of the target spacecraft. For the target spacecraft relative to the inertial frame The true angular velocity in the target spacecraft system The following indicates, To control the time index.

[0084] Let the rotational inertia matrix of the target spacecraft be... The external torque acting on the target spacecraft in the context of the target spacecraft's own system is expressed as: The true angular velocity of the target spacecraft propagates according to the Euler dynamics equations of the target spacecraft's own system:

[0085]

[0086] In implementations where free tumbling or external torque is negligible, let Then we have:

[0087]

[0088] The target spacecraft's true attitude quaternion uses the format specified in the instruction manual. ,in For the scalar part. The true attitude quaternion of the target spacecraft propagates according to the quaternion attitude kinematics equation:

[0089]

[0090] in, Let be the quaternion kinematics matrix constructed from angular velocities. After each integral update, for Normalization is performed to meet the requirements. .

[0091] The target spacecraft's estimated state channel is represented as follows:

[0092]

[0093] in, Estimate the attitude quaternions for the target spacecraft. Estimate the angular velocity of the target spacecraft within the target spacecraft's own system. The following is an indication.

[0094] In the nominal estimation mode, the estimated state is obtained by propagation from a preset nominal initial value and a nominal dynamic model. In the full observation mode, let:

[0095]

[0096] In the noisy estimation mode, let:

[0097]

[0098]

[0099] in, and These are angular velocity estimation perturbations and attitude estimation perturbations, respectively, which can be generated from a preset perturbation range, perturbation basis matrix, or measurement noise model.

[0100] In the above manner, the real state channel is propagated according to the target roll attitude dynamics and used for real dynamic reference state and safety evaluation, while the estimated state channel is generated in the form of nominal estimation, full observation or noisy estimation and used for reference trajectory generation, estimated dynamic reference state reconstruction, policy observation vector construction and deployment stage control.

[0101] At the beginning of each training epoch, the true state and the estimated state are randomized and initialized. Specifically:

[0102] The target's nominal initial angular velocity is taken as (Corresponding to the approximation) The actual initial angular velocity is randomly sampled near the nominal value within a preset perturbation range.

[0103] The initial quaternion for the target nominal value is [0.3826834, 0, 0, 0.9238795] (in the format of...). The true initial quaternion is randomly sampled and normalized near the nominal quaternion according to a preset perturbation range.

[0104] The angular velocity estimation perturbation basis matrix is The quaternion-estimated perturbation basis matrix is: In each training round, angular velocity perturbations are sampled from a zero-mean random distribution. and attitude perturbation The estimated angular velocity is obtained by superimposing the true values. The attitude estimate is obtained by superimposing the true quaternions. The result was obtained by normalization.

[0105] Step 3: Reference trajectory generation and binding. Based on the estimated state of the target spacecraft, the approach point of the target's own system is transformed into the LVLH coordinate system to generate a reference trajectory in the LVLH coordinate system. Then, the reference trajectory in the LVLH coordinate system is transformed by the reverse rotation of the estimated attitude of the target spacecraft and saved as a reference position sequence and a reference velocity sequence in the target spacecraft's own system.

[0106] The generation of the reference trajectory is divided into three stages: first, it is planned in the LVLH coordinate system; then, it is bound to the target's own coordinate system; and finally, it is unfolded at the control time.

[0107] (1) The reference trajectory is generated in the LVLH coordinate system. The terminal velocity is determined based on the estimated angular velocity of the target spacecraft and the rotational induced velocity generated by the target's own system approach point as the target rolls.

[0108] Specifically, set For the target spacecraft system To the LVLH coordinate system The rotation matrix, Estimate the attitude quaternions for the target spacecraft. The target is the approach point of this system. Based on the estimated attitude... The terminal approach position of the target system's approach point in the LVLH coordinate system is:

[0109]

[0110] set up Estimate the angular velocity of the target spacecraft within the target spacecraft's own system. The following indicates, The reference orbital angular velocity. The angular velocity of the target spacecraft's own system relative to the LVLH coordinate system is expressed in the LVLH coordinate system as follows:

[0111]

[0112] The rotational induced velocity generated by the target's approach point as it rolls is the cross product of this relative angular velocity and the terminal approach position. Therefore, the terminal velocity of the reference trajectory in the LVLH coordinate system is determined as follows:

[0113]

[0114] If Defined as the angular velocity of the LVLH coordinate system relative to the target spacecraft's own system, then The sign before the angular velocity should be negative accordingly. The definition of the direction of angular velocity should remain consistent throughout the same instruction manual.

[0115] After generating the reference trajectory in the LVLH coordinate system, the reference position sequence is obtained. and reference velocity sequence To bind the reference trajectory to the target spacecraft's own system, the target's own system reference position sequence can be represented as:

[0116]

[0117] The reference velocity sequence of the target system can be expressed as:

[0118]

[0119] Estimate angular velocity based on the target Calculate the LVLH terminal velocity generated by the target roll at the approach point:

[0120]

[0121] in The angular velocity of the target system relative to the LVLH coordinate system is expressed in the LVLH coordinate system, and the calculation formula is as follows:

[0122]

[0123] Then, a reference trajectory from the current state of the tracking spacecraft to the aforementioned terminal position and terminal velocity is generated in the LVLH coordinate system, resulting in a reference position sequence. and reference velocity sequence ,in, .

[0124] First, subtract the rotational induced velocity component caused by the target roll from the reference velocity in the LVLH coordinate system, and then... The system is then transformed to the target spacecraft's intrinsic frame. Through this process, the saved reference position and velocity sequences are defined within the target spacecraft's intrinsic frame, and can subsequently be re-deployed into a dynamic reference state in the LVLH coordinate system at each control moment based on the estimated or actual attitude.

[0125] (2) Binding to the target system

[0126] The reference trajectory in the LVLH coordinate system is transformed by the inverse rotation of the estimated attitude and saved as a reference position sequence and a reference velocity sequence in the target's native system:

[0127] Target system reference position sequence:

[0128]

[0129] Target system reference velocity sequence:

[0130]

[0131] in It is the transpose of the rotation matrix, that is, the coordinate transformation matrix from the LVLH coordinate system to the target body system.

[0132] Through the above transformation, the reference trajectory is "bound" to the target body. Even if the target continues to roll, the reference point is still fixed near the same geometric position of the target body.

[0133] Step 4: Dual-channel dynamic reconstruction. At each control moment, the estimated state and the actual state of the target spacecraft are used to reconstruct the reference position sequence and reference velocity sequence of the same target spacecraft in the same system into the estimated dynamic reference state and the actual dynamic reference state in the LVLH coordinate system.

[0134] At each control time k, the reference sequence of the same target system is reconstructed into the estimated dynamic reference state and the actual dynamic reference state in the LVLH coordinate system using the estimated state and the actual state, respectively.

[0135] (1) Estimating dynamic reference state reconstruction

[0136] Estimated dynamic reference position:

[0137]

[0138] Estimated dynamic reference speed:

[0139]

[0140] (2) Real dynamic reference state reconstruction

[0141] Actual dynamic reference location:

[0142]

[0143] Real dynamic reference speed:

[0144]

[0145] Where the true angular velocity corresponds to:

[0146]

[0147] The estimated dynamic reference state is used for policy observation vector construction and deployment phase control, while the real dynamic reference state is used for reward calculation, success determination, and security evaluation during the training phase.

[0148] Step 5: Construction of policy observation vector. The policy observation vector is constructed by estimating the dynamic reference state. The policy observation vector is input into the continuous action reinforcement learning policy network. The policy network outputs three-dimensional continuous actions and maps the three-dimensional continuous actions to three-axis thrust applied to the tracking spacecraft.

[0149] Policy observation vector It contains the following components:

[0150] Position error of the spacecraft relative to the estimated dynamic reference state: According to the position scale Normalization;

[0151] The velocity error of tracking the spacecraft relative to the estimated dynamic reference state: via velocity scale Normalization;

[0152] Target spacecraft attitude estimation quaternion ;

[0153] Target spacecraft estimated angular velocity Angular velocity scale Normalization;

[0154] Position increments between adjacent reference points in the target system: According to the position scale Normalization;

[0155] The velocity increment between adjacent reference points in the target system: via velocity scale Normalization;

[0156] Task progress (Current step count / Total steps).

[0157] Right now:

[0158]

[0159] In this implementation, the observation vector dimension is 20.

[0160] Furthermore, the policy network and action output:

[0161] The continuous action reinforcement learning policy network employs the SoftActor-Critic (SAC) algorithm. (Policy network) With observation vector As input, output three-dimensional continuous motion ∈[-1,1]³. After the motion is limited by the hyperbolic tangent function, it is multiplied element-wise with the maximum thrust of the three axes to obtain the actual three-axis thrust command:

[0162]

[0163] Where ⊙ represents element-wise multiplication.

[0164] Policy network structure: Two fully connected hidden layers, each with 256 nodes, using ReLU activation function. The output layer outputs the mean and log-standard deviation parameters of the actions.

[0165] Step 6: Calculate training reward. Calculate the real tracking error using the real dynamic reference state, calculate the degree of violation of the ellipsoidal no-fly zone using the position of the tracked spacecraft in the target spacecraft's own system corresponding to the real target attitude, and calculate the training reward based at least on the real tracking error and the degree of no-fly zone violation to train the reinforcement learning policy network.

[0166] Training rewards The reward function is calculated based on real dynamic reference states and the degree of violation of real no-fly zones. The design is as follows:

[0167]

[0168] in:

[0169] This represents the actual dynamic reference position error.

[0170] This represents the actual dynamic reference speed error.

[0171] Degree of violation of no-fly zone: ,in To track the spacecraft's position within the target's own system corresponding to the actual target's attitude;

[0172] , , , , , Non-negative reward weights;

[0173] To maintain the tracking indicator function, it takes the value 1 when the task progress is close to completion and the error is within the threshold, and 0 otherwise.

[0174] The items in the above reward function correspond to: actual tracking position error penalty, actual tracking velocity error penalty, position error propulsion reward (encouraging error reduction), thrust consumption penalty, no-fly zone violation penalty, and tracking maintenance reward.

[0175] Further modifications in this embodiment enhance the learning and training process:

[0176] The policy network is trained using the SAC algorithm. SAC is a maximum entropy reinforcement learning method that not only maximizes the cumulative reward during policy training but also introduces a policy entropy term, ensuring that the policy continues to explore necessary actions during the training phase.

[0177] The following network and modules are set up during training:

[0178] A policy network The parameters are ;

[0179] Two Action Value Networks and The parameters are respectively , ;

[0180] The corresponding two target action value networks and ,parameter , Initial and same;

[0181] An experience replay pool with a capacity of .

[0182] Each control step will generate tuples ( , , , , Stored in the experience replay pool, where The termination criteria are (rendezvous completion, timeout, or violation of no-fly zone rules). Furthermore, the experience replay pool is updated per training round: at the start of each round, the current round state is reset while retaining historical samples; when the capacity limit is reached, old samples are removed using a first-in, first-out (FIFO) strategy. This embodiment uses uniform random sampling and does not employ priority experience replay. A round terminates upon rendezvous completion, reaching the maximum number of control steps, or entering a no-fly zone. After each control step, the current policy observation vector is... ,action ,award Next strategy observation vector and termination mark Composition of empirical samples:

[0183]

[0184] And store it in the experience replay pool The experience replay pool capacity is... ,when The number of samples exceeded At the same time, the earliest stored sample is removed according to the first-in, first-out (FIFO) method; the experience replay pool is not emptied between training rounds to retain historical experience from samples in different random domains. This implementation method uses uniform random sampling from... Small batches of samples are extracted, and priority experience replay is not used.

[0185] End mark It can be determined as follows:

[0186]

[0187] The conditions for rendezvous completion can be determined jointly by the true dynamic reference position error, the true dynamic reference velocity error, and the mission progress. For example:

[0188]

[0189] In the formula, This is the position error threshold. For speed error threshold, The minimum number of control steps set to avoid premature termination.

[0190] A small batch (size 256) of samples is randomly drawn from the experience replay pool for network updates.

[0191] Policy network loss function:

[0192]

[0193] in The action obtained by reparameterized sampling, This is the temperature coefficient.

[0194] Action Value Network Loss Function:

[0195]

[0196] The target value is: , In this embodiment, the discount factor is used. Take 0.99.

[0197] The target action value network uses soft updates:

[0198]

[0199] in In this embodiment, the soft update coefficient is used. Pick .

[0200] The training rounds are set to 10,000, with a maximum of 600 steps per round (corresponding to a 300-second mission time). At the start of each training round, the target's true attitude / quaternion, true angular velocity, state estimation bias, initial relative position and velocity of the tracking spacecraft, and target rotational inertia perturbation are resampled within a preset random domain. Within the same round, the estimated dynamic reference state is constructed from the sampled estimated state, while the true dynamic reference state and the degree of no-fly zone violation are calculated from the sampled true state.

[0201] Step 7: During the deployment phase, the policy observation vector is constructed using only the real-time estimated target spacecraft state. The trained policy network then forward-infers the three-axis thrust to execute control actions.

[0202] like Figure 3 As shown, the deployment phase control process is as follows:

[0203] Read or estimate the target spacecraft's attitude quaternion in real time and angular velocity And tracking the relative position of spacecraft and speed ;

[0204] Based on the estimated state, an estimated dynamic reference state is generated according to the methods in steps 3 and 4. , ;

[0205] Constructing the strategy observation vector (Same as step 5);

[0206] Will Input the trained policy network Forward reasoning yields normalized actions ;

[0207] Will Mapped to three-axis thrust commands ;

[0208] Maintain constant thrust within one control cycle, propel the spacecraft relative state, and update the target estimated state (e.g., through a filter).

[0209] Repeat the above steps until the rendezvous task is completed or a timeout occurs.

[0210] To facilitate reproduction, the key parameters of this embodiment are summarized in Tables A1 and A2 below.

[0211] Table A1 provides a set of specific implementation parameters consistent with the reinforcement learning rendezvous control environment. This table is used to illustrate the basic task parameters and the position, velocity, attitude, and angular velocity disturbance parameters that can be selected in strong disturbance implementation scenarios. In actual engineering applications, the values ​​in the table can be adjusted according to the target size, mission distance, propulsion capability, and sensor accuracy. The parameter values ​​do not constitute a limitation on the scope of protection of this invention.

[0212] Table A1 Examples of Specific Implementation Parameters

[0213]

[0214]

[0215] Table A2 Examples of Reinforcement Learning Training and Network Structures

[0216]

[0217] Table A3 Simulation Indicators Comparison with Tube MPC Under the Same Conditions

[0218]

[0219] To verify the technical effectiveness of this invention, a benchmark simulation under the same conditions was established. The comparison results are shown in Table A3: This method and the TubeMPC comparison method use the same mission time, the same ellipsoidal no-fly zone, the same target roll nominal state, and the same sampled initial state and uncertainty perturbation.

[0220] Furthermore, such as Figures 5-7 As shown, in a scenario with strong disturbances (superimposed with initial relative position and velocity disturbances, and amplified attitude and angular velocity estimation disturbances), deployment evaluation indicates that: the position error norm remains within approximately 1.2m throughout the entire process, and within 0.4m at the end; the three-axis velocity errors are mainly distributed within the range of approximately 0.05m / s; and there are no violations of no-fly zones. This verifies the closed-loop feasibility of the present invention under conditions of complex uncertainties.

[0221] It should be noted that those skilled in the art should understand that the specific parameters mentioned above (such as orbital parameters, sampling period, mass, inertia, thrust limit, network structure, SAC hyperparameters, etc.) are for illustrative purposes only and are not intended to limit the invention. While maintaining the core mechanism of "dynamic dual-channel reference to the target system," these parameters can be adjusted according to actual mission requirements. Furthermore, in addition to the SAC algorithm, other continuous action reinforcement learning algorithms (such as TD3, PPO, etc.) can also be used to train the policy network, as long as they maintain the functional separation of the estimated channel and the real channel and share the same target system reference sequence.

[0222] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A spacecraft rendezvous control method for non-cooperative targets undergoing tumbling, characterized in that, Includes the following steps: Step 1: Establish a training environment that includes the relative translational dynamics of the tracking spacecraft, the tumbling attitude dynamics of the target spacecraft, the upper limit of the three-axis thrust, and the control sampling period, and define the approach point and ellipsoidal no-fly zone of the target spacecraft's own system under the target spacecraft's own system; Step 2: Maintain two channels in parallel in the training environment: the actual state of the target spacecraft and the estimated state of the target spacecraft. Step 3: Based on the estimated state of the target spacecraft, transform the approach point of the target's own system to the LVLH coordinate system to generate a reference trajectory in the LVLH coordinate system. Then, through the reverse rotation transformation of the estimated attitude of the target spacecraft, save the reference trajectory in the LVLH coordinate system as a reference position sequence and a reference velocity sequence in the target spacecraft's own system. Step 4: At each control moment, using the estimated state and the actual state of the target spacecraft, the reference position sequence and reference velocity sequence under the same target spacecraft system are reconstructed into the estimated dynamic reference state and the actual dynamic reference state under the LVLH coordinate system. Step 5: Construct a policy observation vector using the estimated dynamic reference state, input the policy observation vector into the continuous action reinforcement learning policy network, output the three-dimensional continuous action from the policy network, and map the three-dimensional continuous action as a three-axis thrust applied to the tracking spacecraft. Step 6: Calculate the real tracking error using the real dynamic reference state, calculate the degree of violation of the ellipsoidal no-fly zone using the position of the tracking spacecraft in the target spacecraft's own system corresponding to the real target attitude, and calculate the training reward based at least on the real tracking error and the degree of violation of the no-fly zone in order to train the reinforcement learning policy network. Step 7: During the deployment phase, the policy observation vector is constructed using only the real-time estimated target spacecraft state. The trained policy network then forward-infers the three-axis thrust to execute control actions.

2. The method according to claim 1, characterized in that, In the training environment, the relative translational dynamics of the tracking spacecraft are described by the relative motion equations in the near-circular orbit LVLH coordinate system; the tumbling attitude dynamics of the target spacecraft are described by the Euler dynamics equations or quaternion attitude kinematics equations in the target spacecraft's own system.

3. The method according to claim 1, characterized in that, The estimated state of the target spacecraft is generated by a nominal estimation mode, a fully observed mode, or a noisy estimation mode, and is propagated according to the target tumbling attitude dynamics in the same control cycle as the actual state of the target spacecraft.

4. The method according to claim 1, characterized in that, In step 3, when generating the reference trajectory in the LVLH coordinate system, its terminal velocity is determined based on the estimated angular velocity of the target spacecraft and the rotational induced velocity generated by the target's approach point as it rolls.

5. The method according to claim 4, characterized in that, The reference velocity sequence in the target spacecraft's intrinsic system is generated by subtracting the velocity component in the LVLH coordinate system, which is the cross product of the estimated angular velocity of the target spacecraft and the reference position, from the reference velocity in the LVLH coordinate system. Then, the result is transformed into the target spacecraft's intrinsic system by the transpose of the rotation matrix corresponding to the estimated attitude of the target spacecraft.

6. The method according to claim 1, characterized in that, Step 4, estimating the dynamic reference state and reconstructing the true dynamic reference state, includes: The reconstruction method of the estimated dynamic reference state is as follows: the reference position sequence under the target spacecraft's own system is transformed to the LVLH coordinate system through the rotation matrix corresponding to the estimated attitude of the target spacecraft to obtain the estimated dynamic reference position; the reference velocity sequence under the target spacecraft's own system is transformed to the LVLH coordinate system through the same rotation matrix, and then the velocity component generated by the cross product of the estimated angular velocity of the target spacecraft and the estimated dynamic reference position in the LVLH coordinate system is added to obtain the estimated dynamic reference velocity; The reconstruction method of the true dynamic reference state is as follows: the reference position sequence under the same target spacecraft body system is transformed to the LVLH coordinate system through the rotation matrix corresponding to the true attitude of the target spacecraft to obtain the true dynamic reference position. The reference velocity sequence under the same target spacecraft body system is transformed to the LVLH coordinate system through the same true attitude rotation matrix. Then, the velocity component generated by the cross product of the true angular velocity of the target spacecraft and the true dynamic reference position in the LVLH coordinate system is added to obtain the true dynamic reference velocity.

7. The method according to claim 1, characterized in that, The strategy observation vector includes: the position error of the tracking spacecraft relative to the estimated dynamic reference state, the velocity error of the tracking spacecraft relative to the estimated dynamic reference state, the estimated attitude quaternion of the target spacecraft, the estimated angular velocity and mission progress of the target spacecraft, the position increment between adjacent reference points in the target spacecraft's own system, and the velocity increment between adjacent reference points in the target spacecraft's own system.

8. The method according to claim 1, characterized in that, The training rewards include at least: negative values ​​of the true dynamic reference position error, negative values ​​of the true dynamic reference velocity error, and penalties for violations of the ellipsoidal no-fly zone.

9. The method according to claim 8, characterized in that, The training rewards include any one of the following: a position error propulsion reward consisting of the difference between the norm of the true dynamic reference position error between the previous control time and the current control time; a thrust consumption penalty consisting of the norm of the three-axis thrust; a tracking hold reward to encourage the strategy to maintain a close approach state after successful tracking; and a task progress reward given according to the task completion progress.

10. The method according to claim 1, characterized in that, The continuous action reinforcement learning policy network is trained using the Soft Actor-Critic algorithm. The state input of the Soft Actor-Critic algorithm is the policy observation vector, and the action output is a three-dimensional continuous action. The loss function during training includes a policy entropy term, which is used to balance the maximization of cumulative reward with the degree of policy exploration.

11. The method according to claim 10, characterized in that, When training the reinforcement learning policy network, a policy network, two action value networks, two target action value networks corresponding to the two action value networks, and an experience replay pool are set up. In each control step, the policy observation vector, action, reward, next policy observation vector, and termination flag are stored in the experience replay pool, and a small batch of samples are randomly drawn from the experience replay pool to update the network parameters. The two target action value networks adopt a soft update method, that is, in each training step, the parameters of the current action value network and the parameters of the target action value networks are weighted and averaged.

12. The method according to claim 1, characterized in that, In the training environment, at the beginning of each training round, at least one of the following parameters is resampled in a preset random domain: the target spacecraft's true attitude or true quaternion, the target spacecraft's true angular velocity, the target spacecraft's state estimation bias, the tracking spacecraft's initial relative position, the tracking spacecraft's initial relative velocity, and the target spacecraft's moment of inertia. Within the same training round, the estimated dynamic reference state is constructed from the sampled estimated state, and the true dynamic reference state and the degree of violation of the no-fly zone are calculated from the sampled true state.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.