Spacecraft health estimation and fault-tolerant control method based on mixed Mama and meta Q learning
By combining hybrid Mamba and meta-Q learning, an HMamba-BCAF health estimation model and fault-tolerant control strategy are constructed. This solves the problems of long sequence modeling and cross-scale feature fusion in spacecraft health estimation and fault-tolerant control, achieving faster convergence speed and higher stability, and improving the fault-tolerant control capability of spacecraft in multi-fault scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF ELECTRONICS SCI & TECH OF CHINA
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing spacecraft health estimation and fault-tolerant control methods struggle to balance long-sequence modeling capabilities with cross-scale feature fusion, leading to unstable feature representations. Furthermore, reinforcement learning control methods lack reward design related to system stability, resulting in slow convergence speeds and unstable policy behavior. The disconnect between health state estimation and control policy makes it difficult to compensate for faults in a timely manner.
We construct an HMamba-BCAF health estimation model by combining hybrid Mamba and meta-Q learning. The hybrid Mamba module enables long-sequence state space modeling and bidirectional cross-scale attention fusion. Combined with the fault-tolerant control strategy of meta-Q learning, we embed system stability and actuator stability into the reward structure and develop an end-to-end closed-loop fusion framework to achieve deep coupling between health state estimation and adaptive control.
It improves the accuracy of fault-tolerant control of spacecraft under multi-fault scenarios, has faster convergence speed and higher stability, and can maintain strong robustness under sensor noise and actuator noise conditions, achieving rapid fault recovery and robust control.
Smart Images

Figure CN121902301A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to spacecraft health estimation and fault-tolerant control methods. Background Technology
[0002] With the increasing complexity and duration of modern spacecraft missions, high reliability, stability, and autonomous fault tolerance have become fundamental requirements for attitude control systems. As core energy output and actuators, actuators operate under high loads and strong disturbances, making them susceptible to structural wear and electronic degradation, leading to thrust decay, bias drift, and nonlinear time-varying faults. If these faults are not diagnosed and compensated for in a timely manner, they can result in attitude deviations, excessive energy consumption, and even mission failure. Furthermore, spacecraft fault signals exhibit time non-stationarity and multi-scale dynamic characteristics, making traditional deep learning models prone to gradient degradation and insufficient long-sequence modeling. Reinforcement learning-based control methods also suffer from slow convergence speeds and unstable learning behavior. Therefore, achieving accurate health state estimation and robust fault-tolerant closed-loop integration has become an urgent technical challenge.
[0003] In recent years, intelligent spacecraft control technology has made significant progress in deep learning-based health state estimation and reinforcement learning-based fault-tolerant control. Existing health state estimation methods can be broadly categorized into model-driven and data-driven approaches. Traditional model-driven observers, such as sliding mode observers and Kalman filters, can provide analytical state estimates, but their performance tends to degrade when dealing with high-dimensional and strongly nonlinear dynamic systems. With increasing mission complexity and data scale, data-driven technologies are gradually becoming mainstream. Neural networks possess powerful feature representation capabilities and can be used for effective feature extraction; convolutional neural networks (CNNs) are commonly used for spatial feature extraction, while recurrent neural networks (RNNs) and their variants (such as LSTM and GRU) are suitable for processing time-series information. Meanwhile, correlation modeling methods such as the CANN framework demonstrate that explicitly describing the nonlinear dependencies between multiple subsystems can significantly improve fault detection capabilities, which is particularly important for spacecraft with strong actuator-sensor coupling. Existing research has proposed a variety of multimodal and privacy-preserving data analysis frameworks. For example, the TSIN method proposed by Xu et al. fuses time series and image features to identify multi-source industrial conditions, the privacy-preserving federated learning framework constructed by Lu et al. is used to improve the unbalanced health estimation of wind turbines, and the DNN-CCA method proposed by Xie et al. can capture the spatiotemporal dependencies of multiple subsystems, thereby significantly improving the overall performance of spacecraft health state estimation.
[0004] Fault-tolerant control (FTC) technology mainly includes two types of schemes: passive fault-tolerant control (PFTC) and active fault-tolerant control (AFTC). PFTC relies on fixed control laws and has limited adaptability to sudden or time-varying faults. In contrast, AFTC can adjust the control strategy online based on health status information, thereby achieving timely fault compensation and performance recovery. Although AFTC has made some progress in the spacecraft field, existing methods generally lack a unified fusion mechanism between health estimation and control strategy, and rely heavily on ground station intervention, which is particularly prominent in deep space missions with long communication delays. To address these issues, some research has attempted to propose improvement schemes. For example, Wang et al. proposed a graph neural network method based on physical constraints for attitude fault-tolerant control under sensor failure conditions; Dong et al. proposed the GQRT offline reinforcement learning framework, which can improve long-term decision-making capabilities in complex mission scenarios, thereby alleviating the shortcomings of current AFTC systems in terms of autonomy and adaptability to some extent.
[0005] Although deep temporal modeling techniques have made some progress in spacecraft fault diagnosis and control, existing methods still have significant limitations in dynamic operating conditions and multi-fault environments, thus affecting their engineering applicability. Zhang et al. proposed a multi-scale Mamba architecture diagnostic model, but its perception-oriented structure struggles to generate the time-stability, cross-scale correlation, and control-related fault features required under non-stationary actuator fault conditions, failing to meet the downstream control requirements for high-quality health features. Liu et al. used reinforcement learning methods for satellite control research, but their reward function did not include system stability and actuator stability information, resulting in slow training convergence speed and unstable policy update process under different fault scenarios. Shao et al. studied a robust fault-tolerant attitude tracking method, but its control structure treats fault estimation and control independently, lacking an end-to-end closed-loop fusion mechanism that integrates perception and decision-making. In summary, the performance of existing data-driven health estimation methods and reinforcement learning-based fault-tolerant control techniques remains limited in long-sequence dependencies and multi-fault scenarios, requiring further improvement.
[0006] Based on the above phenomena, the motivation for proposing this invention can be summarized as follows:
[0007] Existing time series models struggle to simultaneously characterize long-range dependency features and multi-scale actuator failure modes, leading to unstable feature representations required for downstream control. This necessitates a modeling mechanism capable of long-sequence modeling and adaptive cross-scale feature fusion.
[0008] Reinforcement learning-based control methods often lack reward design related to system stability, resulting in slow convergence speed and unstable policy behavior under different actuator failure conditions. Therefore, a meta-learning control policy that is stability-oriented and can be adaptively updated is needed.
[0009] Meanwhile, existing spacecraft fault-tolerant control schemes generally separate health status estimation from control strategies, making it difficult to update strategies and compensate for faults in a timely manner. Therefore, it is necessary to construct an overall framework that integrates health status perception and adaptive strategy optimization to achieve more robust closed-loop fault-tolerant performance. Summary of the Invention
[0010] The purpose of this invention is to address the problem that existing spacecraft health estimation and fault-tolerant control are difficult to balance, resulting in low accuracy of fault-tolerant control. Therefore, this invention proposes a spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning.
[0011] The specific process of the spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning is as follows:
[0012] Step 1: Obtain the rigid body spacecraft operation dataset;
[0013] The rigid body spacecraft operational dataset consists of a time series window. ;
[0014] Time step samples The corresponding actual fault value is ;
[0015] in, This represents a sample of rigid body spacecraft operational data at time step 1. This represents a sample of rigid body spacecraft operational data at time step 2; Indicates time step Samples of operational data from rigid-body spacecraft; Indicates time step Samples of operational data from rigid-body spacecraft; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The actual value of the axial direction actuator offset fault; Indicates time step of The actual value of the axial direction actuator offset fault; Indicates time step of The actual value of the axial direction actuator offset fault; axis, axis, The axis is in the spacecraft's body coordinate system , , The axis, the direction of advancement of the principal axis of symmetry is axis, Axis perpendicular to axis, Axis perpendicular to flat;
[0016] Step 2: Construct a hybrid HMamba-BCAF health estimation model; obtain the trained hybrid HMamba-BCAF health estimation model; the hybrid HMamba-BCAF health estimation model includes: a Hybrid Mamba module and a BCAF module consisting of parallel branches of BIE and GFU;
[0017] Step 3: Obtain the operational data sample of the rigid body spacecraft under test, input the operational data sample of the rigid body spacecraft under test into the trained hybrid HMamba-BCAF health estimation model, the trained hybrid HMamba-BCAF health estimation model outputs the health estimation result, and control compensation is performed on the health estimation result based on the fault-tolerant control method.
[0018] The beneficial effects of this invention are as follows:
[0019] This invention constructs a unified closed-loop method that organically integrates long-sequence fault feature representation, stability-oriented adaptive control strategy, and deep coupling of health state estimation and fault-tolerant control. This invention addresses the challenge of existing technologies in simultaneously addressing three aspects: deep temporal modeling, meta-reinforcement learning, and estimation-control co-optimization. Specifically:
[0020] First, a hybrid HMamba-BCAF health estimation architecture is proposed, combining stable long-sequence state-space modeling with a bidirectional cross-scale attention fusion mechanism. This maintains temporal continuity while effectively integrating multi-source fault-related features across multiple time scales. Second, a fault-tolerant control strategy based on meta-Q learning is constructed, embedding system stability and actuator stability indices into the reward structure to improve the reliability of the control strategy under different fault conditions. A meta-adaptive mechanism enables rapid policy reconstruction. Finally, an end-to-end closed-loop fusion framework is developed, coupling the health state estimation module with an adaptive meta-Q learning controller to achieve continuous state awareness and online policy updates. Experimental results for four types of actuator fault scenarios demonstrate that the framework exhibits faster convergence speed, higher stability, and strong robustness to sensor and actuator noise, improving the accuracy of fault-tolerant control.
[0021] This invention constructs a closed-loop fault management framework for spacecraft missions, deeply integrating the HMamba-BCAF health state estimator with a meta-Q-learning-based fault-tolerant controller to address the issues of insufficient long-sequence modeling capabilities and limited control adaptability. Specifically, the HMamba-BCAF estimator achieves stable characterization of actuator fault dynamics through a hybrid Mamba backbone, bidirectional information interaction, and a multi-scale gating fusion structure; the meta-Q-learning controller utilizes meta-level priors from multiple fault scenarios, enabling rapid reconfiguration of the control strategy under new faults or external disturbances. Experimental results under constant faults, time-varying faults, and noisy multi-scenario conditions demonstrate that this closed-loop framework achieves faster convergence, higher stability, and stronger fault recovery capabilities compared to baseline methods. Overall, the proposed framework significantly enhances the collaborative capabilities of spacecraft from perception to control, laying the foundation for cross-domain generalization and collaborative control in more complex mission environments in the future.
[0022] The stability and safety of spacecraft systems largely depend on accurate health status estimation and reliable fault-tolerant control techniques. Traditional estimation algorithms have inherent limitations, including insufficient ability to model long-term features, low sensitivity to early or weak faults, and a tendency to generate redundant or unstable feature representations. These shortcomings reduce diagnostic reliability, hinder subsequent fault-tolerant control, and lead to delays in compensation and handling processes. Therefore, there is an urgent need to construct an advanced estimation algorithm framework capable of outputting stable and reliable diagnostic features to support the implementation of active fault-tolerant control. To address the above technical problems, this invention proposes a health status estimation framework (HMamba-BCAF) that combines a hybrid Mamba module with bidirectional cross-scale attention fusion. The hybrid Mamba module enhances the extraction capability of long-term dependencies, achieving stable temporal feature modeling; the bidirectional cross-scale attention mechanism improves the expressive power of fine-grained fault features; and the multi-scale adaptive gating fusion unit further balances strong and weak feature responses, improving the discriminative performance of health status estimation. Simultaneously, this invention also proposes a corresponding active fault-tolerant control strategy. This strategy is designed based on the meta-Q learning mechanism. By combining stability evaluation metrics and explicit stability constraint functions, rapid control strategy reconstruction and compensation under fault conditions are achieved. Extensive experimental results based on various representative thruster fault scenarios demonstrate that the proposed HMamba-BCAF estimator and meta-Q learning fault-tolerant controller can achieve model-independent rapid fault recovery and high control stability, thus validating the effectiveness of the integrated fault-tolerant framework. Attached Figure Description
[0023] Figure 1 This is a flowchart of the present invention; Figure 2 Network diagram for hybrid Mamba modules; Figure 3 This is a schematic diagram of a bidirectional cross-scale attention fusion module; Figure 4 A diagram of the meta-Q learning framework; Figure 5 The graph shows the health status estimation results of HMamba-BCAF, with the horizontal axis representing the round and the vertical axis representing health factors. Figure 6 The graph shows the comparison of actuator stability in round 1 under disturbance conditions between the fault-tolerant control method based on meta-Q learning and the traditional Q learning method; the horizontal axis represents the round number, and the vertical axis represents the actuator stability. Figure 7 The graph shows the comparison of the system stability of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under disturbance conditions in round 1; the horizontal axis represents the round number, and the vertical axis represents the system stability. Figure 8 The figure shows the comparison results of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under disturbance conditions in round 2. Figure 9 The figure shows a comparison of the system stability of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under disturbance conditions in round 2. Figure 10 The figure shows the comparison results of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under disturbance conditions in round 3. Figure 11 The figure shows the comparison of the system stability of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under disturbance conditions in round 3. Figure 12 The figure shows a comparison of the actuator stability in round 4 under disturbance conditions between the fault-tolerant control method based on meta-Q learning and the traditional Q learning method. Figure 13 The figure shows the comparison of the system stability of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under disturbance conditions in round 4. Figure 14 The figure shows a comparison of the fault-tolerant control performance of the meta-Q learning-based fault-tolerant control method and the traditional Q learning method under the influence of injected Gaussian noise in round 1. Figure 15 The figure shows a comparison of the fault-tolerant control performance of the meta-Q learning-based fault-tolerant control method and the traditional Q learning method under the influence of injected Gaussian noise in round 1. Figure 16 The figure shows a comparison of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under the influence of injected Gaussian noise in round 2. Figure 17 The figure shows a comparison of the fault-tolerant control method based on meta-Q learning and the traditional Q learning method under the influence of injected Gaussian noise in round 2. Figure 18 The figure shows a comparison of the fault-tolerant control performance of the meta-Q learning-based fault-tolerant control method and the traditional Q learning method under the influence of injected Gaussian noise in round 3. Figure 19 The figure shows a comparison of the fault-tolerant control performance of the meta-Q learning-based fault-tolerant control method and the traditional Q learning method under the influence of injected Gaussian noise in round 3. Figure 20 The figure shows a comparison of the fault-tolerant control performance of the meta-Q learning-based fault-tolerant control method and the traditional Q learning method under the influence of injected Gaussian noise in round 4. Figure 21 The figure shows a comparison of the fault-tolerant control performance of the meta-Q learning-based fault-tolerant control method and the traditional Q learning method under the influence of injected Gaussian noise in round 4. Detailed Implementation
[0024] Specific implementation method one: Combining Figure 1 This implementation method, based on a hybrid Mamba and meta-Q learning approach for spacecraft health estimation and fault-tolerant control, is described as follows:
[0025] Step 1: Obtain the rigid body spacecraft operation dataset; the rigid body spacecraft operation dataset consists of a time series window. Time step samples The corresponding actual fault value is ;
[0026] in, This represents a sample of rigid body spacecraft operational data at time step 1. This represents a sample of rigid body spacecraft operational data at time step 2; Indicates time step Samples of operational data from rigid-body spacecraft; Indicates time step Samples of operational data from rigid-body spacecraft; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The actual value of the axial direction actuator offset fault; Indicates time step of The actual value of the axial direction actuator offset fault; Indicates time step of The actual value of the axial direction actuator offset fault; axis, axis, The axis is in the spacecraft's body coordinate system , , The axis, the direction of advancement of the principal axis of symmetry is axis, Axis perpendicular to axis, Axis perpendicular to flat;
[0027] Step 2: Construct a hybrid HMamba-BCAF health estimation model; obtain the trained hybrid HMamba-BCAF health estimation model; the hybrid HMamba-BCAF health estimation model includes: a Hybrid Mamba module and a BCAF module consisting of parallel branches of BIE and GFU;
[0028] Step 3: Obtain the operational data sample of the rigid body spacecraft under test, input the operational data sample of the rigid body spacecraft under test into the trained hybrid HMamba-BCAF health estimation model, the trained hybrid HMamba-BCAF health estimation model outputs the health estimation result, and control compensation is performed on the health estimation result based on the fault-tolerant control method.
[0029] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that, in step one, the time step... samples Includes six-dimensional input ;
[0030] in, This represents the dynamic residual vector calculated from the angular velocity measurement and the inertia matrix. Components along the axial direction; This represents the dynamic residual vector calculated from the angular velocity measurement and the inertia matrix. Components along the axial direction; This represents the dynamic residual vector calculated from the angular velocity measurement and the inertia matrix. Components along the axial direction; ;in, Represents the dynamic residual term of the system. The inertia matrix, Angular velocity is the variable. The control torque output by the actuator. These parameters, which are external disturbances that vary over time, together form the basis of the system's dynamic response under fault and disturbance conditions. This indicates the desired control torque output by the controller. Components along the axial direction; This indicates the desired control torque output by the controller. Components along the axial direction; This indicates the desired control torque output by the controller. Components along the axial direction; axis, axis, The axis is in the spacecraft's body coordinate system , , The axis, the direction of advancement of the principal axis of symmetry is axis, Axis perpendicular to axis, Axis perpendicular to flat.
[0031] The other steps and parameters are the same as in Specific Implementation Method 1.
[0032] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that the working process of the hybrid HMamba-BCAF health estimation model in step two is as follows:
[0033] Rigid body spacecraft operation dataset Each sample Input a Hybrid Mamba module, output characteristics of the Hybrid Mamba module ;feature Input to the BCAF module, the BCAF module outputs samples. Corresponding predicted fault value Until the rigid body spacecraft operational dataset is completed. The health estimates of all samples are processed to obtain a trained hybrid HMamba-BCAF health estimation model. Other steps and parameters are the same as in implementation method one or two.
[0034] Specific Implementation Method Four: This implementation method differs from one of the specific implementation methods one to three in that the HybridMamba module includes: a first normalized Norm layer, a first convolutional layer, a first SiLU activation function layer, a first fully connected Linear layer, a state space model SSM, a learnable gated convolution operator, a second fully connected Linear layer, and a second SiLU activation function layer.
[0035] The learnable gated convolution operator includes: a second convolutional layer, a third SiLU activation function layer, a third convolutional layer, and a sigmoid activation function layer. For example... Figure 2 As shown. Other steps and parameters are the same as in any of the specific implementation methods one to three.
[0036] Specific Implementation Method Five: This implementation method differs from one of Specific Implementation Methods One to Four in that the rigid spacecraft operation dataset in step two... Each sample Input a Hybrid Mamba module, output characteristics of the Hybrid Mamba module The specific process is as follows:
[0037] sample Input to the first normalized Norm layer, output features from the first normalized Norm layer. ;feature Input to the first convolutional layer, output features from the first convolutional layer ;feature The first SiLU activation function layer is input, and the first SiLU activation function layer outputs features. ;feature Input to the first fully connected Linear layer, output features from the first fully connected Linear layer. ;feature Input State-Space Model (SSM), Output Features of State-Space Model (SSM) ;feature Extracting local mutation features from learnable gated convolution operators ;
[0038] feature and characteristics Perform weighted fusion to obtain features ; indicates as: ;
[0039] feature Input to a second fully connected Linear layer; output features from the second fully connected Linear layer. ;
[0040] feature The input is the second SiLU activation function layer, and the output of the second SiLU activation function layer is the feature. ;
[0041] Features and samples Perform weighted fusion to obtain features ; indicates as:
[0042]
[0043] feature As an output feature of the Hybrid Mamba module.
[0044] The other steps and parameters are the same as those in one of the specific implementation methods one to four.
[0045] Specific Implementation Method Six: This implementation method differs from one of Specific Implementation Methods One to Five in that the described feature... Extracting local mutation features from learnable gated convolution operators Specifically:
[0046] feature The input is to the second convolutional layer, and the output of the second convolutional layer is to the third SiLU activation function layer, and the output of the third SiLU activation function layer is to the third SiLU activation function layer. ;feature Input to the third convolutional layer, the third convolutional layer outputs features. Input to the sigmoid activation function layer, the sigmoid activation function layer outputs features. ; Features and characteristics Perform element-wise dot product to obtain the features. ;
[0047]
[0048] in, The input feature sequence; and These are the weights of the convolutional layer; and This refers to the bias term of the convolutional layer; This is an element-wise pointwise multiplication operation. The other steps and parameters are the same as in any of the specific implementation methods one to five.
[0049] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One to Six in that the BCAF module includes: a second normalized Norm layer, a third fully connected Linear layer, a first ReLU activation function layer, a third normalized Norm layer, a fourth fully connected Linear layer, a second ReLU activation function layer, a cascaded Concat layer, a fifth fully connected Linear layer, a sixth fully connected Linear layer, a first feedforward neural network (FFN), a second feedforward neural network (FFN), a first multi-head attention mechanism model, a seventh fully connected Linear layer, a first pooling layer, a second pooling layer, a third pooling layer, a fourth pooling layer, an eighth fully connected layer, a first softmax activation function layer, a ninth fully connected layer, a second softmax activation function layer, and a fusion module. (Example) Figure 3 As shown.
[0050] The other steps and parameters are the same as those in one of the specific implementation methods one to six.
[0051] Specific Implementation Method Eight: This implementation method differs from one of Specific Implementation Methods One to Seven in that the described feature... Input to the BCAF module, the BCAF module outputs samples. Corresponding predicted fault value ;
[0052] The specific process is as follows:
[0053] 1) Output characteristics of the Hybrid Mamba module The input is the second normalized Norm layer, and the output of the second normalized Norm layer is the feature. ;feature Input to a third fully connected Linear layer; output features from the third fully connected Linear layer. ;feature Input to the first ReLU activation function layer, output features from the first ReLU activation function layer ;
[0054] Hybrid Mamba module output characteristics Input to the third normalized Norm layer, output features from the third normalized Norm layer. ;feature Input to the fourth fully connected Linear layer, output features from the fourth fully connected Linear layer. ;feature The input is a second ReLU activation function layer, and the output of the second ReLU activation function layer is a feature. ;
[0055] feature and characteristics Perform concatenation to obtain features. ;feature Input to the fifth fully connected Linear layer, output features from the fifth fully connected Linear layer. ;
[0056] feature Input to the sixth fully connected Linear layer, output features from the sixth fully connected Linear layer. ;
[0057] Features and characteristics By performing element-wise summation, we obtain the characteristics. ;feature Input to the first feedforward neural network (FFN), the first feedforward neural network (FFN) outputs features ;
[0058] feature The second feedforward neural network (FFN) is input, and the second feedforward neural network (FFN) outputs features. ;
[0059] feature Input to the first multi-head attention mechanism model, the first multi-head attention mechanism model outputs features ;
[0060] feature Input to the first multi-head attention mechanism model, the first multi-head attention mechanism model outputs features ;
[0061] Features and characteristics Perform concatenation to obtain features. ;feature Input to the seventh fully connected Linear layer, output features from the seventh fully connected Linear layer. ;
[0062] 2) Output characteristics of the Hybrid Mamba module Input to the first pooling layer (Pooling), output features from the first pooling layer (Pooling). ; for features Perform a 1x downsampling to obtain features ;feature Input to the second pooling layer (Pooling), output features from the second pooling layer (Pooling). ; for features Perform 0.5x downsampling to obtain features. Hybrid Mamba module output characteristics Input to the third pooling layer (Pooling), output features from the third pooling layer (Pooling). ; for features Perform a 1x downsampling to obtain features ;feature Input to the fourth pooling layer (Pooling), output features from the fourth pooling layer (Pooling). , for features Perform 0.5x downsampling to obtain features. ;feature The layers sequentially pass through the eighth fully connected layer and the first softmax activation function layer. The first softmax activation function layer outputs the features. ;feature The layers sequentially pass through the ninth fully connected layer and the second softmax activation function layer. The second softmax activation function layer outputs the features. ; Features and characteristics Input to the fusion module, output features from the fusion module ;
[0063] 3) Features With features Weighted fusion is performed to obtain the output samples of the BCAF module. Corresponding predicted fault value ; indicates as:
[0064] .
[0065] The other steps and parameters are the same as those in any of the specific implementation methods one to seven.
[0066] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One to Eight in that the feature... and characteristics Input to the fusion module, output features from the fusion module The specific process is as follows:
[0067] Features ,feature ,feature and characteristics The difference, and the characteristics of fusion Composition sequence ;Change the sequence The input layers are sequentially a linear layer, a ReLU activation function layer, another linear layer, and a Sigmoid activation function layer. The Sigmoid activation function layer outputs features. Based on features ,feature and characteristics , to obtain features ; indicates as:
[0068]
[0069] in, and Represents forward and backward features; The difference represents the feature; Linear layer; This is a linear layer used to generate the gating weights and transformation results required for fusion. Other steps and parameters are the same as in specific implementation methods one through eight.
[0070] Specific Implementation Method Ten: This implementation method differs from Specific Implementation Methods One through Nine in that, in step three, the operational data sample of the rigid-body spacecraft under test is obtained. This operational data sample is then input into the trained hybrid HMamba-BCAF health estimation model. The trained hybrid HMamba-BCAF health estimation model outputs a health estimation result, and control compensation is applied to the health estimation result based on a fault-tolerant control method. The specific process is as follows:
[0071] 1) Output health estimation results for the trained hybrid HMamba-BCAF health estimation model. Perform meta-learning to obtain Meta-learning initial parameters at time step ;
[0072] 2) Order ;
[0073] 3) Obtain state of time ; state of time It is expressed as follows:
[0074]
[0075] in, express angular velocity at time t, express Scalar components of the quaternion error at time step; Represents the L2 norm, Represents absolute value;
[0076] 4) Execution state of time Get Moment of action ;
[0077] 5) Based on state of time and Moment of action conduct Value evaluation, obtained The beginning of time function , and Momentary Rewards The specific process is as follows:
[0078]
[0079]
[0080]
[0081] in, express The state at any given moment, express Actions at any moment express Initial parameters at time; express Initial parameters at time;
[0082] , , , , , express function;
[0083] express Momentary rewards; Indicates basic environmental rewards; express System stability indicators at any given time; express Actuator stability metrics at any given time; This indicates the reward items for successful task completion; The weighting coefficients represent the system stability index; The weighting coefficients represent the stability index of the actuator;
[0084] in
[0085]
[0086]
[0087]
[0088] in, The variance representing the rate of change of angular velocity; The modulus representing the average angular velocity; The weighting coefficients represent the system stability index; The weighting coefficients represent the stability index of the actuator; This represents the stability component extracted from the scalar term in the attitude error quaternion; The mean square term representing the actuator output torque; coefficient and These correspond to the attitude stability weight factor and the energy smoothness weight factor, respectively, and are used to adjust the relative importance of attitude convergence and torque change stability in the reward function; This indicates a successful terminal reward. and These represent the threshold values for attitude quaternions and angular velocity, respectively.
[0089] 6) Based on Optimal action at any given moment Get state of time , ;in, Optimal action at any given moment The superscript T indicates transpose.
[0090] 7) Execution state of time Get Moment of action ;
[0091] 8) Based on state of time and Moment of action conduct Value evaluation, obtained Moment function , and Momentary Rewards The specific process is as follows:
[0092]
[0093]
[0094]
[0095] in, express Parameters at time; express Parameters at time; , express Moment Value function value; , express Moment Value function value; express Momentary rewards; Indicates basic environmental rewards; express System stability indicators at any given time; express Actuator stability metrics at any given time; This indicates the reward items for successful task completion; The weighting coefficients represent the system stability index; The weighting coefficients represent the stability index of the actuator;
[0096] 9) Based on Moment Value function value and Correction Moment Value function value and The specific process is as follows:
[0097]
[0098] in, Indicates the learning rate; Indicates the discount factor;
[0099] , express Moment Value function value;
[0100]
[0101]
[0102] express The optimal action at any given moment;
[0103]
[0104] 10) Order ;
[0105] Repeat steps 6) through 9), using the Reptile optimization strategy to adjust the parameters. and Update until the rigid body spacecraft operational dataset is complete. All samples The control compensation is used to obtain the trained hybrid HMamba-BCAF health estimation model and the parameters of the corresponding Q-learning module. and .like Figure 4 As shown.
[0106] The other steps and parameters are the same as those in any of the specific implementation methods one to nine.
[0107] Reptile optimization strategy affects parameters and The update is formally expressed as follows:
[0108]
[0109] in, Indicates globally shared parameters. For the inner loop learning rate, Update the step size for the outer loop element. Indicates in the task Below Value loss function.
[0110] Problem Statement: Spacecraft actuators are prone to performance degradation, bias drift, and external disturbances under complex mission conditions, which significantly impact system stability. To achieve an organic combination of health state estimation and fault-tolerant control compensation in dynamic environments, this invention, based on the concept of dynamic residual modeling, systematically analyzes the closed-loop fusion mechanism between estimation and control to ensure the consistency and real-time performance of fault perception and control strategy updates over time.
[0111] When actuator output is correlated with control commands, the system will experience continuous efficiency degradation and additional bias in a specific control direction, leading to long-term performance deterioration and cumulative control error in that direction. In contrast, external disturbances are transient and irregular, and their effects are independent of control commands. This invention summarizes and concludes the satellite dynamics and kinematic equations under actuator failure and external disturbance conditions, providing a theoretical basis for subsequent health state estimation and fault-tolerant control design. .in, Represents the dynamic residual term of the system. The inertia matrix, Angular velocity is the variable. The control torque output by the actuator. These parameters, representing time-varying external disturbances, collectively form the basis of the system's dynamic response under fault and disturbance conditions. The attitude error quaternion is defined as... Its kinematic evolution process satisfies the corresponding quaternion update law, which is used to describe the dynamic characteristics of attitude error changing over time; ;in, Represents the quaternion projection matrix;
[0112] In this invention, known actuator efficiency losses and unknown external disturbances are decoupled through system dynamics equations, thereby forming fault characteristics that are more conducive to neural network learning; the fault characteristics are: ;
[0113] in, Represents the identity matrix. The actuator fault distribution matrix, Indicates the nominal control torque. As inputs for fault compensation control, the aforementioned variables are used together to describe the control allocation and compensation mechanism under fault conditions. This invention evaluates four typical fault scenarios, including: (1) simultaneous occurrence of two fault forms; (2) time-varying efficiency degradation under constant additive bias; (3) time-varying bias under constant efficiency degradation; and (4) simultaneous time-varying efficiency loss and additive bias. In actual operating environments, spacecraft actuators may suffer from sudden faults or external disturbances, accompanied by performance degradation that gradually accumulates over time during the mission. The time-varying fault types considered in this invention can effectively reflect the dynamic characteristics and uncertainties of actuators under real operating conditions, thereby providing a more representative fault model basis for health status estimation and fault-tolerant control strategy design.
[0114] In this invention, fault torque scaling, bias drift, and external disturbances are all modeled as dynamic processes that obey exponential recovery characteristics; fault injection follows a physical mechanism combining instantaneous impact and exponential recovery. For any actuator When the fault occurs in discrete time steps When it occurs, its residual thrust ratio Offset torque and external disturbance vector for:
[0115]
[0116] in, Indicates the minimum remaining thrust ratio. This represents the initial amplitude of the bias fault. It is a time constant. The initial peak vector represents the external disturbance. The parameters mentioned above collectively describe the characteristics of the fault or disturbance at the initial moment and its recovery over time. Based on the above fault description, this invention characterizes the actuator state in a compact health factor form, defined as: .
[0117] In summary, HMamba-BCAF can output health factors in real time, while meta-Q learning is used to continuously optimize the control strategy, thus forming a highly integrated closed-loop framework, which significantly improves both convergence stability and overall robustness.
[0118] III. Closed-loop fault-tolerant control health status estimation theory:
[0119] A. Health Status Estimation: 1. Hybrid Mamba with Bidirectional Cross-Scale Attention Fusion; To address the efficiency decay and bias drift issues of spacecraft actuators, this invention proposes a bidirectional cross-scale attention fusion hybrid Mamba framework for intelligent health status estimation. This model uses a hybrid Mamba structure as the backbone for time series modeling and introduces a bidirectional information exchange (BIE) module and a gated fusion unit (GFU) to jointly construct a bidirectional cross-scale attention fusion (BCAF) structure, achieving multi-level feature interaction and multi-scale adaptive fusion. The model output integrates the features from both branches through a weighted fusion mechanism to generate the final health status estimation result. This architecture can provide stable feature representations for subsequent fault-tolerant control, comprehensively capturing long-range dependencies, bidirectional time consistency, and adaptive multi-scale dynamic features. In this framework, the hybrid Mamba module is used to extract temporal features from sequence data including angular velocity, quaternions, and actuator commands. This module combines the long-range filtering capability of the State-Space Model (SSM) with a gated convolutional structure, thereby achieving a more expressive global modeling of fault signals. The resulting temporal representations are then processed in parallel by the BIE branch and the GFU branch to capture richer cross-scale interaction information. Specifically, the BIE module achieves bidirectional interaction between features by constructing a multi-head gated channel between the forward and backward information flows; its adaptive weights... and The intensity of information exchange is adjusted based on contextual relevance to ensure consistency over time. The GFU module utilizes a multi-scale gating matrix. Adaptive alignment is performed on the feature scale, and simultaneously, the difference term is adjusted. With interactive items Modeling is employed to capture transient and steady-state characteristics during actuator failure. Joint processing of these two parameters effectively improves the stability of feature representations under sudden disturbances. In health state estimation, angular velocity, quaternions, and command torque are used to calculate dynamic residuals and fed as inputs into a hybrid Mamba backbone network for time encoding. Subsequently, the BIE and GFU modules perform bidirectional feature interaction and cross-scale fusion processing. The fused representation is then passed to a discriminative estimation unit to obtain an estimate of the actuator efficiency loss. and bias fault estimation value This invention enables real-time detection and regression of multiple types of faults. Through cross-scale dual-channel modeling, the method improves temporal continuity and fault discrimination capabilities, obtaining high-precision health status estimation results. Formally, the mapping relationship between temporal features and efficiency decay and bias fault estimates is expressed as follows:
[0120]
[0121] in, and These represent the weights and bias parameters of the output layer, respectively. , and They represent the times respectively. Angular velocity, attitude quaternion, and actuator control commands; state Features including attitude error and angular velocity are used to characterize the system at time steps. To improve the rapid recovery capability and smooth response performance of the meta-Q-learning control strategy under random actuator failures and external disturbances, this invention designs a reward function based on system stability and actuator stability indices. Unlike methods that directly rely on attitude error or angular velocity, this scheme constructs a system stability index... This transforms the system's convergence trend and the actuator's smoothness into learnable reinforcement signals, thereby achieving adaptive optimization of control performance. System stability is assessed through angular velocity sequences. The incremental variance and mean are quantified, and the actuator stability is determined by the variance of adjacent control torque changes, the mean of control torque, and the attitude error quaternion. Conduct an evaluation. System stability indicators. Composed of two parts, the first part reflects the dynamic smoothness of the system by suppressing angular velocity fluctuations; the second part characterizes the convergence capability of the attitude by constraining the average level of angular velocity, thus jointly achieving a comprehensive quantification of system stability. Compared with traditional reward design methods that rely on instantaneous errors, the stability index construction method proposed in this invention can dynamically capture the system's response characteristics to disturbances, thereby enhancing the temporal correlation between reward signals and control behavior, enabling the control strategy to obtain a more stable and consistent optimization direction in the continuous time domain; during control execution, the system stability index... Used to evaluate the overall convergence characteristics and steady-state smoothness of a spacecraft, its calculations simultaneously consider attitude error, angular velocity, and their rates of change, thus reflecting the system's dynamic response capability under disturbances; complementary to this, actuator stability indices... This metric characterizes the actuator's response quality under control commands, especially in the event of faults or external disturbances. It describes the smoothness of actuator behavior by measuring the magnitude of changes in the control input, the trend of control effort, and the consistency between the attitude error quaternion and the control deviation. The synergistic effect of these two stability metrics allows the reward function to simultaneously promote dynamic stability at the system level and robust operation at the actuator level, thereby achieving more reliable attitude control performance within a closed-loop fault-tolerant control framework. This reward mechanism enables the system to achieve rapid response in the early stages of training, while gradually increasing the focus on overall system stability and precise actuator control in later stages. Through this adaptive balancing approach, the training process maintains effective control performance under various disturbance conditions, further enhancing the robustness and reliability of the proposed algorithm. By enhancing the dynamic tracking accuracy of the actuator during attitude adjustment, the control system can converge quickly even with minor deviations. In the steady-state phase, this stability term penalizes unnecessary control torque to reduce energy consumption and mechanical load, thereby guiding the actuator to perform tasks in a more consistent, reliable, and long-lasting manner.
[0122] Closed-Loop System Analysis: During spacecraft operation, the proposed HMamba-BCAF module serves as a front-end health state estimator, forming a closed-loop structure with a meta-Q-learning-based fault-tolerant controller. This detection module continuously outputs estimates of actuator efficiency losses and biases, which the controller uses to perform adaptive compensation and update the control strategy within defined limits. This closed-loop interaction mechanism effectively suppresses the effects of faults and external disturbances, thereby maintaining the stability of attitude control. By combining deep temporal modeling with adaptive fault-tolerant control, this framework provides real-time adaptability and scalability for spacecraft autonomous health management, exhibiting higher robustness compared to traditional observer-based reconstruction methods. The overall closed-loop structure is as follows: Figure 4 As shown.
[0123] Example:
[0124] A. Spacecraft System Dataset Description and Evaluation Index Analysis; Spacecraft attitude motion follows strongly coupled nonlinear dynamics, where actuator efficiency losses, bias drift, and external disturbances can all weaken control capabilities and affect the stability of the rigid body system. Within this dynamic framework, health state estimation is used to identify actuator performance degradation in real time; fault-tolerant control adaptively adjusts the command torque based on these estimations to maintain attitude stability. Together, they constitute the closed-loop mechanism necessary for reliable spacecraft operation. This dataset originates from a rigid body spacecraft simulation model, injecting actuator efficiency losses into attitude dynamics. Additive bias fault Generation. Each sample consists of a time-series window containing six-dimensional input. ; Indicates at time step The input time series vector. Its corresponding regression objective is the fault value. The dataset contains both constant-value faults and time-varying faults, thus providing diverse training modes for health status estimation. The source of this dataset can be found in the references. At time steps... The details of the network analysis are as follows:
[0125] To evaluate the overall performance of the proposed health state estimation network and fault-tolerant control method, its hyperparameter configuration is shown in Table 1. To verify the accuracy of the health state estimation model in fault numerical estimation, this invention selects mean absolute error (MAE) and mean square error (MSE) as evaluation metrics. Their mathematical definitions are shown below.
[0126] Table 1. Main hyperparameter configurations for health status estimation
[0127]
[0128] in, Indicates the actual fault value;
[0129] This represents the estimated fault value.
[0130] In this invention, parameters The number of parameters (millions), model size (kb), and inference time (ms) are used as comprehensive evaluation metrics for health status estimation networks. The HMamba-BCAF estimator and the meta-Q-learning controller are integrated into the same closed-loop system and validated in a dynamic spacecraft simulation environment that includes rigid body attitude dynamics and exponentially decaying actuator failures. The system uses historical attitude data, dynamic residual signals, and command torque as inputs to generate real-time efficiency loss and bias estimates. The controller then uses these estimates to perform adaptive compensation and update the strategy, thereby maintaining the stable operation of the spacecraft.
[0131] B. Health Status Assessment Experiments: 1. Ablation Experiments: To verify the superior performance of the proposed HMamba-BCAF model in the health status estimation task, a series of systematic ablation experiments were conducted. The model comprises three core components: the hybrid Mamba backbone, the BIE module, and the GFU module. Their respective contributions were evaluated through a series of comparative experiments. Specifically, the hybrid Mamba backbone was replaced with LSTM and GRU structures, respectively, and the BIE module or GFU module was removed for independent comparative testing. The relevant experimental results are summarized in Table 2. The results show that, compared with all baseline configurations, HMamba-BCAF exhibits a significant improvement in fault estimation accuracy.
[0132] Table 2 Performance comparison of different models on fault diagnosis tasks
[0133] Replacing the hybrid Mamba backbone with LSTM or GRU results in a significant performance degradation. This degradation manifests as increased mean squared error and mean absolute error. The decrease in inference time and the increase in model size indicate that the hybrid Mamba architecture can more effectively capture long-range temporal dependencies while maintaining a compact representation of fault-related features. Removing the BIE module causes the largest increase in MSE and leads to... The significant decrease indicates that bidirectional temporal interactions play a crucial role in characterizing the asymmetric dynamic behavior during fault activation and recovery. While removing the BIE provides a slight improvement in inference time and model size, only the full HMamba–BCAF configuration achieves a good balance between accuracy and computational efficiency. Removing the GFU module results in the highest MAE and the lowest... Furthermore, it has minimal impact on inference speed and model size, further validating the importance of adaptive weighting and multi-scale feature selection mechanisms in enhancing diagnostic capabilities when dealing with complex and time-varying fault conditions. The HMamba-BCAF architecture achieves accurate and stable health state estimation, ensuring that the predicted results remain highly consistent with the actual fault values, and continues to play a core sensing module role in fault-tolerant control systems. Related results are as follows... Figure 5As shown in Table 3. 2. Comparative Experiments: To comprehensively evaluate the performance of the proposed HMamba-BCAF model in health status estimation, this invention conducted a series of comparative experiments with several representative baseline methods, including Crossformer, Inception, LSTM, TCN, Transformer, PatchTST, iTransformer, and LightTS. All models were trained and tested under the same data partitioning conditions and with consistent parameter settings to ensure the fairness and reliability of the comparison results. The experimental results are summarized in Table 3.
[0134] Table 3. Performance comparison of different time series models in spacecraft fault diagnosis
[0135]
[0136] Among all the comparative models, the proposed HMamba-BCAF achieved the best performance on the health status estimation task, with the lowest mean squared error (MSE) and mean absolute error (MAE), and the lowest coefficient of determination. The results demonstrate the model's superior ability to capture long-term dependencies and short-term dynamic changes. Thanks to the Mamba hierarchical structure and bidirectional cross-scale attention fusion mechanism, HMamba-BCAF can more accurately characterize the temporal features of actuator faults and external disturbances, achieving high-precision health status identification. In contrast, LSTM exhibits higher errors due to its difficulty in modeling complex correlations in long sequences; while Transformer series models (including Crossformer, PatchTST, and iTransformer) outperform recurrent networks, they still lag behind this model in handling transient mutations and time-varying faults. Overall, HMamba-BCAF demonstrates stable convergence and extremely high estimation accuracy, laying a reliable foundation for subsequent fault-tolerant control in health perception. The comparison results show that the LSTM model has a significantly higher error level, due to its structural limitations that make it difficult to effectively capture complex cross-scale temporal dependencies, especially exhibiting significant lag under non-stationary actuator faults and disturbances. Transformer-type models (Crossformer, PatchTST, iTransformer) have stronger global dependency modeling capabilities compared to recurrent structures, thus outperforming traditional RNNs overall. However, their attention mechanisms are not sensitive enough to sudden fault signals, making it difficult to maintain stable feature representations in multi-scale dynamic scenarios, thus their performance still lags behind HMamba-BCAF. Comprehensive experimental results show that HMamba-BCAF can achieve robust, fast, and high-accuracy health state estimation, achieving the best performance among all compared methods. This structure achieves accurate characterization of complex fault modes through Mamba's global modeling capabilities, GFU's multi-scale adaptive weight selection, and BIE's bidirectional temporal consistency enhancement. These reliable estimation outputs not only improve the accuracy of health monitoring but also provide continuous and reliable state inputs for subsequent active fault-tolerant control, significantly enhancing the stability and practicality of the overall fault-tolerant framework.
[0137] B. Active Fault-Tolerant Control: Based on the health state estimation results, this invention proposes an active fault-tolerant control method based on meta-Q learning. 1. Meta-Q Learning: To evaluate the stability optimization capability of the proposed method, fault impact disturbances are introduced at the 6000th, 9000th, and 11000th training rounds under actuator fault and external disturbance conditions. By applying disturbances at different training stages, the convergence and maintenance capability and recovery speed of the control strategy in dynamic environments can be verified. Figures 6-13This paper presents a comparison between the meta-Q-learning-based fault-tolerant control method and the traditional Q-learning method under the aforementioned disturbance conditions, verifying the superior performance of the control strategy of this invention in terms of stability maintenance and disturbance suppression. Experimental results show that the meta-learning mechanism can significantly improve the convergence efficiency and stability of fault-tolerant control. Compared with the baseline controller, the meta-Q-learning method exhibits higher actuator stability and smoother system recovery capability under the same disturbance conditions. When a fault impact disturbance is applied, the actuator stability index only shows a slight decrease, while the meta-Q-learning controller can recover to a stable state more quickly. These results clearly demonstrate that in the early stages of a disturbance, the meta-learning mechanism can adaptively adjust the learning rate, rapidly correct state distribution shifts, and continuously maintain dynamic consistency in the closed-loop system, thereby ensuring the stable operation of the system under fault and external disturbance conditions. 2. Reward Mechanism Ablation: To verify the effectiveness of the proposed reward function in active fault-tolerant control, this invention conducted a systematic ablation experiment, setting four reward configurations for comparative analysis: (1) without stability indicators; (2) only including actuator stability indicators; (3) only including system stability indicators; (4) a complete version including both types of stability indicators. All models were retrained under the same parameter conditions, and their performance was evaluated under four representative actuator failure scenarios. The experimental results are summarized in Table 4, which is used to compare the impact of different reward designs on the performance of the control strategy.
[0138] Table 4 Comparison of convergence performance under different stability configurations
[0139] This invention employs two metrics to evaluate the convergence performance of the algorithm under various reward configurations: the number of convergence iterations and the convergence duration. The number of convergence iterations is defined as the training epoch in which system stability first exceeds 90% and actuator stability first exceeds 50%, after which both metrics must remain above these thresholds. The convergence duration represents the total number of training epochs in which both stability metrics continuously satisfy the aforementioned convergence conditions. In a horizontal comparison of the four reward configurations, the joint reward design, which includes both system stability and actuator stability metrics, consistently exhibits the fastest convergence speed and the longest stability maintenance time, fully demonstrating the crucial role of dual stability terms in improving the performance of active fault-tolerant control.
[0140] Overall, integrating stability metrics into the composite reward function significantly improves the controller's convergence efficiency, enhances closed-loop stability, and improves its generalization ability in diverse and challenging fault scenarios. These results clearly demonstrate that in practical and safety-critical spacecraft applications, a well-designed reward mechanism that balances actuator smoothness with system dynamic stability is crucial for achieving reliable, robust, and fault-tolerant control performance.
[0141] 3. Noise Experiment: To verify the robustness of the proposed active fault-tolerant control framework under noise interference conditions, this study injected zero-mean Gaussian noise with a standard deviation of 0.02 into the system observations and actuator commands to simulate common factors in actual spacecraft, such as gyroscope drift and actuator interference. Under four noise scenarios (e.g., ... Figures 14-21 As shown in the figure, the closed-loop system consisting of the HMamba-BCAF health estimator and the meta-Q-learning controller remained stable, exhibiting only a slight performance degradation compared to the noise-free baseline. The system stability curve was almost identical to that under ideal sensing conditions, indicating that attitude dynamics itself is robust to measurement disturbances, and that HMamba-BCAF can effectively suppress health state reconstruction biases caused by noise.
[0142] Under noise, actuator stability slightly decreases due to noise-driven fluctuations in the compensation torque. However, the meta-Q-learning controller still exhibits significant advantages, recovering stability faster after a fault, reducing transient oscillations, and achieving a smoother actuator output than standard Q-learning during the steady-state phase. Overall, the meta-Q-learning strategy maintains accurate fault compensation capabilities under conditions of strong perceived noise and actuator noise, effectively avoiding the amplification effect of disturbances in the control closed loop, and significantly improving the consistency of the system during dynamic recovery.
[0143] Experimental results show that the proposed HMamba-BCAF + meta-Q-learning closed-loop fusion framework exhibits significant noise immunity under various complex noise conditions, and can stably maintain dual stability at both the actuator and system levels. Regardless of sensor disturbances, actuator uncertainties, or complex noise interference environments, this closed-loop framework demonstrates robust and reliable fault-tolerant behavior, proving its strong engineering applicability and safety assurance capabilities in practical spacecraft missions.
[0144] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning, characterized by: The specific process of the method is as follows: Step 1: Obtain the rigid body spacecraft operation dataset; The rigid body spacecraft operational dataset consists of a time series window. ; Time step samples The corresponding actual fault value is ; in, This represents a sample of rigid body spacecraft operational data at time step 1. This represents a sample of rigid body spacecraft operational data at time step 2; Indicates time step Samples of operational data from rigid-body spacecraft; Indicates time step Samples of operational data from rigid-body spacecraft; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The true value of the axial actuator performance loss fault; Indicates time step of The actual value of the axial direction actuator offset fault; Indicates time step of The actual value of the axial direction actuator offset fault; Indicates time step of The actual value of the axial direction actuator offset fault; axis, axis, The axis is in the spacecraft's body coordinate system , , The axis, the direction of advancement of the principal axis of symmetry is axis, Axis perpendicular to axis, Axis perpendicular to flat; Step 2: Construct a hybrid HMamba-BCAF health estimation model; Obtain a well-trained hybrid HMamba-BCAF health estimation model; The hybrid HMamba-BCAF health estimation model includes: a Hybrid Mamba module and a BCAF module consisting of parallel branches of BIE and GFU; Step 3: Obtain the operational data sample of the rigid body spacecraft under test, input the operational data sample of the rigid body spacecraft under test into the trained hybrid HMamba-BCAF health estimation model, the trained hybrid HMamba-BCAF health estimation model outputs the health estimation result, and control compensation is performed on the health estimation result based on the fault-tolerant control method.
2. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 1, characterized in that: In step one, at the time step samples Includes six-dimensional input ; in, This represents the dynamic residual vector calculated from the angular velocity measurement and the inertia matrix. Components along the axial direction; This represents the dynamic residual vector calculated from the angular velocity measurement and the inertia matrix. Components along the axial direction; This represents the dynamic residual vector calculated from the angular velocity measurement and the inertia matrix. Components along the axial direction; This indicates the desired control torque output by the controller. Components along the axial direction; This indicates the desired control torque output by the controller. Components along the axial direction; This indicates the desired control torque output by the controller. Components along the axial direction; axis, axis, The axis is in the spacecraft's body coordinate system , , The axis, the direction of advancement of the principal axis of symmetry is axis, Axis perpendicular to axis, Axis perpendicular to flat.
3. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 2, characterized in that: The working process of the hybrid HMamba-BCAF health estimation model in step two is as follows: Rigid body spacecraft operation dataset Each sample Input a Hybrid Mamba module, output characteristics of the Hybrid Mamba module ; feature Input to the BCAF module, the BCAF module outputs samples. Corresponding predicted fault value ; Until the rigid body spacecraft operational dataset is completed The health estimates of all samples are used to obtain a trained hybrid HMamba-BCAF health estimation model.
4. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 3, characterized in that: The Hybrid Mamba module includes: First normalized Norm layer, first convolutional layer, first SiLU activation function layer, first fully connected Linear layer, state space model SSM, learnable gated convolution operator, second fully connected Linear layer, second SiLU activation function layer; Learnable gated convolution operators include: Second convolutional layer, third SiLU activation function layer, third convolutional layer, sigmoid activation function layer.
5. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 4, characterized in that: The rigid body spacecraft operation dataset in step two Each sample Input a Hybrid Mamba module, output characteristics of the Hybrid Mamba module ; The specific process is as follows: sample Input to the first normalized Norm layer, output features from the first normalized Norm layer. ; feature Input to the first convolutional layer, output features from the first convolutional layer ; feature The first SiLU activation function layer is input, and the first SiLU activation function layer outputs features. ; feature Input to the first fully connected Linear layer, output features from the first fully connected Linear layer. ; feature Input State-Space Model (SSM), Output Features of State-Space Model (SSM) ; feature Extracting local mutation features from learnable gated convolution operators ; feature and characteristics Weighted fusion is performed to obtain features ; indicates as: feature Input to a second fully connected Linear layer; output features from the second fully connected Linear layer. ; feature The input is the second SiLU activation function layer, and the output of the second SiLU activation function layer is the feature. ; Features and samples Weighted fusion is performed to obtain features ; indicates as: feature As an output feature of the Hybrid Mamba module.
6. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 5, characterized in that: The features Extracting local mutation features from learnable gated convolution operators Specifically: feature The input is to the second convolutional layer, and the output of the second convolutional layer is to the third SiLU activation function layer, and the output of the third SiLU activation function layer is to the third SiLU activation function layer. ; feature Input to the third convolutional layer, the third convolutional layer outputs features. Input to the sigmoid activation function layer, the sigmoid activation function layer outputs features. ; Features and characteristics Perform element-wise dot product to obtain the features. .
7. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 6, characterized in that: The BCAF module includes: a second normalized Norm layer, a third fully connected Linear layer, a first ReLU activation function layer, a third normalized Norm layer, a fourth fully connected Linear layer, a second ReLU activation function layer, a cascaded Concat layer, a fifth fully connected Linear layer, a sixth fully connected Linear layer, a first feedforward neural network (FFN), a second feedforward neural network (FFN), a first multi-head attention mechanism model, a seventh fully connected Linear layer, a first pooling layer, a second pooling layer, a third pooling layer, a fourth pooling layer, an eighth fully connected layer, a first softmax activation function layer, a ninth fully connected layer, a second softmax activation function layer, and a fusion module.
8. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 7, characterized in that: The features Input to the BCAF module, the BCAF module outputs samples. Corresponding predicted fault value ; The specific process is as follows: 1) Output characteristics of the Hybrid Mamba module The input is the second normalized Norm layer, and the output of the second normalized Norm layer is the feature. ; feature Input to a third fully connected Linear layer; output features from the third fully connected Linear layer. ; feature Input to the first ReLU activation function layer, output features from the first ReLU activation function layer ; Hybrid Mamba module output characteristics Input to the third normalized Norm layer, output features from the third normalized Norm layer. ; feature Input to the fourth fully connected Linear layer, output features from the fourth fully connected Linear layer. ; feature The input is a second ReLU activation function layer, and the output of the second ReLU activation function layer is a feature. ; feature and characteristics Perform concatenation to obtain features. ; feature Input to the fifth fully connected Linear layer, output features from the fifth fully connected Linear layer. ; feature Input to the sixth fully connected Linear layer, output features from the sixth fully connected Linear layer. ; Features and characteristics By performing element-wise summation, we obtain the characteristics. ; feature Input to the first feedforward neural network (FFN), the first feedforward neural network (FFN) outputs features ; feature The second feedforward neural network (FFN) is input, and the second feedforward neural network (FFN) outputs features. ; feature Input to the first multi-head attention mechanism model, the first multi-head attention mechanism model outputs features ; feature Input to the first multi-head attention mechanism model, the first multi-head attention mechanism model outputs features ; Features and characteristics Perform concatenation to obtain features. ; feature Input to the seventh fully connected Linear layer, output features from the seventh fully connected Linear layer. ; 2) Output characteristics of the Hybrid Mamba module Input to the first pooling layer (Pooling), output features from the first pooling layer (Pooling). ; Features Perform a 1x downsampling to obtain features ;feature Input to the second pooling layer (Pooling), output features from the second pooling layer (Pooling). ; Features Perform 0.5x downsampling to obtain features. ; Hybrid Mamba module output characteristics Input to the third pooling layer (Pooling), output features from the third pooling layer (Pooling). ; Features Perform a 1x downsampling to obtain features ;feature Input to the fourth pooling layer (Pooling), output features from the fourth pooling layer (Pooling). , for features Perform 0.5x downsampling to obtain features. ; feature The layers sequentially pass through the eighth fully connected layer and the first softmax activation function layer. The first softmax activation function layer outputs the features. ; feature The layers sequentially pass through the ninth fully connected layer and the second softmax activation function layer. The second softmax activation function layer outputs the features. ; Features and characteristics Input to the fusion module, output features from the fusion module ; 3) Features With features Weighted fusion is performed to obtain the output samples of the BCAF module. Corresponding predicted fault value ; indicates as: 。 9. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 8, characterized in that: The features and characteristics Input to the fusion module, output features from the fusion module ; The specific process is as follows: Features ,feature ,feature and characteristics The difference and the characteristics of fusion Composition sequence ; will sequence The input layers are sequentially a linear layer, a ReLU activation function layer, another linear layer, and a Sigmoid activation function layer. The Sigmoid activation function layer outputs features. ; Based on features ,feature and characteristics , to obtain features ; indicates as: in, and Represents forward and backward features; The difference represents the feature; Linear layer; It is a linear layer.
10. The spacecraft health estimation and fault-tolerant control method based on hybrid Mamba and meta-Q learning according to claim 9, characterized in that: In step three, the operational data sample of the rigid body spacecraft under test is obtained. The operational data sample of the rigid body spacecraft under test is input into the trained hybrid HMamba-BCAF health estimation model. The trained hybrid HMamba-BCAF health estimation model outputs the health estimation result. The health estimation result is controlled and compensated based on the fault-tolerant control method. The specific process is as follows: 1) Output health estimation results for the trained hybrid HMamba-BCAF health estimation model. Perform meta-learning to obtain Meta-learning initial parameters at time step ; 2) Order ; 3) Obtain state of time ; state of time It is expressed as follows: in, express angular velocity at time t, express Scalar components of the quaternion error at time step; Represents the L2 norm, Represents absolute value; 4) Execution state of time Get Moment of action ; 5) Based on state of time and Moment of action conduct Value evaluation, obtained The beginning of time function , and Momentary Rewards The specific process is as follows: in, express The state at any given moment, express Momentary actions express Initial parameters at time; express Initial parameters at time; , , , , , express function; express Momentary rewards; Indicates basic environmental rewards; express System stability indicators at any given time; express Actuator stability metrics at any given time; This indicates the reward items for successful task completion; The weighting coefficients represent the system stability index; The weighting coefficients represent the stability index of the actuator; in in, The variance representing the rate of change of angular velocity; The modulus representing the average angular velocity; The weighting coefficients represent the system stability index; The weighting coefficients represent the stability index of the actuator; This represents the stability component extracted from the scalar term in the attitude error quaternion; The mean square term representing the actuator output torque; coefficient and These correspond to the attitude stability weighting factor and the energy smoothness weighting factor, respectively. This indicates a successful terminal reward. and These represent the threshold values for attitude quaternions and angular velocity, respectively. 6) Based on Optimal action at any given moment Get state of time , ; in, Optimal action at any given moment ; The superscript T indicates transpose; 7) Execution state of time Get Moment of action ; 8) Based on state of time and Moment of action conduct Value evaluation, obtained Moment function , and Momentary Rewards The specific process is as follows: in, express Parameters at time; express Parameters at time; , express Moment Value function value; , express Moment Value function value; express Momentary rewards; Indicates basic environmental rewards; express System stability indicators at any given time; express Actuator stability metrics at any given time; This indicates the reward items for successful task completion; The weighting coefficients represent the system stability index; The weighting coefficients represent the stability index of the actuator; 9) Based on Moment Value function value and Correction Moment Value function value and The specific process is as follows: in, Indicates the learning rate; Indicates the discount factor; , express Moment Value function value; express The optimal action at any given moment; 10) Order ; Repeat steps 6) through 9), using the Reptile optimization strategy to adjust the parameters. and Update until the rigid body spacecraft operational dataset is complete. All samples The control compensation is used to obtain the trained hybrid HMamba-BCAF health estimation model and the parameters of the corresponding Q-learning module. and .