A space operation and maintenance prediction method and system based on multi-modal data
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-11
AI Technical Summary
当前运维多依赖单一模态数据(遥测、图像等),存在信息维度不足、跨部件耦合关系捕捉缺失、易受噪声干扰、健康评估不准、故障误判率高等问题;同时仅能碎片化预测故障、寿命或位置偏差,缺乏多指标联合预测能力,干预策略被动,难以适配空间设备高可靠、高自主的运维需求
基于结构因果模型识别跨模态因果关系图,并根据跨模态因果关系图对多模态数据进行时间同步、空间对齐以及因果关联映射,构建因果语义数字孪生体,使得数字孪生体不仅包含设备的几何与物理状态,还内嵌了跨模态的因果驱动关系,提升了状态感知全面性与准确性;通过跨模态因果注意力机制,生成多部件耦合的健康状态隐空间向量,能准确表征多部件之间的耦合退化关系,同时通过TOPSIS法计算健康度指数,能实时、灵敏地反映设备当前的健康水平及其在级联退化路径上的动态变化,从而避免突发故障导致的安全事故,显著提升设备在轨运行的可靠性与稳定性;通过包含特征耦合编码器和交互预测解码器的混合预测引擎,输出故障概率、剩余使用寿命与位置漂移的联合概率分布,能实现故障、寿命与位置偏差的耦合联动预测,突破单指标碎片化预测局限,提升预测完整性与精度,为全局运维决策提供可靠依据;基于联合概率分布及当前资源状态,由预先训练的风险感知元强化学习智能体生成干预策略并执行,能实现风险预判、资源优化与自主干预,降低在轨故障发生率与运行风险,最大程度保障设备在轨安全稳定运行。
Smart Images

Figure CN122548645A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of equipment operation and maintenance prediction technology, and in particular to a spatial operation and maintenance prediction method and system based on multimodal data. Background Technology
[0002] Space equipment operates in harsh environments with strong radiation and extreme temperature differences, making its components prone to aging and degradation. Furthermore, it cannot be maintained in orbit and carries high operational risks, necessitating precise and forward-looking maintenance prediction technologies. Current maintenance methods largely rely on single-modal data (telemetry, images, etc.), which suffer from insufficient information dimensions, lack of capture of cross-component coupling relationships, susceptibility to noise interference, inaccurate health assessments, and high rates of fault misdiagnosis. Moreover, they can only predict faults, lifespan, or location deviations in a fragmented manner, lacking the ability to jointly predict multiple indicators, resulting in passive intervention strategies that are ill-suited to the high reliability and autonomy requirements of space equipment maintenance.
[0003] With the development of multimodal sensing technology, data such as 3D point clouds, high-definition images, and multi-source telemetry can comprehensively characterize equipment status. However, existing technologies lack deep multimodal data fusion mechanisms and causal modeling capabilities, failing to form an integrated operation and maintenance system and fully leveraging the complementary advantages of multimodal data. Therefore, there is an urgent need for a multimodal spatial operation and maintenance prediction scheme that integrates causal reasoning and intelligent algorithms to ensure the safe and stable operation of equipment in orbit. Summary of the Invention
[0004] Therefore, it is necessary to provide a space operation and maintenance prediction method and system based on multimodal data that can ensure the safe and stable operation of equipment in orbit, addressing the aforementioned technical problems.
[0005] Firstly, this application provides a spatial operation and maintenance prediction method based on multimodal data, the method comprising: Multimodal data from space devices are collected, cross-modal causal relationship graphs are identified based on structural causal models, and time synchronization, spatial alignment, and causal association mapping are performed on the multimodal data according to the cross-modal causal relationship graphs to construct a causal semantic digital twin. In the causal semantic digital twin, features extracted from the multimodal data are fused through a cross-modal causal attention mechanism to generate a multi-component coupled health status latent space vector, and a health index is calculated based on the cross-modal causal relationship graph using the TOPSIS method. A hybrid prediction engine comprising a feature-coupled encoder and an interactive prediction decoder is constructed, which combines the health state latent space vector and the health index to output the joint probability distribution of failure probability, remaining lifetime and position drift. Based on the joint probability distribution and the current resource status extracted from the causal semantic digital twin, an intervention strategy is generated and executed by a pre-trained risk-aware meta-reinforcement learning agent, wherein the current resource status includes remaining fuel, power, available redundant components, and task priority.
[0006] In one embodiment, the construction of a hybrid prediction engine comprising a feature-coupled encoder and an interactive predictive decoder, combining the health state latent space vector and the health index, and outputting a joint probability distribution of failure probability, remaining lifetime, and location drift includes: Construct a hybrid prediction engine that includes a feature-coupled encoder and an interactive prediction decoder; The historical trajectory dataset and spatial environment parameters are obtained, and the health status latent space vector, the health index, the historical trajectory dataset and the spatial environment parameters are time-aligned and then input into the feature coupling encoder to generate coupled features. The coupling features are input into the interactive prediction decoder, which includes a fault prediction branch and a location change prediction branch. Through iterative interaction between the fault prediction branch and the location change prediction branch, the joint probability distribution of fault probability, remaining service life, and location drift is finally output.
[0007] In one embodiment, the step of iteratively interacting the fault prediction branch and the location change prediction branch to finally output the joint probability distribution of fault probability, remaining useful life, and location drift includes: Based on the coupling characteristics, fault evolution inference is performed through the fault prediction branch, the change in physical parameters caused by component degradation is output, and the change in physical parameters is transmitted to the position change prediction branch. The position change prediction branch corrects the internal dynamic model parameters based on the received changes in physical parameters, and performs trajectory and attitude inference based on the corrected dynamic model parameters, outputting vibration parameters and attitude deviation parameters. The vibration parameters and the attitude deviation parameters are used as acceleration factors and fed back from the position change prediction branch to the fault prediction branch to adjust the degradation accumulation rate. The fault prediction branch and the location change prediction branch interact in multiple rounds of iteration until the preset convergence condition is met, and finally output the joint probability distribution of fault probability, remaining service life and location drift.
[0008] In one embodiment, the position change prediction branch corrects the internal dynamic model parameters based on the received changes in physical parameters, and performs trajectory and attitude inference based on the corrected dynamic model parameters, outputting vibration parameters and attitude deviation parameters, including: The position change prediction branch receives the physical parameter changes transmitted by the fault prediction branch, the physical parameter changes including the centroid offset and inertia tensor changes caused by component degradation; Based on the centroid offset and the inertia tensor change, correct the centroid position parameters and inertia matrix parameters in the internal dynamic model of the position change prediction branch; Based on the corrected centroid position parameters and inertia matrix parameters, combined with the space environment parameters, orbit extrapolation and attitude evolution calculations are performed on the space equipment to obtain the predicted position and attitude values in the future time domain. The predicted position and the predicted attitude are compared with the nominal values, and the vibration amplitude and vibration frequency components are extracted as vibration parameters, and the attitude angle deviation and attitude drift rate are extracted as attitude deviation parameters. Output the vibration parameters and the attitude deviation parameters.
[0009] In one embodiment, the causal semantic digital twin, by fusing features extracted from the multimodal data through a cross-modal causal attention mechanism to generate a multi-component coupled health state latent space vector, and calculating a health index based on the cross-modal causal relationship graph using the TOPSIS method, includes: In the causal semantic digital twin, geometric deformation features of point clouds, surface degradation features of images, and time-frequency domain statistical features of telemetry data are extracted from the multimodal data, respectively. Based on the cross-modal causal relationship diagram, determine whether there is a causal connection between different modal feature channels of the geometric deformation features of the point cloud, the surface degradation features of the image, and the time-frequency domain statistical features of the telemetry data; Through a cross-modal causal attention mechanism, attention weights are calculated only for feature channels with causal connections, and features extracted from the multimodal data are weighted and fused based on the attention weights to generate a multi-component coupled health state latent space vector. Based on the latent space vector of the health status, the health index is calculated using the TOPSIS method through the cross-modal causal relationship graph.
[0010] In one embodiment, the step of calculating attention weights only for feature channels with causal connections through a cross-modal causal attention mechanism, and then weighting and fusing the features extracted from the multimodal data based on these attention weights to generate a multi-component coupled health state latent space vector includes: By using a cross-modal causal attention mechanism, the causal connections between different modal feature channels of the point cloud's geometric deformation features, the image's surface degradation features, and the time-frequency domain statistical features of the telemetry data are obtained from the cross-modal causal relationship graph. Based on the causal connections, a query vector, key vector, and value vector are constructed for each feature channel, and attention scores are calculated only for feature channels with causal connections to generate an attention weight matrix with causal constraints. Based on the attention weight matrix, the value vectors of each feature channel extracted from the multimodal data are weighted and summed to obtain the health state latent space vector of the multi-component coupling.
[0011] In one embodiment, the step of calculating the health index using the TOPSIS method based on the health status latent space vector and the cross-modal causal relationship graph includes: An evaluation matrix is constructed using each dimension of the latent space vector of the health status as an evaluation index and the states of each component in the multimodal data at different times as evaluation objects. The causal connection strength between each modal feature channel is extracted from the cross-modal causal relationship graph, and the entropy weight of each evaluation index is dynamically determined based on the causal connection strength. The evaluation matrix is weighted and standardized based on the entropy weight, and the relative alignment of each component with the positive and negative ideal solutions is calculated using the TOPSIS method to obtain the health index.
[0012] In one embodiment, the step of generating and executing an intervention strategy by a pre-trained risk-aware meta-reinforcement learning agent based on the joint probability distribution and the current resource state extracted from the causal semantic digital twin includes: The predicted values of failure probability, remaining useful life, and location drift in the joint probability distribution are concatenated with the current resource state extracted from the causal semantic digital twin to form a decision state vector. The decision state vector is input into a pre-trained risk-aware meta-reinforcement learning agent. The risk-aware meta-reinforcement learning agent aims to maximize the task success rate and minimize the intervention cost and tail risk. It performs multi-step deduction in the virtual simulation environment composed of the causal semantic digital twin to generate an intervention strategy. After performing a security verification on the intervention strategy, the intervention strategy is converted into an executable instruction sequence and sent to the space device for execution.
[0013] Secondly, this application also provides a space operation and maintenance prediction system based on multimodal data, the system comprising: Multimodal data acquisition module, used to acquire multimodal data from space equipment; A digital twin construction module is used to identify cross-modal causal relationship graphs based on structural causal models, and to perform time synchronization, spatial alignment, and causal association mapping on the multimodal data according to the cross-modal causal relationship graphs to construct a causal semantic digital twin. The feature fusion and health assessment module is used to fuse features extracted from the multimodal data in the causal semantic digital twin through a cross-modal causal attention mechanism, generate a multi-component coupled health status latent space vector, and calculate the health index based on the cross-modal causal relationship graph using the TOPSIS method. The hybrid prediction engine module includes a feature-coupled encoder and an interactive prediction decoder, which combine the health state latent space vector and the health index to output a joint probability distribution of failure probability, remaining lifetime and location drift. The decision-making and intervention module is used to generate and execute intervention strategies by a pre-trained risk-aware reinforcement learning agent based on the joint probability distribution and the current resource status extracted from the causal semantic digital twin. The current resource status includes remaining fuel, power, available redundant components, and task priority.
[0014] In summary, this application includes the following beneficial technical effects: Based on a structural causal model, cross-modal causal relationship graphs are identified. These graphs are then used to perform time synchronization, spatial alignment, and causal association mapping on multimodal data, constructing a causal semantic digital twin. This digital twin not only includes the device's geometric and physical state but also embeds cross-modal causal driving relationships, improving the comprehensiveness and accuracy of state perception. A cross-modal causal attention mechanism generates a latent health state vector that accurately represents the coupling degradation relationships between multiple components. Furthermore, the TOPSIS method is used to calculate a health index, which can reflect the device's current health level and its dynamic changes along cascading degradation paths in real time and with high sensitivity, thereby preventing sudden failures from leading to… This significantly improves the reliability and stability of on-orbit operation of equipment, preventing safety accidents. Through a hybrid prediction engine comprising a feature-coupled encoder and an interactive predictive decoder, it outputs a joint probability distribution of fault probability, remaining service life, and position drift. This enables coupled and linked prediction of faults, service life, and position deviation, overcoming the limitations of fragmented single-indicator prediction, improving prediction completeness and accuracy, and providing a reliable basis for global operation and maintenance decisions. Based on the joint probability distribution and current resource status, a pre-trained risk-aware reinforcement learning agent generates and executes intervention strategies, enabling risk prediction, resource optimization, and autonomous intervention. This reduces the on-orbit fault rate and operational risks, maximizing the safe and stable operation of equipment in orbit. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating a spatial operation and maintenance prediction method based on multimodal data in one embodiment; Figure 2 This is a flowchart illustrating a spatial operation and maintenance prediction method based on multimodal data in another embodiment; Figure 3 This is a block diagram of a space operation and maintenance prediction system based on multimodal data in one embodiment. Detailed Implementation
[0016] This invention provides a spatial operation and maintenance prediction method and system based on multimodal data.
[0017] The embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0018] In the description of the embodiments disclosed in this invention, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0019] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the spatial operation and maintenance prediction method and system based on multimodal data in this invention includes: S100 collects multimodal data from space equipment, identifies cross-modal causal relationship graphs based on structural causal models, and performs time synchronization, spatial alignment, and causal association mapping on multimodal data according to the cross-modal causal relationship graphs to construct a causal semantic digital twin.
[0020] Specifically, multimodal data is first collected using various sensors deployed on space equipment. This multimodal data includes, but is not limited to, 3D point cloud data obtained from lidar scanning, image data captured by visible light or infrared cameras, and various physical parameter data recorded by the internal telemetry system. The raw data collected cannot be directly used for subsequent causal analysis due to differences in sampling frequency, coordinate systems, and noise interference. Therefore, causal structure learning is first performed on historical multimodal data based on structural causal models (such as PC algorithms, fast causal inference algorithms, or scoring-based search algorithms) to identify directed causal dependencies between different modal characteristic variables, forming a cross-modal causal relationship graph. This graph uses nodes to represent each modal characteristic (such as temperature, deformation, vibration amplitude, etc.) and directed edges to represent the causal driving direction (such as "temperature increase → thermal expansion → geometric deformation"). Subsequently, based on the causal relationship diagram, the multimodal data undergoes time synchronization processing, aligning data with different sampling frequencies to a unified time axis through interpolation or resampling; spatial alignment processing is performed, unifying point cloud data, image data, etc., to the ontological coordinate system of the space device through coordinate transformation; causal association mapping is then performed, that is, labeling each data channel with its causal role (dependent variable or effect variable) and the dependencies between them according to the paths in the causal relationship diagram. Finally, the processed multimodal data and its causal structure are injected into a pre-constructed digital twin framework to form a causal semantic digital twin. This digital twin not only includes traditional digital twin elements such as device geometry, physics, and behavior, but also embeds cross-modal causal driving relationships, thereby more realistically reflecting the collaborative degradation and fault propagation mechanisms between various components of the space device, providing a high-fidelity virtual mirror environment for subsequent health status assessment and prediction.
[0021] S200, in the causal semantic digital twin, fuses features extracted from multimodal data through a cross-modal causal attention mechanism to generate a multi-component coupled health status latent space vector, and calculates the health index based on the cross-modal causal relationship graph using the TOPSIS method.
[0022] Specifically, in the constructed causal semantic digital twin, features such as geometric deformation, surface degradation, and telemetry time-frequency domain corresponding to multimodal data are extracted. Using a cross-modal causal attention mechanism, attention weights are calculated and weighted fusion is performed only on feature channels with causal relationships to generate a latent space vector of health status that can represent the coupling relationship of multiple components. At the same time, based on the cross-modal causal relationship graph, the TOPSIS (Approximation Ideal Solution Ranking Method) method is used to quantitatively evaluate the latent space vector of health status and calculate a health index that can intuitively reflect the overall health level of the equipment, so as to achieve accurate and reliable quantitative representation of the coupling health status of multiple components of space equipment.
[0023] S300 constructs a hybrid prediction engine that includes a feature-coupled encoder and an interactive predictive decoder. It combines the latent space vector of health status and the health index to output the joint probability distribution of failure probability, remaining lifetime and location drift.
[0024] Specifically, a hybrid prediction engine integrating a feature-coupled encoder and an interactive predictive decoder is constructed. The generated health status latent space vector and health index are combined with historical trajectory data and spatial environment parameters and then time-aligned before being input into the encoder. After feature coupling encoding, coupled features that integrate multi-dimensional information are obtained. The coupled features are then input into the interactive predictive decoder. Through the linkage deduction of fault prediction and position change prediction, the joint probability distribution of fault probability, remaining service life and position drift is finally output, realizing the integrated coupled prediction of equipment degradation, lifespan decay and position deviation.
[0025] S400, based on the joint probability distribution and the current resource state extracted from the causal semantic digital twin, generates and executes intervention strategies by a pre-trained risk-aware meta-reinforcement learning agent.
[0026] Specifically, after obtaining the joint probability distribution, it needs to be transformed into executable operational intervention decisions. Considering the highly limited resources of space equipment in orbit (such as limited fuel, power, and spare parts) and the dynamic changes in mission priorities, a pre-trained risk-aware meta-reinforcement learning agent is introduced, which operates in a sequential decision-making manner. First, the predicted information in the joint probability distribution is concatenated with the current resource status (remaining fuel, current power level, list and quantity of available redundant parts, current mission priority level, etc.) extracted in real time from the causal semantic digital twin to form a comprehensive decision state vector. This vector includes both the future risk situation and the current capability constraints. Then, this decision state vector is input into the agent, which is pre-trained using a model-independent meta-learning algorithm (MAML) or a similar meta-reinforcement learning framework. Simultaneously, a conditional risk value term is introduced into the agent's reward function, making it more sensitive to tail-end extreme risks (i.e., low-probability but severe events) in the joint probability distribution. The intelligent agent performs multi-step deduction in a virtual simulation environment composed of a causal semantic digital twin, evaluates the expected cumulative reward of different candidate intervention strategies under risk constraints, and selects the optimal strategy.
[0027] In one embodiment, such as Figure 2 As shown, S300 includes: S310, constructing a hybrid prediction engine that includes a feature-coupled encoder and an interactive prediction decoder; S320: Obtain historical trajectory dataset and spatial environment parameters, and after time alignment of health status latent space vector, health index, historical trajectory dataset and spatial environment parameters, input them together into feature coupling encoder to generate coupled features; S330, inputs the coupled features to the interactive predictive decoder; S340, through the iterative interaction of the fault prediction branch and the location change prediction branch, finally outputs the joint probability distribution of fault probability, remaining useful life and location drift.
[0028] Specifically, firstly, a hybrid prediction engine is constructed, consisting of two core modules: a feature-coupled encoder and an interactive predictive decoder. The feature-coupled encoder deeply fuses multi-source input information into a unified coupled feature representation; the interactive predictive decoder includes a fault prediction branch and a position change prediction branch, which can interact to collaboratively output multi-dimensional prediction results. After the engine is built, historical trajectory datasets and current spatial environment parameters are obtained from the causal semantic digital twin. The historical trajectory dataset records the pose information of the space device over a period of time, such as time-series data of position coordinates, velocity, acceleration, attitude angle, and angular velocity. The spatial environment parameters reflect the spatial environmental conditions in which the device is located, including but not limited to atmospheric density, solar radiation flux, geomagnetic index, and gravity gradient. At the same time, the health state latent space vector (a high-dimensional feature vector representing the coupled degradation state of multiple components) and the health index (a scalar value used to quantify the overall health level of the device) obtained in the previous steps are extracted. Since the four types of data mentioned above may have different sampling frequencies, time bases, or data formats, time alignment processing is required first. Specifically, methods such as linear interpolation, spline interpolation, or nearest neighbor interpolation can be used to unify all data sequences to the same timestamp, ensuring that the health status, historical trajectory, and environmental parameters at the same moment can be accurately matched and fused. The time-aligned health status latent space vector, health index, historical trajectory dataset, and spatial environmental parameters are then input into the feature-coupled encoder. The feature-coupled encoder uses a cross-attention mechanism to map these inputs to a common coupled feature space, generating a coupled feature that includes degradation information, position information, and environmental information. This coupled feature retains the key information from each input source and establishes high-order correlations between them. Subsequently, the generated coupled feature is input into the interactive predictive decoder. The interactive predictive decoder has a fault prediction branch and a position change prediction branch. The fault prediction branch simulates the performance degradation process of equipment or components over time, while the position change prediction branch simulates the evolution of the orbit and attitude of the space equipment over time under the combined effects of environmental disturbances and internal state changes. The two branches do not operate independently, but rather simulate the physical two-way coupling effect through multiple iterative interactions. Specifically, the fault prediction branch passes its output degradation-related information to the position change prediction branch to help correct the latter's pose evolution inference; the position change prediction branch feeds back its output pose deviation and related dynamic parameters to the fault prediction branch to adjust the degradation process.After iterative interaction between the fault prediction branch and the location change prediction branch, until the preset convergence condition is met or the preset number of iterations is completed, the interactive prediction decoder finally outputs a joint probability distribution. This joint probability distribution describes the dependency relationship between the three variables of fault probability, remaining service life and location drift, thereby providing complete uncertainty quantification information for subsequent operation and maintenance decisions.
[0029] In one embodiment, through iterative interaction between the fault prediction branch and the location change prediction branch, the final output joint probability distribution of fault probability, remaining useful life, and location drift includes: Based on the coupling characteristics, fault evolution inference is performed through the fault prediction branch, which outputs the changes in physical parameters caused by component degradation and transmits these changes to the position change prediction branch. The position change prediction branch then corrects the internal dynamic model parameters based on the received changes in physical parameters and performs trajectory and attitude inference based on the corrected dynamic model parameters, outputting vibration parameters and attitude deviation parameters. These vibration parameters and attitude deviation parameters are used as acceleration factors and fed back from the position change prediction branch to the fault prediction branch to adjust the degradation accumulation rate. The fault prediction branch and the position change prediction branch interact in multiple rounds until the preset convergence condition is met, and finally outputs the joint probability distribution of fault probability, remaining service life, and position drift.
[0030] Specifically, after inputting the coupling features into the interactive predictive decoder, the fault prediction branch infers fault evolution based on these coupling features. The fault prediction branch simulates the degradation process of a component over time based on information such as the health state latent space vector, health index, historical trajectory, and spatial environment parameters contained in the coupling features. It outputs the changes in physical parameters caused by component degradation. These changes characterize the changes in the mechanical or physical properties of the component due to wear, fatigue, aging, etc., such as changes in mass distribution, stiffness degradation, and friction coefficient. The fault prediction branch then transmits these changes in physical parameters to the position change prediction branch. After receiving the changes in physical parameters from the fault prediction branch, the position change prediction branch uses this information to correct its internal dynamic model parameters. Specifically, the position change prediction branch originally maintains a dynamic model describing the orbital and attitude evolution of the space equipment. Some parameters in this model (such as the center of mass position, inertia distribution, and elastic coefficients) use nominal values when the equipment is healthy. When it receives changes in physical parameters due to component degradation, the position change prediction branch corrects the corresponding model parameters so that the dynamic model can reflect the impact of degradation on the equipment's posture and motion. Subsequently, based on the corrected dynamic model parameters, combined with the input space environment parameters and historical trajectory data, the position change prediction branch performs trajectory and attitude inference calculations, predicting the position changes, attitude evolution, and resulting dynamic responses of the space equipment over a future period, and outputs vibration parameters and attitude deviation parameters. The vibration parameters reflect the mechanical vibration characteristics of the equipment caused by internal degradation or external disturbances, such as vibration amplitude and vibration frequency components; the attitude deviation parameters reflect the degree of deviation between the actual attitude and the desired attitude, such as attitude angle deviation and attitude drift rate. The position change prediction branch feeds back the output vibration and attitude deviation parameters as acceleration factors to the fault prediction branch. Upon receiving these feedback parameters, the fault prediction branch incorporates them into the calculation of the degradation accumulation rate. Specifically, vibration and attitude deviation exacerbate friction, impact, and fatigue between components, thereby accelerating further degradation. The fault prediction branch dynamically adjusts the degradation accumulation rate based on the feedback acceleration factors, ensuring that subsequent fault evolution inferences reflect the positive feedback effect of pose disturbances on the degradation process. This process is not completed in a single execution but requires multiple iterative interactions. In each iteration, the fault prediction branch updates the degradation accumulation rate based on the vibration and attitude deviation parameters from the previous iteration and outputs new physical parameter changes. The position change prediction branch then re-adjusts the dynamic model parameters based on these new physical parameter changes and outputs new vibration and attitude deviation parameters. The fault prediction branch and the location change prediction branch exchange information repeatedly in this manner until the preset convergence condition is met. The convergence condition can be that the change in the joint probability distribution of the output between two consecutive iterations is less than a set threshold, or that the preset maximum number of iterations is reached.After the iterative interaction is completed, the interactive prediction decoder finally outputs a joint probability distribution, which describes the dependency relationship between the three variables of failure probability, remaining service life and position drift. This accurately captures the bidirectional coupling effect between component degradation and pose evolution, providing high-precision risk situation prediction results for subsequent operation and maintenance decisions.
[0031] In one embodiment, the position change prediction branch corrects the internal dynamic model parameters based on the received changes in physical parameters, and performs trajectory and attitude inference based on the corrected dynamic model parameters, outputting vibration parameters and attitude deviation parameters, including: The position change prediction branch receives the physical parameter changes transmitted by the fault prediction branch. These physical parameter changes include the centroid offset and inertia tensor changes caused by component degradation. Based on the centroid offset and inertia tensor changes, the centroid position parameters and inertia matrix parameters in the internal dynamic model of the position change prediction branch are corrected. Based on the corrected centroid position parameters and inertia matrix parameters, combined with space environment parameters, orbit extrapolation and attitude evolution calculations are performed on the space equipment to obtain the predicted position and attitude values in the future time domain. The predicted position and attitude values are compared with the nominal values, and the vibration amplitude and vibration frequency components are extracted as vibration parameters, and the attitude angle deviation and attitude drift rate are extracted as attitude deviation parameters. The vibration parameters and attitude deviation parameters are then output.
[0032] Specifically, the position change prediction branch first receives the physical parameter changes transmitted by the fault prediction branch. These physical parameter changes include the centroid offset and inertia tensor changes caused by component degradation. The centroid offset reflects the shift of the mass distribution center relative to the design nominal position due to factors such as localized wear, material shedding, or component displacement. The inertia tensor change reflects the change in the overall rotational inertia of the equipment due to component degradation, such as changes in the moments of inertia and products of inertia around each coordinate axis caused by mass redistribution. Subsequently, the position change prediction branch corrects the corresponding parameters in its internally maintained dynamic model based on the received centroid offset and inertia tensor changes. Specifically, the original centroid position parameter is updated to the actual centroid position after adding the centroid offset to the nominal centroid position; the original inertia matrix parameter is updated to the actual inertia matrix after adding the inertia tensor change to the nominal inertia matrix. Through these corrections, the dynamic model can accurately reflect the impact of component degradation on the equipment's mass distribution and rotational characteristics. Building upon this foundation, the position change prediction branch performs orbit extrapolation and attitude evolution calculations for the space equipment based on the corrected center-of-mass position parameters and inertia matrix parameters, combined with space environment parameters (such as atmospheric density, solar radiation flux, geomagnetic index, and gravity gradient) obtained from the causal semantic digital twin. Orbit extrapolation refers to predicting the position and velocity time series over a future period based on the current orbital state and external forces (including Earth's gravity, atmospheric drag, and solar radiation pressure), using numerical integration methods of the dynamic equations. Attitude evolution calculation refers to predicting the changes in attitude angle and angular velocity over a future period based on the current attitude state and external torques acting on the equipment (including gravity gradient torque, aerodynamic torque, and solar radiation pressure torque). After these calculations, the predicted position and attitude values of the equipment in the future time domain are obtained. After obtaining the predicted position and attitude values, the position change prediction branch compares them with the corresponding nominal values, which are the theoretical values of position and attitude calculated based on the ideal dynamic model under the condition that the equipment is completely healthy, without degradation, and without abnormal disturbances. By calculating the difference between the predicted and nominal position values, vibration amplitude and frequency components are extracted as vibration parameters. Similarly, by calculating the difference between the predicted and nominal attitude values, attitude angle deviation and attitude drift rate are extracted as attitude deviation parameters. Vibration amplitude reflects the severity of the equipment's deviation from the nominal trajectory, while the vibration frequency component reflects the main frequency domain characteristics of the vibration. Attitude angle deviation reflects the angular error between the actual and desired orientation of the equipment, and attitude drift rate reflects how quickly this deviation changes over time. Finally, the position change prediction branch outputs the calculated vibration parameters (including vibration amplitude and frequency components) and attitude deviation parameters (including attitude angle deviation and attitude drift rate) and feeds them back to the fault prediction branch as acceleration factors for adjusting the degradation accumulation rate in subsequent iterations.
[0033] In one embodiment, within a causal semantic digital twin, features extracted from multimodal data are fused using a cross-modal causal attention mechanism to generate a multi-component coupled health state latent space vector. Based on the cross-modal causal relationship graph, a health index is calculated using the TOPSIS method, including: In the causal semantic digital twin, geometric deformation features of point clouds, surface degradation features of images, and time-frequency domain statistical features of telemetry data are extracted from multimodal data. Based on the cross-modal causal relationship graph, it is determined whether there are causal connections between different modal feature channels of the geometric deformation features of point clouds, surface degradation features of images, and time-frequency domain statistical features of telemetry data. Through a cross-modal causal attention mechanism, attention weights are calculated only for feature channels with causal connections, and the features extracted from multimodal data are weighted and fused based on the attention weights to generate a multi-component coupled health state latent space vector. Based on the health state latent space vector, the health index is calculated using the TOPSIS method with the cross-modal causal relationship graph.
[0034] Specifically, firstly, different types of high-level features are extracted from the collected multimodal data. Specifically, geometric deformation features reflecting structural deformation and surface curvature changes are extracted from 3D point cloud data acquired by LiDAR; surface degradation features reflecting wear, cracks, contamination, and other appearance degradation are extracted from image data captured by visible light or infrared cameras; and time-frequency domain statistical features, such as mean, variance, spectral peak value, and wavelet energy, are extracted from various physical parameter data recorded by the telemetry system as time-frequency domain statistical features of the telemetry data. Then, based on a pre-constructed cross-modal causal relationship graph, it is determined whether there are causal connections between the extracted different modal feature channels, i.e., which feature channels have direct causal driving relationships. On this basis, a cross-modal causal attention mechanism is introduced. The core constraint of this mechanism is that attention weights are calculated only for feature channels with causal connections in the causal relationship graph, while the attention weights for feature channels without causal connections are forcibly set to zero. In this way, attention weights based on causal priors are calculated, and these weights are used to weightedly fuse features from different modalities, generating a low-dimensional, dense, multi-component coupled health state latent space vector. This vector couples degradation information from different components and modalities together, accurately reflecting the mutual influence and cascading degradation paths between multiple components. Finally, based on this health state latent space vector, a health index is calculated using the TOPSIS method with a cross-modal causal relationship graph. Specifically, each dimension of the health state latent space vector is used as an evaluation indicator, and the states of different time points or different components are used as evaluation objects to construct an evaluation matrix. The causal connections implied in the causal relationship graph are used to weight the evaluation process, and finally, the relative closeness of each evaluation object to the positive and negative ideal solutions is calculated. This closeness is the health index, which typically ranges from 0 to 1 and is used to quantitatively reflect the current comprehensive health level of the device.
[0035] In one embodiment, an attention weight is calculated only for feature channels with causal connections using a cross-modal causal attention mechanism, and features extracted from multimodal data are weighted and fused based on these attention weights to generate a multi-component coupled health state latent space vector, including: By employing a cross-modal causal attention mechanism, causal connections between different modal feature channels—geometric deformation features of point clouds, surface degradation features of images, and time-frequency domain statistical features of telemetry data—are obtained from the cross-modal causal relationship graph. Based on these causal connections, query vectors, key vectors, and value vectors are constructed for each feature channel, and attention scores are calculated only for feature channels with causal connections, generating an attention weight matrix with causal constraints. Based on the attention weight matrix, the value vectors of each feature channel extracted from the multimodal data are weighted and summed to obtain the latent space vector of the health status of multiple components coupled together.
[0036] Specifically, firstly, a pre-constructed cross-modal causal relationship graph is read through a cross-modal causal attention mechanism to obtain causal connections between different modal feature channels, such as geometric deformation features of point clouds, surface degradation features of images, and time-frequency domain statistical features of telemetry data. Specifically, the causal relationship graph uses nodes to represent each feature channel and directed edges to represent the causal driving relationships between channels. Based on this, the cross-modal causal attention mechanism generates a causal adjacency matrix, identifying which feature channel pairs have causal connections. Then, a query vector, key vector, and value vector are constructed for each feature channel. The query vector represents the attention requirement of the current channel for other channels, the key vector represents the feature identifier of the current channel, and the value vector carries the feature value of the current channel. When calculating the attention score, causal constraints are strictly followed; attention scores are only allowed to be calculated when there is a causal connection between two feature channels. If no causal connection exists, the corresponding attention score is directly set to negative infinity or zero, thus not generating effective weights in the subsequent Softmax normalization. Based on the attention scores under the aforementioned causal constraints, a causal-constrained attention weight matrix is generated, where non-zero elements appear only between channel pairs with causal connections. Finally, based on this attention weight matrix, the value vectors of each feature channel extracted from the multimodal data are weighted and summed. Specifically, for each target channel, the value vectors of all source channels with which it has a causal connection are weighted and summed according to their corresponding attention weights to obtain the updated representation of that target channel. The updated representations of all channels together constitute the health state latent space vector of the multi-component coupling. Through this approach, the attention mechanism focuses only on cross-modal features with physically causal driving relationships, effectively suppressing spurious fusion caused by noise or random correlations, and improving the interpretability and robustness of the health state latent space vector.
[0037] In one embodiment, the health index is calculated using the TOPSIS method based on the latent space vector of health status and a cross-modal causal relationship graph, including: Using the dimensions of the latent space vector of health status as evaluation indicators and the states of each component at different times in multimodal data as evaluation objects, an evaluation matrix is constructed. The causal connection strength between each modal feature channel is extracted from the cross-modal causal relationship graph, and the entropy weight of each evaluation indicator is dynamically determined according to the causal connection strength. The evaluation matrix is weighted and standardized based on the entropy weight, and the relative alignment of each component with the positive ideal solution and the negative ideal solution is calculated using the TOPSIS method to obtain the health index.
[0038] Specifically, firstly, each dimension of the latent space vector of health status is used as an evaluation index. Simultaneously, the states of each component in the multimodal data at different times are used as evaluation objects; for example, the overall state of the i-th component at the t-th sampling time can be considered an evaluation object. Based on these evaluation indices and objects, an initial evaluation matrix is constructed, where each element represents the value of a certain evaluation object on a certain evaluation index. Then, the causal connection strength between each modal feature channel is extracted from the cross-modal causal relationship graph. This causal connection strength can be the weight value of directed edges, which can be learned from historical data through a structural causal model, reflecting the magnitude of the driving effect of causal features on outcome features. Based on these causal connection strengths, the entropy weight of each evaluation index is dynamically determined. Specifically, for each evaluation index, the sum of its causal connection strengths when it is an outcome variable in the causal relationship graph is analyzed, or the cumulative strength of its related causal paths is analyzed. This is used as prior information to correct the traditional entropy weight method, allowing those in key driving positions in the causal chain or those subject to more causal influence to receive higher weights. After determining the entropy weights of each evaluation index, the initial evaluation matrix is weighted and standardized based on these entropy weights to eliminate the influence of different index dimensions and numerical ranges, and to assign corresponding weights to the indexes. Finally, the TOPSIS method is used to calculate the relative fit progress of each evaluation object (i.e., the state of each component at each time step) with the positive and negative ideal solutions. The positive ideal solution is composed of the optimal values of each evaluation index, and the negative ideal solution is composed of the worst values of each evaluation index. The closer the relative fit progress is to 1, the closer the evaluation object is to the positive ideal solution, i.e., the healthier the equipment; the closer it is to 0, the more severely degraded or near-failure the equipment is. This relative fit progress is the final health index. Through this method, the health index not only reflects the current health level of the equipment but also incorporates cross-modal causal structure knowledge, enabling it to more sensitively capture dynamic changes on cascading degradation paths.
[0039] In one embodiment, based on a joint probability distribution and the current resource state extracted from a causal semantic digital twin, a pre-trained risk-aware meta-reinforcement learning agent generates and executes an intervention policy, including: The predicted values of failure probability, remaining useful life, and position drift in the joint probability distribution are concatenated with the current resource state extracted from the causal semantic digital twin to form a decision state vector. The decision state vector is then input into a pre-trained risk-aware reinforcement learning agent. The risk-aware reinforcement learning agent, with the goal of maximizing the task success rate and minimizing intervention costs and tail risks, performs multi-step deductions in a virtual simulation environment composed of the causal semantic digital twin to generate an intervention strategy. After the intervention strategy is safety verified, it is converted into an executable instruction sequence and sent to the space device for execution.
[0040] Specifically, firstly, the predicted information from the joint probability distribution is fused with the current resource status extracted from the causal semantic digital twin. Specifically, predicted values for failure probability, remaining useful life, and location drift are extracted from the joint probability distribution, serving as key indicators reflecting future risk situations. Simultaneously, the current resource status is extracted in real-time from the causal semantic digital twin. This current resource status includes, but is not limited to, remaining fuel, current power level, available redundant component list and quantity, and current task priority. The predicted values from the joint probability distribution are then concatenated with the parameters in the current resource status to form a comprehensive decision state vector. This vector simultaneously encompasses both the future risk situation and current capability constraints. Next, this decision state vector is input into a pre-trained risk-aware meta-reinforcement learning agent. This agent is pre-trained using a meta-reinforcement learning framework, enabling it to quickly adapt to new and unseen environmental conditions through minimal learning. Furthermore, a conditional risk value term is introduced into the agent's reward function, making it more sensitive to tail-end extreme risks (i.e., low-probability but severe events) in the joint probability distribution, thus proactively avoiding high-risk scenarios during decision-making. The agent performs multi-step deductions within a virtual simulation environment comprised of a causal semantic digital twin. This involves simulating the future evolutionary trajectory after executing different candidate intervention strategies in the virtual environment and evaluating the expected cumulative reward of each strategy considering risk constraints. The agent's decision-making objective is to maximize the task success rate while minimizing intervention costs (such as fuel consumption, power consumption, and redundant component usage) and tail risks. Through deduction and comparison, the agent selects the optimal intervention strategy under the current risk situation and resource constraints. After strategy generation, the generated intervention strategy undergoes safety verification. Safety verification includes, but is not limited to, checking whether the instructions exceed the physical limitations of the actuators (e.g., maximum thrust of the thrusters, range of motion of the robotic arms), whether they conflict with the currently executed task, and whether they will trigger other safety cascading risks. If the verification finds unsafe factors in the strategy, the strategy is modified or regenerated. Finally, the safety-verified intervention strategy is converted into an executable instruction sequence. This sequence includes specific control commands for each actuator of the space equipment (such as thrusters, robotic arms, power controllers, and communication equipment), organized in a temporal order. Command sequences are sent to the corresponding actuators of the space equipment via communication links, and the actuators then perform intervention operations. This method achieves a closed-loop decision-making process from risk perception and resource optimization to autonomous intervention, minimizing on-orbit failure rates and operational risks, and ensuring the long-term safe and stable operation of space equipment in complex and unknown environments.
[0041] In one embodiment, such as Figure 3As shown, a spatial operation and maintenance prediction system based on multimodal data is provided, including: Multimodal data acquisition module 10 is used to acquire multimodal data from space equipment; The digital twin construction module 20 is used to identify cross-modal causal relationship graphs based on structural causal models, and to perform time synchronization, spatial alignment and causal association mapping on multimodal data according to the cross-modal causal relationship graphs to construct a causal semantic digital twin. The feature fusion and health assessment module 30 is used to fuse features extracted from multimodal data in the causal semantic digital twin through a cross-modal causal attention mechanism, generate a multi-component coupled health status latent space vector, and calculate the health index based on the cross-modal causal relationship graph using the TOPSIS method. The hybrid prediction engine module 40 includes a feature-coupled encoder and an interactive prediction decoder, which combine the health status latent space vector and health index to output the joint probability distribution of failure probability, remaining lifetime and location drift. The decision-making and intervention module 50 is used to generate and execute intervention strategies by a pre-trained risk-aware reinforcement learning agent based on the joint probability distribution and the current resource status extracted from the causal semantic digital twin. The current resource status includes remaining fuel, electricity, available redundant components, and task priority.
[0042] Specifically, the space operation and maintenance prediction system based on multimodal data includes a multimodal data acquisition module 10, a digital twin construction module 20, a feature fusion and health assessment module 30, a hybrid prediction engine module 40, and a decision-making and intervention module 50. The multimodal data acquisition module 10 is used to collect multimodal data from space equipment. This multimodal data may include 3D point cloud data acquired by LiDAR, image data captured by visible light or infrared cameras, and various physical parameter data recorded by the internal telemetry system, providing a raw data foundation for subsequent state perception and causal analysis. The digital twin construction module 20 is connected to the multimodal data acquisition module 10. It is used to identify cross-modal causal relationship graphs based on structural causal models, and to perform time synchronization, spatial alignment, and causal association mapping on multimodal data according to the cross-modal causal relationship graphs, thereby constructing a causal semantic digital twin. Specifically, the module extracts the directed causal dependencies between different modal feature variables from historical multimodal data through a causal structure learning algorithm to form a cross-modal causal relationship graph. Then, it performs time axis alignment, coordinate system unification, and causal role labeling on the acquired multimodal data, and finally generates a digital twin with embedded cross-modal causal driving relationships, providing a high-fidelity virtual mirror environment for subsequent health assessment and prediction. The feature fusion and health assessment module 30 is connected to the digital twin construction module 20. It is used to fuse features extracted from multimodal data in the causal semantic digital twin through a cross-modal causal attention mechanism to generate a health status latent space vector coupled with multiple components. Based on the cross-modal causal relationship graph, the module calculates the health index using the TOPSIS method. This module first extracts geometric deformation features, surface degradation features, and time-frequency domain statistical features from point cloud, image, and telemetry data, respectively. Then, it uses the causal attention mechanism to calculate attention weights only for feature channels with causal connections and performs weighted fusion to obtain a low-dimensional dense health status latent space vector. Then, it uses each dimension of this vector as an evaluation index, takes the state of each component at different times as the evaluation object, and dynamically determines the entropy weight by combining the causal connection strength in the causal relationship graph. The health index reflecting the overall health level of the equipment is calculated using the TOPSIS method. The hybrid prediction engine module 40 is connected to the feature fusion and health assessment module 30. This module includes a feature coupling encoder and an interactive prediction decoder, which combine the health status latent space vector and the health index to output a joint probability distribution of failure probability, remaining useful life, and location drift. The feature coupling encoder performs time alignment and deep fusion of the health status latent space vector, health index, historical trajectory dataset, and spatial environment parameters to generate coupled features. The failure prediction branch and the location change prediction branch in the interactive prediction decoder interact through multiple rounds of iteration to finally output a joint probability distribution that simultaneously describes the dependencies of the three variables: failure probability, remaining useful life, and location drift.The decision-making and intervention module 50 is connected to the hybrid prediction engine module 40 and the digital twin construction module 20. It is used to generate and execute intervention strategies by a pre-trained risk-aware reinforcement learning agent based on the joint probability distribution and the current resource status extracted from the causal semantic digital twin. The current resource status includes remaining fuel, power, available redundant components and task priority. This module concatenates the predicted values in the joint probability distribution with the current resource status to form a decision state vector, which is input into the pre-trained risk-aware reinforcement learning agent. The agent performs multi-step deduction in the virtual simulation environment composed of the causal semantic digital twin to generate the optimal intervention strategy with the goal of maximizing the task success rate and minimizing the intervention cost and tail risk. After safety verification, it is converted into an executable instruction sequence and sent to the space device for execution.
[0043] Through the collaborative work of the above five modules, the space operation and maintenance prediction system based on multimodal data can realize a complete closed loop from data acquisition, digital twin construction, health status assessment, multi-indicator joint prediction to autonomous decision-making intervention, effectively improving the reliability and autonomy of space equipment in orbit.
[0044] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for space operation and maintenance prediction based on multi-modal data, characterized in that, include: Multimodal data from space devices are collected, cross-modal causal relationship graphs are identified based on structural causal models, and time synchronization, spatial alignment, and causal association mapping are performed on the multimodal data according to the cross-modal causal relationship graphs to construct a causal semantic digital twin. In the causal semantic digital twin, features extracted from the multimodal data are fused through a cross-modal causal attention mechanism to generate a multi-component coupled health status latent space vector, and a health index is calculated based on the cross-modal causal relationship graph using the TOPSIS method. A hybrid prediction engine comprising a feature-coupled encoder and an interactive prediction decoder is constructed, which combines the health state latent space vector and the health index to output the joint probability distribution of failure probability, remaining lifetime and position drift. Based on the joint probability distribution and the current resource status extracted from the causal semantic digital twin, an intervention strategy is generated and executed by a pre-trained risk-aware meta-reinforcement learning agent, wherein the current resource status includes remaining fuel, power, available redundant components, and task priority.
2. The method of claim 1, wherein, The construction of a hybrid prediction engine comprising a feature-coupled encoder and an interactive predictive decoder, combining the health state latent space vector and the health index, outputs a joint probability distribution of failure probability, remaining lifetime, and location drift, including: Construct a hybrid prediction engine that includes a feature-coupled encoder and an interactive prediction decoder; The historical trajectory dataset and spatial environment parameters are obtained, and the health status latent space vector, the health index, the historical trajectory dataset and the spatial environment parameters are time-aligned and then input into the feature coupling encoder to generate coupled features. The coupling features are input into the interactive prediction decoder, which includes a fault prediction branch and a location change prediction branch. Through iterative interaction between the fault prediction branch and the location change prediction branch, the joint probability distribution of fault probability, remaining service life, and location drift is finally output.
3. The method of claim 2, wherein, The iterative interaction between the fault prediction branch and the location change prediction branch, ultimately outputting the joint probability distribution of fault probability, remaining useful life, and location drift, includes: Based on the coupling characteristics, fault evolution inference is performed through the fault prediction branch, the change in physical parameters caused by component degradation is output, and the change in physical parameters is transmitted to the position change prediction branch. The position change prediction branch corrects the internal dynamic model parameters based on the received changes in physical parameters, and performs trajectory and attitude inference based on the corrected dynamic model parameters, outputting vibration parameters and attitude deviation parameters. The vibration parameters and the attitude deviation parameters are used as acceleration factors and fed back from the position change prediction branch to the fault prediction branch to adjust the degradation accumulation rate. The fault prediction branch and the location change prediction branch interact in multiple rounds of iteration until the preset convergence condition is met, and finally output the joint probability distribution of fault probability, remaining service life and location drift.
4. The method of claim 3, wherein, The position change prediction branch corrects its internal dynamic model parameters based on the received changes in physical parameters, and performs trajectory and attitude inference based on the corrected dynamic model parameters, outputting vibration parameters and attitude deviation parameters, including: The position change prediction branch receives the physical parameter changes transmitted by the fault prediction branch, the physical parameter changes including the centroid offset and inertia tensor changes caused by component degradation; Based on the centroid offset and the inertia tensor change, correct the centroid position parameters and inertia matrix parameters in the internal dynamic model of the position change prediction branch; Based on the corrected centroid position parameters and inertia matrix parameters, combined with the space environment parameters, orbit extrapolation and attitude evolution calculations are performed on the space equipment to obtain the predicted position and attitude values in the future time domain. The predicted position and the predicted attitude are compared with the nominal values, and the vibration amplitude and vibration frequency components are extracted as vibration parameters, and the attitude angle deviation and attitude drift rate are extracted as attitude deviation parameters. Output the vibration parameters and the attitude deviation parameters.
5. The spatial operation and maintenance prediction method based on multimodal data according to claim 1, characterized in that, In the causal semantic digital twin, features extracted from the multimodal data are fused using a cross-modal causal attention mechanism to generate a multi-component coupled health state latent space vector. Based on the cross-modal causal relationship graph, a health index is calculated using the TOPSIS method, including: In the causal semantic digital twin, geometric deformation features of point clouds, surface degradation features of images, and time-frequency domain statistical features of telemetry data are extracted from the multimodal data, respectively. Based on the cross-modal causal relationship diagram, determine whether there is a causal connection between different modal feature channels of the geometric deformation features of the point cloud, the surface degradation features of the image, and the time-frequency domain statistical features of the telemetry data; Through a cross-modal causal attention mechanism, attention weights are calculated only for feature channels with causal connections, and features extracted from the multimodal data are weighted and fused based on the attention weights to generate a multi-component coupled health state latent space vector. Based on the latent space vector of the health status, the health index is calculated using the TOPSIS method through the cross-modal causal relationship graph.
6. The spatial operation and maintenance prediction method based on multi-modal data according to claim 5, characterized in that, The step of using a cross-modal causal attention mechanism to calculate attention weights only for feature channels with causal connections, and then weighting and fusing features extracted from the multimodal data based on these attention weights to generate a multi-component coupled health state latent space vector includes: By using a cross-modal causal attention mechanism, the causal connections between different modal feature channels of the point cloud's geometric deformation features, the image's surface degradation features, and the time-frequency domain statistical features of the telemetry data are obtained from the cross-modal causal relationship graph. Based on the causal connections, a query vector, key vector, and value vector are constructed for each feature channel, and attention scores are calculated only for feature channels with causal connections to generate an attention weight matrix with causal constraints. Based on the attention weight matrix, the value vectors of each feature channel extracted from the multimodal data are weighted and summed to obtain the health state latent space vector of the multi-component coupling.
7. The method of claim 5, wherein, The calculation of the health index based on the latent space vector of the health status and using the cross-modal causal relationship graph via the TOPSIS method includes: An evaluation matrix is constructed using each dimension of the latent space vector of the health status as an evaluation index and the states of each component in the multimodal data at different times as evaluation objects. The causal connection strength between each modal feature channel is extracted from the cross-modal causal relationship graph, and the entropy weight of each evaluation index is dynamically determined based on the causal connection strength. The evaluation matrix is weighted and standardized based on the entropy weight, and the relative alignment of each component with the positive and negative ideal solutions is calculated using the TOPSIS method to obtain the health index.
8. The method of claim 1, wherein, The step of generating and executing an intervention strategy by a pre-trained risk-aware meta-reinforcement learning agent based on the joint probability distribution and the current resource state extracted from the causal semantic digital twin includes: The predicted values of failure probability, remaining useful life, and location drift in the joint probability distribution are concatenated with the current resource state extracted from the causal semantic digital twin to form a decision state vector. The decision state vector is input into a pre-trained risk-aware meta-reinforcement learning agent. The risk-aware meta-reinforcement learning agent aims to maximize the task success rate and minimize the intervention cost and tail risk. It performs multi-step deduction in the virtual simulation environment composed of the causal semantic digital twin to generate an intervention strategy. After performing a security verification on the intervention strategy, the intervention strategy is converted into an executable instruction sequence and sent to the space device for execution. 9.A space operation and maintenance prediction system based on multi-modal data, characterized in that, include: Multimodal data acquisition module, used to acquire multimodal data from space equipment; A digital twin construction module is used to identify cross-modal causal relationship graphs based on structural causal models, and to perform time synchronization, spatial alignment, and causal association mapping on the multimodal data according to the cross-modal causal relationship graphs to construct a causal semantic digital twin. The feature fusion and health assessment module is used to fuse features extracted from the multimodal data in the causal semantic digital twin through a cross-modal causal attention mechanism, generate a multi-component coupled health status latent space vector, and calculate the health index based on the cross-modal causal relationship graph using the TOPSIS method. The hybrid prediction engine module includes a feature-coupled encoder and an interactive prediction decoder, which combine the health state latent space vector and the health index to output a joint probability distribution of failure probability, remaining lifetime and location drift. The decision-making and intervention module is used to generate and execute intervention strategies by a pre-trained risk-aware reinforcement learning agent based on the joint probability distribution and the current resource status extracted from the causal semantic digital twin. The current resource status includes remaining fuel, power, available redundant components, and task priority.