Satellite-borne mechanical arm state prediction model training method and device, medium and equipment
Patent Information
- Application Number
- CN202611327396.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
然而,如果AI的决策结果不符合太空微重力环境下的物理动力学规律,则容易导致星载机械臂任务执行失败,甚至发生碰撞造成毁损
[0009]借由上述技术方案,本申请提供的一种星载机械臂的状态预测模型训练方法、装置、介质及设备,与现有技术相比,本申请通过在损失函数中引入轨道动力学约束损失和多体运动学约束损失,在模型决策和星载机械臂执行之间构建一个物理推演层,能够使决策在执行前经过动力学一致性验证,与此同时,本申请还在预测阶段对模型输出的预测状态序列进行动力学残差修正,通过训练阶段和预测阶段的双重物理约束机制,本申请能够使模型决策结果符合太空微重力环境下的物理动力学规律,从而能够提升星载机械臂任务执行的成功率,避免发生碰撞造成星载机械臂损毁。此外,本申请构建的状态预测模型不仅能够推演动作序列的执行后果,还能够预测动作序列的执行奖励,依据该预测奖励,能够快速锁定最优候选动作,从而能够提高星载机械臂的控制精度和控制效率。
Smart Images

Figure CN122840173A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular to a method, apparatus, medium and equipment for training a state prediction model of a spaceborne robotic arm. Background Technology
[0002] On-Orbit Servicing (OOS) refers to a technological system that utilizes spacecraft to perform maintenance, repair, resupply, life extension, and deorbiting operations on other spacecraft or space targets while in orbit. With the rapid increase in the number of satellites in low Earth orbit and the growing severity of space debris issues, on-orbit servicing technology has become an important development direction in the aerospace field. Among these, the onboard robotic arm, as the core execution mechanism for on-orbit servicing, undertakes key tasks such as space debris removal, satellite repair, on-orbit assembly, and refueling. Its autonomous control capability directly determines the efficiency and safety of on-orbit servicing.
[0003] Currently, AI (Artificial Intelligence) decision-making modules typically output control commands directly. However, if the AI's decisions do not conform to the physical dynamics of the microgravity environment in space, it can easily lead to the failure of the spaceborne robotic arm's mission, or even collisions causing damage. Summary of the Invention
[0004] In view of this, this application provides a method, device, medium and equipment for training a state prediction model of a spaceborne robotic arm, which mainly enables the model decision results to conform to the physical dynamic laws of the microgravity environment in space, thereby improving the success rate of the spaceborne robotic arm in performing tasks.
[0005] According to a first aspect of this application, a method for training a state prediction model for a spaceborne robotic arm is provided, the method comprising: Collect state transition samples when the spaceborne robotic arm grasps target objects in different scenarios. The state transition samples include sample states, sample action sequences, and actual state sequences. An initial state prediction model is constructed, and the sample state and the sample action sequence are input into the initial state prediction model to perform state prediction, thereby obtaining the training prediction state sequence and training prediction reward corresponding to the sample action sequence. A multi-dimensional reward evaluation is performed on the training prediction state sequence to obtain the actual reward corresponding to the training prediction state sequence. Based on the training predicted state sequence and the actual state sequence, as well as the training predicted reward and the actual reward, a total loss function is constructed, wherein the total loss function embeds a dynamic constraint loss, which includes orbital dynamics constraint loss and multibody kinematics constraint loss. Based on the total loss function, the initial state prediction model is iteratively trained to construct a preset state prediction model. In the prediction phase, the predicted state sequence output by the preset state prediction model is obtained after dynamic residual correction.
[0006] According to a second aspect of this application, a state prediction model training device for a spaceborne robotic arm is provided, the device comprising: The acquisition unit is used to acquire state transition samples when the spaceborne robotic arm grasps the target object in different scenarios. The state transition samples include sample states, sample action sequences, and actual state sequences. The prediction unit is used to construct an initial state prediction model and input the sample state and the sample action sequence into the initial state prediction model to perform state prediction, thereby obtaining the training prediction state sequence and training prediction reward corresponding to the sample action sequence. An evaluation unit is used to perform multi-dimensional reward evaluation on the training prediction state sequence to obtain the actual reward corresponding to the training prediction state sequence. A construction unit is used to construct a total loss function based on the training predicted state sequence and the actual state sequence, as well as the training predicted reward and the actual reward, wherein the total loss function embeds a dynamic constraint loss, which includes orbital dynamics constraint loss and multibody kinematics constraint loss; The training unit is used to iteratively train the initial state prediction model according to the total loss function to construct a preset state prediction model, wherein the predicted state sequence output by the preset state prediction model in the prediction stage is obtained after dynamic residual correction.
[0007] According to a third aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described method for training the state prediction model of a spaceborne robotic arm.
[0008] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described method for training a state prediction model for a spaceborne robotic arm.
[0009] By employing the above technical solutions, this application provides a method, apparatus, medium, and device for training a state prediction model for a spaceborne robotic arm. Compared with existing technologies, this application introduces orbital dynamics constraint loss and multibody kinematics constraint loss into the loss function, constructing a physical deduction layer between model decision-making and spaceborne robotic arm execution. This allows decisions to undergo dynamic consistency verification before execution. Simultaneously, this application also performs dynamic residual correction on the predicted state sequence output by the model during the prediction phase. Through the dual physical constraint mechanism of the training and prediction phases, this application ensures that the model decision results conform to the physical dynamic laws of the microgravity environment in space, thereby improving the success rate of spaceborne robotic arm mission execution and preventing collisions that could damage the spaceborne robotic arm. Furthermore, the state prediction model constructed in this application can not only deduce the execution consequences of action sequences but also predict the execution rewards of action sequences. Based on this predicted reward, the optimal candidate action can be quickly identified, thereby improving the control accuracy and efficiency of the spaceborne robotic arm.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a state prediction model training method for a spaceborne robotic arm provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating the preset state prediction model training method provided in an embodiment of this application is shown. Figure 3 A flowchart illustrating the actual reward calculation method provided in the embodiments of this application is shown; Figure 4 A schematic diagram of the system architecture for state prediction provided in an embodiment of this application is shown; Figure 5 This paper illustrates a flowchart of the execution deviation correction process provided in an embodiment of this application. Figure 6 This paper shows a schematic diagram of the structure of a state prediction model training device for a spaceborne robotic arm provided in an embodiment of this application. Detailed Implementation
[0012] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0013] The decision-making results of existing AI technology do not conform to the physical dynamics of the microgravity environment in space, which can easily lead to the failure of the spaceborne robotic arm mission, or even collisions that cause damage.
[0014] To address the aforementioned problems, embodiments of the present invention provide a method for training a state prediction model for a spaceborne robotic arm, such as... Figure 1 As shown, the method includes: Step 10: Collect state transition samples when the spaceborne robotic arm grasps the target object in different scenarios. The state transition samples include sample states, sample action sequences, and actual state sequences.
[0015] The target objects specifically include defunct satellites, space debris, etc. It should be noted that the target objects in this embodiment of the invention are not limited to the above-listed examples; they can be any object in space that needs to be grasped by the spaceborne robotic arm. Furthermore, the different scenarios in this embodiment of the invention include not only scenarios where the spaceborne robotic arm grasps the target object, but also all scenarios of interaction between the spaceborne robotic arm and the target object, such as alignment and striking. The sample action sequence includes multiple consecutive time steps of sample actions, specifically including retraction, extension, approach, contact, and grasping. The sample state is the state vector of the spaceborne robotic arm and the target object in any training time step. The actual state sequence is the state vector of the spaceborne robotic arm and the target object in subsequent training time steps. For example, if the sample state is the state vector of training time step 1, the actual state sequence includes the state vectors of training time steps 2, 3, 4, and 5. This state vector is used to represent the state of the target object and the spaceborne robotic arm, specifically including the contour information, pose, velocity, and angular velocity of the target object, the angle, angular velocity, and end-effector pose of each joint of the spaceborne robotic arm, and the contact force and torque between the end-effector of the spaceborne robotic arm and the target object.
[0016] Before training the model, sample data is collected at different training time steps. Specifically, visual images of the target object are acquired using a camera device (such as a spaceborne visual camera). The target object in the visual image is detected, and based on the target detection results, the target region in the visual image is located. Then, edge detection is performed on the target region to obtain the shape contour information of the target object. Simultaneously, the target region image is cropped and input into a preset pose estimation network for pose estimation, obtaining the position components of the target object's centroid on the x, y, and z axes, as well as its three-dimensional rotational pose (3 pose components). The preset pose estimation network specifically includes a CNN convolutional network and a two-branch MLP multilayer perceptron. The two-branch MLP multilayer perceptron is used to output the position and pose components, respectively. Based on these position and pose components, the pose of the target object can be determined. Furthermore, based on the position components of the target object in the current frame and the previous frame, the velocity of the target object can be determined. Similarly, based on the pose components of the target object in the current frame and the previous frame, the angular velocity of the target object can be determined. Thus, the contour information, pose, velocity, and angular velocity of the target object can be obtained.
[0017] Simultaneously, angle sensors are used to collect the angles of each joint of the spaceborne robotic arm, and the angular velocity of each joint is calculated based on the angles of two adjacent training time steps. After obtaining the angles of each joint, these angles are substituted into the position forward kinematics function. The three-dimensional position coordinates of the end effector of the spaceborne robotic arm are obtained. Similarly, the angles of each joint are substituted into the attitude forward kinematics function. The attitude of the spaceborne robotic arm's end effector is obtained, representing its orientation, such as roll, pitch, and yaw. Based on the three-dimensional position coordinates and attitude of the end effector, its pose can be determined. Furthermore, the contact force and torque between the end effector and the target object are collected using an end effector torque sensor. This allows for the acquisition of the angles, angular velocities, and end effector pose of each joint of the spaceborne robotic arm, as well as the contact force and torque between the end effector and the target object.
[0018] In addition, three-dimensional point cloud data of the target object and the distance to the target object are obtained through lidar or millimeter-wave radar, and the attitude and angular velocity of the spacecraft body are obtained through inertial measurement unit (IMU).
[0019] After collecting multimodal sample data within the above-mentioned different training time steps, the multimodal sample data is processed by time synchronization, coordinate transformation, denoising and normalization. The processed sample data is then encoded into state vectors of the spaceborne robotic arm and the target object, and then divided into sample states and actual state sequences according to the training time steps.
[0020] Step 20: Construct an initial state prediction model, and input the sample state and the sample action sequence into the initial state prediction model to perform state prediction, so as to obtain the training prediction state sequence and training prediction reward corresponding to the sample action sequence.
[0021] The initial state prediction model includes an initial world model and an initial reward prediction model. The initial world model specifically includes an initial state encoder, an initial transition prediction model, and an initial state decoder. The initial state encoder includes fully connected layers, and the structure of the initial state decoder corresponds to that of the encoder. The initial transition prediction model specifically includes recurrent neural networks (such as GRU / LSTM) or state-space models (such as Mamba / S4), and the initial reward prediction model specifically includes small MLP networks.
[0022] Specifically, such as Figure 2 As shown, the initial state encoder includes a fully connected layer, which is used to process the input sample state vector. s ( t Compressed into low-dimensional sample hidden state vectors z ( t For example, the dimension of the hidden state vector of this low-dimensional sample is 256 or 512. Then, for the sample action sequence { , , ..., The initial transition prediction model is based on the sample hidden state vector. z ( t And for each step in the sample action sequence, recursively predict the future. H The hidden state sequence at each training time step { , , ..., The specific formula is as follows:
[0023] in, This is a random disturbance term used to model environmental uncertainties. For network parameters, Based on network parameters The initial transition prediction model is obtained. Therefore, by recursively applying the above formula, the hidden state vector for each future training time step can be obtained.
[0024] Next, the initial state decoder is used to process the hidden state sequence { , , ..., } Decoded into an interpretable state sequence, i.e., the training and prediction state sequence { , , ..., The training prediction state sequence includes the predicted pose (including predicted position and predicted attitude), predicted velocity and predicted angular velocity of the target object in each future training time step, the predicted angle, predicted angular velocity and predicted end-effector pose (including predicted position and predicted attitude) of each joint of the spaceborne robotic arm, and the predicted contact force and predicted torque between the end-effector of the spaceborne robotic arm and the target object.
[0025] Step 30: Perform multi-dimensional reward evaluation on the training prediction state sequence to obtain the actual reward corresponding to the training prediction state sequence.
[0026] The preset state prediction model constructed in this embodiment of the invention can not only predict the candidate state sequence corresponding to each candidate action sequence, but also predict the predicted reward of the candidate state sequence. Based on this, this embodiment of the invention needs to introduce a reward loss when constructing the total loss function, and the calculation of the reward loss needs to be combined with the actual reward. Regarding the calculation process of the actual reward, as follows... Figure 3 As shown, it includes: Step 31: Based on the training prediction state sequence, determine the predicted angle and predicted torque of each joint of the spaceborne robotic arm, the predicted position of the end effector of the spaceborne robotic arm, the predicted position of the target object, and the number of task completion time steps within each training time step.
[0027] In this embodiment of the invention, the predicted angles and torques of each joint of the spaceborne robotic arm, the predicted position of the end effector of the spaceborne robotic arm, the predicted position of the target object, and the number of task completion time steps can be obtained from the training prediction state sequence output by the initial state prediction model.
[0028] Step 32: Calculate the task completion reward based on the predicted position of the end effector of the spaceborne robotic arm and the predicted position of the target object.
[0029] In this embodiment of the invention, based on the predicted position of the end effector of the spaceborne robotic arm and the predicted position of the target object, the negative value of the Euclidean distance between the end effector position and the target object position is calculated. This negative value is the task completion reward. The smaller the distance between the two, the higher the task completion reward. When the end effector is close to the target object, an additional positive reward can be given.
[0030] Step 33: Calculate the collision probability based on the predicted angles of each joint of the spaceborne robotic arm, and determine the collision safety reward based on the collision probability.
[0031] In this embodiment of the invention, the collision safety reward is equal to the negative of the collision probability. That is, the higher the collision probability, the higher the risk of a collision and the lower the collision safety reward. Conversely, the lower the collision probability, the lower the risk of a collision and the higher the collision safety reward.
[0032] When calculating the collision probability, the fixed structural parameters of the spaceborne robotic arm and the geometric dimensions of each link, as well as the three-dimensional point cloud data of the target object, are obtained. Based on the fixed structural parameters, the geometric dimensions of each link, and the predicted angles of each joint of the spaceborne robotic arm, a directed bounding box is constructed for each link. Based on the three-dimensional point cloud data, a point cloud model of the target object is constructed. Based on the directed bounding box of each link and the point cloud model of the target object, the minimum distance between each link and the target object is calculated. Based on the minimum distance, the collision probability is calculated.
[0033] When constructing the directed bounding box for each link, the transformation matrix of adjacent links of the spaceborne robotic arm is calculated using a preset forward kinematics function based on the fixed structural parameters and the predicted angles of each joint of the spaceborne robotic arm. Based on the transformation matrix of adjacent links, the pose of each link in the world base coordinate system is calculated cumulatively. Based on the pose of each link in the world base coordinate system and the geometric dimensions of each link, the directed bounding box of each link is constructed.
[0034] Specifically, the formula for calculating the transformation matrix of adjacent links is as follows:
[0035] in, The transformation matrix of adjacent links is used to represent the link coordinate system. i Relative to the link coordinate system i -1 position and orientation; Represents the kinematic function; The fixed structural parameters representing the spaceborne robotic arm include the link length, which is determined by the factory settings. The training prediction angles represent the joints of the spaceborne robotic arm.
[0036] After calculating the transformation matrix of adjacent links according to the above formula, the pose of each link in the world base coordinate system is calculated cumulatively. For example, This represents the pose of link 1 in the world base coordinate system; , This represents the pose of link 2 in the world base coordinate system; , This represents the pose of link 3 in the world base coordinate system. Similarly, the pose of each link in the world base coordinate system can be calculated.
[0037] Next, the position and orientation of each link are extracted from its pose in the world base coordinate system to obtain its spatial position and orientation. Then, based on the geometric dimensions (such as length, width, and height) of each link as described in the onboard robotic arm's manual, along with its spatial position and orientation, a directed bounding box is constructed for each link. Simultaneously, a point cloud model of the target object is built using the acquired 3D point cloud data. Finally, based on the directed bounding box of each link and the point cloud model of the target object, the minimum distance between each link and the target object is calculated, thereby calculating the collision probability. The specific formula for calculating the collision probability is as follows. P = 1 / (1 + exp( k ·( d min - d threshold ))) in, P Represents the probability of collision; k This represents the distance coefficient, which can be set according to actual business needs; d min Represents the minimum distance; d threshold This represents the distance threshold, which can be set according to actual business needs.
[0038] Step 34: Calculate the energy consumption reward based on the predicted torque of each joint of the spaceborne robotic arm.
[0039] In this embodiment of the invention, the energy reward is equal to the negative of the sum of the squares of the predicted torques of each joint of the spaceborne robotic arm, thereby encouraging the selection of low-energy motion sequences.
[0040] Step 35: Calculate the time efficiency bonus based on the number of time steps and time step length for completing the task; In this embodiment of the invention, the time efficiency reward is equal to the negative value of the product of the task completion time steps and the time step length, thereby encouraging the completion of the task in a shorter time.
[0041] Step 36: Calculate the actual reward based on the task completion reward, the collision safety reward, the energy consumption reward, and the time efficiency reward.
[0042] In this embodiment of the invention, after calculating the collision safety reward, task completion reward, energy consumption reward, and time efficiency reward, the actual reward corresponding to the sample action sequence is calculated based on these rewards. The specific calculation formula is as follows.
[0043] in, Represents actual reward; Rewards represent task completion percentages; Represents a collision safety award; Represents energy consumption reward; Represents time efficiency rewards; , , and These are all weighting coefficients, which can be dynamically adjusted according to the mission type and safety level. In on-orbit operations with extremely high safety requirements... Much larger , and In urgent emergency missions where time is of the essence, It can be increased appropriately.
[0044] Step 40: Construct a total loss function based on the training predicted state sequence and the actual state sequence, as well as the training predicted reward and the actual reward.
[0045] The total loss function includes embedded dynamic constraint loss, which includes orbital dynamic constraint loss and multibody kinematic constraint loss.
[0046] In this embodiment of the invention, when calculating the total loss function, a reconstruction loss is constructed based on the training predicted state sequence and the actual state sequence; a divergence constraint loss is calculated based on the hidden state sequence corresponding to the training predicted state sequence; a dynamic constraint loss is calculated based on the training predicted state sequence and the sample state; a reward loss is calculated based on the training predicted reward and the actual reward; and a total loss function is constructed based on the reconstruction loss, the divergence constraint loss, the dynamic constraint loss, and the reward loss. The specific formula for the total loss function is as follows.
[0047] in, Represents the total loss; Represents reconstruction losses; Represents the divergence constraint loss; Represents dynamic constraint loss; Represents a loss of reward; , , and These represent weighting coefficients, which can be set according to actual business needs.
[0048] For reconstruction loss, based on the training predicted state sequence and the actual state sequence, the pose loss, velocity loss, and angular velocity loss of the target object, the angle loss, angular velocity loss, and end-effector pose loss of each joint of the spaceborne robotic arm, and the contact force loss between the end-effector of the spaceborne robotic arm and the target object are determined for each training time step. These various losses are accumulated over multiple training time steps, and the accumulated losses of each type are concatenated into a loss vector. The Euclidean norm 2 of this loss vector is calculated to obtain the reconstruction loss. The divergence constraint loss describes the divergence between the hidden state sequence corresponding to the training predicted state sequence and the prior distribution. By making the hidden state sequence closer to the prior distribution, generalization can be improved. The reward loss is the absolute value of the difference between the training predicted reward and the actual reward. The dynamic constraint loss is the sum of the orbital dynamic constraint loss and the multibody kinematics constraint loss. The calculation of the orbital dynamic constraint loss and the multibody kinematics constraint loss is mainly to ensure that the predicted state sequence conforms to physical constraints.
[0049] When calculating the dynamic constraint loss, based on the training prediction state sequence, the predicted pose of the end effector of the spaceborne robotic arm, as well as the predicted position and velocity of the target object, are determined within each training time step; according to the sample state, the updated position and velocity of the target object are calculated using the orbital dynamics equations; based on the predicted position and velocity of the target object, and the updated position and velocity of the target object, the orbital dynamic constraint loss is calculated; based on the training prediction state sequence, the updated position and updated attitude of the end effector of the spaceborne robotic arm are calculated using the positive kinematics function of position and the positive kinematics function of attitude, respectively; based on the predicted pose of the end effector of the spaceborne robotic arm, and the updated position and updated attitude of the end effector of the spaceborne robotic arm, the multibody kinematics constraint loss is calculated; based on the orbital dynamic constraint loss and the multibody kinematics constraint loss, the dynamic constraint loss is calculated.
[0050] When calculating the orbital dynamics constraint loss, the velocity and position of the target object can be updated using the orbital dynamics equations, as shown in the following formula:
[0051] in, Represents the velocity of the target object within any training time step in the sample state; This represents the update rate of the target object within the next training time step; Representing perturbation acceleration, in the context of low-Earth orbit satellites, You can choose 1-5. mm / ; This represents the training time step; It represents the position of the target object within any training time step in the sample state; This represents the updated position of the target object within the next training time step.
[0052] After determining the update velocity and update position of the target object, based on the predicted position and velocity of the target object in the training prediction state sequence, as well as the update velocity and update position of the target object, the position dynamics constraint loss and velocity dynamics constraint loss of the target object can be calculated. The position dynamics constraint loss and velocity dynamics constraint loss of the target object are accumulated over multiple training time steps to obtain the total position dynamics constraint loss and total velocity dynamics constraint loss of the target object. Then, the total position dynamics constraint loss and total velocity dynamics constraint loss are concatenated into a dynamics constraint vector, and the Euclidean L2 norm of the dynamics constraint vector is calculated to obtain the orbital dynamics constraint loss of the target object.
[0053] Meanwhile, when calculating the multibody kinematics constraint loss, the position and attitude of the spaceborne robotic arm can be updated using the position positive kinematics function and the attitude positive kinematics function, as shown in the following formulas.
[0054]
[0055] in, This represents the updated position of the end effector of the spaceborne robotic arm; Represents the updated posture of the end effector of the spaceborne robotic arm; This represents the predicted angles of each joint of the spaceborne robotic arm within each training time step in the training predicted state sequence. n Represents the number of joints; and These represent the positive kinematics function of position and the positive kinematics function of attitude, respectively.
[0056] After determining the updated position and updated posture of the end effector of the spaceborne robotic arm, based on the training predicted pose (including training predicted position and training predicted posture) of the end effector in the training predicted state sequence, as well as the updated position and updated posture of the end effector, the position kinematic constraint loss and posture kinematic constraint loss of the end effector can be calculated. The position kinematic constraint loss and posture kinematic constraint loss of the end effector of the spaceborne robotic arm are accumulated over multiple training time steps to obtain the total position kinematic constraint loss and total posture kinematic constraint loss of the end effector of the spaceborne robotic arm. Then, the total position kinematic constraint loss and the total posture kinematic constraint loss are concatenated into a kinematic constraint vector, and the Euclidean norm 2 of the kinematic constraint vector is calculated to obtain the multibody kinematic constraint loss of the spaceborne robotic arm.
[0057] Step 50: Based on the total loss function, iteratively train the initial state prediction model to construct a preset state prediction model.
[0058] In the prediction phase, the predicted state sequence output by the preset state prediction model is obtained after dynamic residual correction.
[0059] In this embodiment of the invention, based on the constructed total loss function, the initial state prediction model is iteratively trained until a preset number of iterations is reached or the total loss is less than a preset threshold, at which point iteration stops and a preset state prediction model is output. This embodiment of the invention allows for continuous training of the complete preset state prediction model using large-scale algorithms on the ground, and periodically compresses the model parameters into a lightweight version through knowledge distillation. The weights of the state prediction model in the onboard AI unit are then updated via uplink. During the distillation process, the core dynamics of the preset state prediction model are preserved, while the number of parameters is significantly reduced to adapt to onboard computing power constraints.
[0060] After training the preset state prediction model, state prediction can be performed using this model. The method further includes: obtaining the current state representation vector shared by the spaceborne robotic arm and the target object, and multiple candidate action sequences of the spaceborne robotic arm; inputting the current state representation vector and each candidate action sequence into the preset state prediction model for state prediction, obtaining a predicted state sequence and a predicted reward for each candidate action sequence; determining a target action sequence from the multiple candidate action sequences based on the predicted reward; solving for the control parameters when the end effector of the spaceborne robotic arm contacts the target object based on the target predicted state sequence corresponding to the target action sequence; controlling the spaceborne robotic arm to grasp the target object based on the control parameters and the target action sequence, and collecting the actual state sequence shared by the spaceborne robotic arm and the target object; calculating a multidimensional execution deviation based on the target predicted state sequence and the actual state sequence, and correcting the multidimensional execution deviation using a corresponding correction strategy.
[0061] The current state representation vector characterizes the current state of the spaceborne robotic arm and the target object. Specifically, it includes the target object's contour information, pose, velocity, and angular velocity; the angles, angular velocities, and end-effector poses of each joint of the spaceborne robotic arm; the contact force and torque between the end-effector and the target object; and the distance to the target object, the attitude, and angular velocity of the spacecraft itself. Each candidate action sequence includes multiple consecutive time steps of action, specifically including retraction, extension, approach, contact, and grasping. Furthermore, the preset state prediction model includes a state encoder, a transition prediction model, a state decoder, and a reward prediction model.
[0062] To achieve control of the spaceborne robotic arm, embodiments of the present invention provide a spaceborne robotic arm control system, such as... Figure 4As shown, the system comprises a perception module, a deduction module, and a decision-making and execution module. The perception module acquires multimodal state information of the target object and the onboard robotic arm, fusing and encoding it into a unified representation state for use by the deduction module. The deduction module, the core module, internally deduces possible predicted states over multiple future time steps based on the current states of the onboard robotic arm and the target object, as well as candidate action sequences, providing a forward-looking basis for the decision-making and execution module. The decision-making and execution module evaluates the risks and benefits of each candidate action sequence based on the deduction results, selecting the optimal action sequence for the robotic arm to execute, while continuously monitoring and correcting prediction deviations. These three modules are tightly coupled through data flow; the output of the perception module is the input of the deduction module, the output of the deduction module is the basis for the decision-making and execution module, and the execution feedback from the decision-making and execution module flows back to the perception module, forming a complete closed loop. In this embodiment of the invention, the output of the deduction module is no longer a simple control command, but a future state sequence of the spaceborne robotic arm. The state data of the robotic arm during the actual execution process serves as the constraint condition for the deduction module to predict the state at the next moment, thereby realizing the data coupling of deduction, execution and feedback.
[0063] When the spaceborne robotic arm performs a task, it can collect multimodal raw data of the target object and the spaceborne robotic arm based on the sensor data acquisition method in step 10. Then, the multimodal raw data is processed by time synchronization, coordinate transformation, denoising and normalization, and the processed data is encoded into the current state representation vector of the spaceborne robotic arm and the target object.
[0064] After obtaining the current state representation vector, the current state representation vector is input into a preset state prediction model for state prediction. Specifically, during state prediction, the current state representation vector is input into the state encoder for dimensionality reduction to obtain the hidden state vector corresponding to the current state representation vector; based on the hidden state vector and each candidate action sequence, the transition prediction model predicts the future hidden state sequence; the state decoder decodes the hidden state sequence into an interpretable state sequence; dynamic residual correction is applied to the interpretable state sequence to obtain the predicted state sequence corresponding to each candidate action sequence; the predicted state sequence corresponding to each candidate action sequence is input into the reward prediction model for reward prediction to obtain the predicted reward corresponding to each candidate action sequence.
[0065] Because the transfer prediction model in this embodiment of the invention incorporates microgravity dynamics priors, it ensures that the predicted hidden state evolution trajectory conforms to physical laws during the recursive process. Furthermore, the preset state prediction model in this embodiment can not only predict the candidate state sequence corresponding to each candidate action sequence, but also predict the prediction reward of the candidate state sequence. Based on this prediction reward, the optimal candidate action can be quickly identified, thereby improving the control accuracy and efficiency of the spaceborne robotic arm.
[0066] To ensure that the derived interpretable state sequence conforms to the physical dynamics of a microgravity environment in space, this invention not only embeds orbital dynamics constraints and multibody kinematics constraints during model training, but also performs dynamic residual correction on the interpretable state sequence output by the model during prediction. During residual correction, based on the interpretable state sequence, the angular velocity and angular acceleration of the target object at each future time step are determined; based on the angular velocity and angular acceleration of the target object at each future time step, the dynamic residual of the target object at each future time step is calculated; based on the dynamic residual, the angular velocity of the target object at each future time step is corrected to obtain the corrected angular velocity; and based on the corrected angular velocity, the predicted state sequence corresponding to each candidate action sequence is determined.
[0067] In the specific calculation of the dynamic residual, a visual image acquired by a camera device is obtained, the target object in the visual image is detected, and the target region image is extracted from the visual image based on the target detection result; the target region image is input into a target classification model for classification prediction to obtain the category information of the target object; based on the category information of the target object, a preset rotational inertia table is consulted to determine the rotational inertia corresponding to the target object; based on the angular velocity and rotational inertia of the target object at each future time step, and the rotational inertia corresponding to the target object, the dynamic residual of the target object at each future time step is calculated.
[0068] Specifically, this embodiment of the invention mainly utilizes attitude dynamics equations to calculate dynamic residuals. In the specific calculation, the moment of inertia of the target object needs to be known. To obtain the moment of inertia of the target object, this embodiment of the invention pre-constructs a target classification model and a preset moment of inertia table. The target classification model is used to classify target objects, and the preset moment of inertia table records the moment of inertia corresponding to different types of target objects. Based on the target object category information output by the target classification model, the preset moment of inertia table can be queried to determine the moment of inertia of the target object. Then, based on the moment of inertia, angular velocity, and angular acceleration of the target object, the dynamic residuals are calculated. The specific calculation formula for the dynamic residuals is as follows.
[0069] in, Represents dynamic residuals; Represents the moment of inertia; Represents the predicted angular velocity; Represents the predicted angular acceleration; The torque representing the target object.
[0070] The dynamic residual can be calculated using the above formula. Then, based on the calculated dynamic residual, the angular velocity of the target object is corrected. The specific formula is as follows.
[0071] in, This represents the corrected angular velocity; This represents the correction coefficient, which can be set according to actual business needs. The state sequence after angular velocity correction is the predicted state sequence.
[0072] The embodiments of the present invention can ensure that the predicted state sequence conforms to physical laws by calculating and correcting the dynamic residual, thereby ensuring the accuracy of the prediction results.
[0073] Furthermore, the predicted state sequence corresponding to each candidate action sequence is input into the reward prediction model to predict the reward, thereby obtaining the predicted reward for each candidate action sequence. A larger predicted reward indicates lower risk and higher reward for the candidate action sequence, meaning a better evaluation effect; conversely, a smaller predicted reward indicates higher risk and lower reward for the candidate action sequence, meaning a worse evaluation effect.
[0074] Furthermore, after determining the predicted rewards for each of the multiple candidate action sequences, the candidate action sequence with the highest predicted reward is selected as the target action sequence. It should be noted that if the predicted rewards for all candidate action sequences are below a safety threshold, a safety fallback strategy is triggered, such as controlling the onboard robotic arm to hover, move backward, or stop abruptly.
[0075] After determining the target candidate action sequence, to ensure that the onboard robotic arm does not collide with the target object and thus suffer damage when executing the target action sequence, this embodiment of the invention calculates control parameters for compliant grasping. Specifically, in calculating the control parameters, a target time step with a contact force greater than zero is determined based on the target predicted state sequence; based on the positional deviation, velocity deviation, and contact force between the onboard robotic arm and the target object within the target time step, and the positional deviation, velocity deviation, and contact force between the onboard robotic arm and the target object within the previous time step corresponding to the target time step, a Cartesian impedance equation is constructed; based on the Cartesian impedance equation, the stiffness and damping parameters corresponding to the target time step are solved; and based on the stiffness and damping parameters corresponding to the target time step, the control parameters when the end effector of the onboard robotic arm contacts the target object are determined.
[0076] Specifically, when there is a positional deviation between the spaceborne robotic arm and the target object, a contact force is generated. Therefore, a target time step with a contact force greater than zero is determined from the target predicted state sequence. Based on the position of the spaceborne robotic arm's end effector and the target object within the target time step, the positional deviation between the spaceborne robotic arm and the target object within the target time step is determined. Based on the position of the spaceborne robotic arm's end effector and the target object within the previous time step corresponding to the target time step, the positional deviation between the spaceborne robotic arm and the target object within the previous time step is determined. Simultaneously, based on the position and time step length of the spaceborne robotic arm's end effector in each future time step, the velocity of the spaceborne robotic arm's end effector in each future time step can be predicted. Then, based on the velocity of the spaceborne robotic arm's end effector and the velocity of the target object within the target time step, the velocity deviation between the spaceborne robotic arm and the target object within the target time step is determined. Based on the velocity of the spaceborne robotic arm's end effector and the velocity of the target object within the previous time step, the velocity deviation between the spaceborne robotic arm and the target object within the previous time step is determined. Next, based on the positional deviation, velocity deviation, and contact force between the spaceborne robotic arm and the target object in the target time step and the previous time step, the Cartesian impedance equation is constructed, and the control parameters are solved. The specific formulas are as follows.
[0077] in, , , These represent the positional deviation, velocity deviation, and contact force between the spaceborne robotic arm and the target object within the target time step, respectively. , , These represent the positional deviation, velocity deviation, and contact force between the spaceborne robotic arm and the target object in the previous time step, respectively. Represents the damping parameter; The stiffness parameter represents the control parameters to be solved, while the damping parameter and stiffness parameter represent the control parameters to be solved.
[0078] This invention, through the construction of the Cartesian impedance equation, solves for the stiffness and damping parameters of the spaceborne robotic arm when grasping a target object, thereby preventing the robotic arm from colliding with the target object and causing damage to the spaceborne robotic arm.
[0079] Furthermore, after determining the target action sequence and calculating the control parameters, the onboard robotic arm is controlled to execute according to the target action sequence, and to perform compliant grasping based on the control parameters when grasping the target object. During the execution of the onboard robotic arm, sensors are used to collect the actual state sequence of the onboard robotic arm and the target object, and the multi-dimensional execution deviation between the target predicted state sequence corresponding to the target action sequence and the actual state sequence is calculated, thereby employing corresponding correction strategies for correction. To address this correction process, the method includes: determining, based on the target predicted state sequence, the predicted pose of the target object, the predicted angles of each joint of the spaceborne robotic arm, and the predicted contact force between the target object and the spaceborne robotic arm in each future time step; determining, based on the actual state sequence, the actual pose of the target object, the actual angles of each joint of the spaceborne robotic arm, and the actual contact force between the target object and the spaceborne robotic arm in each time step during execution; calculating, based on the predicted pose, the predicted angles, and the predicted contact force, and the actual pose, the actual angles, and the actual contact force, the pose deviation of the target object, the angle deviation of each joint of the spaceborne robotic arm, and the contact force deviation between the target object and the spaceborne robotic arm, respectively; determining, based on the pose deviation, the angle deviation, and the contact force deviation, the multidimensional execution deviation; measuring, based on the multidimensional execution deviation, the state distance between the target predicted state sequence and the actual state sequence, and correcting, based on the state distance, using a corresponding correction strategy.
[0080] When making corrections based on the state distance using a corresponding correction strategy, if the state distance is less than the first preset distance, the target action sequence continues to be executed; if the state distance is greater than or equal to the first preset distance and less than the second preset distance, the target action sequence is redefined using the actual state at the corresponding time step as the current state representation vector; if the state distance is greater than or equal to the second preset distance, an emergency stop is triggered, and the spaceborne robotic arm is controlled to move to a preset safe posture.
[0081] The first and second preset distances can be set according to actual business needs.
[0082] Specifically, such as Figure 5As shown, the predicted pose of the target object, the predicted angles of each joint of the spaceborne robotic arm, and the predicted contact force between the target object and the spaceborne robotic arm are determined in the target predicted state sequence corresponding to the target action sequence. Simultaneously, the actual pose of the target object, the actual angles of each joint of the spaceborne robotic arm, and the actual contact force between the target object and the spaceborne robotic arm are determined in the actual state sequence. Then, the pose deviation of the target object is calculated. e 1. Angular deviations of various joints of the spaceborne robotic arm e 2. And the contact force deviation between the target object and the spaceborne robotic arm. e 3. This yields the multidimensional execution bias. The multidimensional execution biases are then concatenated into a bias vector. e The Euclidean norm 2 of the deviation vector is calculated to measure the state distance between the predicted state sequence and the actual state sequence. The specific formula is as follows.
[0083] Furthermore, if the calculated state distance is less than the first preset distance, it indicates that the predicted state sequence is basically consistent with the actual state sequence, and the target action sequence can continue to be executed without correction. If the calculated state distance is greater than or equal to the first preset distance and less than the second preset distance, it indicates that there is a tolerable deviation between the predicted state sequence and the actual state sequence. At this time, the current actual state sequence is used as a new starting point to re-determine and select the optimal action sequence (target action sequence). If the calculated state distance is greater than or equal to the second preset distance, it indicates that the predicted state sequence deviates significantly from the actual state sequence, which may be due to drastic changes in the environment or serious inaccuracies in the state prediction model. At this time, an emergency stop needs to be triggered immediately to control the spaceborne robotic arm to move to the preset safe posture.
[0084] Furthermore, upon triggering an emergency stop, the onboard robotic arm immediately ceases its current movement, retracting to a retracted or hovering posture along a preset safe trajectory. A complete snapshot of the state at the moment of the emergency stop is recorded and stored in on-orbit memory. Simultaneously, the emergency stop event and deviation data are reported to the ground station via an available communication link. In a safe state, the recorded deviation data is used to fine-tune the state prediction model using LoRA, adapting the model to the new environment. After fine-tuning, mission execution can only resume after safety verification.
[0085] It should be noted that the values of the first and second preset distances are not fixed, but are dynamically adjusted according to the task stage and the prediction confidence of the state prediction model. In the early stage of the task or when the prediction confidence of the state prediction model is low, the first and second preset distances are set to lower values to make the correction strategy more sensitive. In the later stage of the task and when the state prediction model has adapted to the current environment, the first and second preset distances can be appropriately widened to reduce unnecessary replanning frequency.
[0086] The spaceborne robotic arm in this invention no longer passively waits for ground commands or mechanically executes preset rules. Instead, it autonomously deduces multiple candidate solutions, assesses risks and benefits, and selects the optimal action sequence. This fundamentally improves on-orbit automatic control capabilities and the ability to respond to emergencies. Even in scenarios of communication interruption or high latency between space and ground, this invention can still ensure the robotic arm completes its on-orbit tasks safely and efficiently.
[0087] This invention provides a method for training a state prediction model for a spaceborne robotic arm. By introducing orbital dynamics constraint loss and multibody kinematics constraint loss into the loss function, a physical deduction layer is constructed between model decision-making and spaceborne robotic arm execution. This allows decisions to undergo dynamic consistency verification before execution. Simultaneously, this invention also performs dynamic residual correction on the predicted state sequence output by the model during the prediction phase. Through this dual physical constraint mechanism of the training and prediction phases, this invention ensures that the model's decision results conform to the physical and dynamic laws of the microgravity environment in space, thereby improving the success rate of spaceborne robotic arm mission execution and preventing collisions that could damage the robotic arm. Furthermore, the state prediction model constructed in this invention can not only deduce the execution consequences of action sequences but also predict the execution rewards. Based on this predicted reward, the optimal candidate action can be quickly identified, thereby improving the control accuracy and efficiency of the spaceborne robotic arm.
[0088] Furthermore, as Figure 1 and Figure 3 The specific implementation of the method shown in this embodiment provides a state prediction model training device for a spaceborne robotic arm, such as... Figure 6 As shown, the device includes: a data acquisition unit 101, a prediction unit 102, an evaluation unit 103, a construction unit 104, and a training unit 105.
[0089] The acquisition unit 101 can be used to acquire state transition samples when the spaceborne robotic arm grasps the target object in different scenarios. The state transition samples include sample state, sample action sequence and actual state sequence.
[0090] The prediction unit 102 can be used to construct an initial state prediction model, and input the sample state and the sample action sequence into the initial state prediction model to perform state prediction, so as to obtain the training prediction state sequence and training prediction reward corresponding to the sample action sequence.
[0091] The evaluation unit 103 can be used to perform multi-dimensional reward evaluation on the training prediction state sequence to obtain the actual reward corresponding to the training prediction state sequence.
[0092] The construction unit 104 can be used to construct a total loss function based on the training predicted state sequence and the actual state sequence, as well as the training predicted reward and the actual reward. The total loss function embeds a dynamic constraint loss, which includes orbital dynamics constraint loss and multibody kinematics constraint loss.
[0093] The training unit 105 can be used to iteratively train the initial state prediction model according to the total loss function to construct a preset state prediction model, wherein the predicted state sequence output by the preset state prediction model in the prediction stage is obtained after dynamic residual correction.
[0094] In some embodiments, the evaluation unit 103 includes a determination module and a first calculation module.
[0095] The determining module can be used to determine, based on the training prediction state sequence, the predicted angle and predicted torque of each joint of the spaceborne robotic arm, the predicted position of the end effector of the spaceborne robotic arm, the predicted position of the target object, and the number of task completion time steps within each training time step.
[0096] The calculation module can be used to calculate the task completion reward based on the predicted position of the end effector of the spaceborne robotic arm and the predicted position of the target object.
[0097] The first calculation module can also be used to calculate the collision probability based on the predicted angles of each joint of the spaceborne robotic arm, and determine the collision safety reward based on the collision probability.
[0098] The first calculation module can also be used to calculate energy consumption reward based on the predicted torque of each joint of the spaceborne robotic arm.
[0099] The first calculation module can also be used to calculate a time efficiency reward based on the number of time steps and the time step length of the task completion.
[0100] The first calculation module can also be used to calculate the actual reward based on the task completion reward, the collision safety reward, the energy consumption reward, and the time efficiency reward.
[0101] In some embodiments, the first computing module includes: an acquisition submodule, a construction submodule, and a computing submodule.
[0102] The acquisition submodule can be used to acquire the fixed structural parameters of the spaceborne robotic arm and the geometric dimensions of each link, as well as the three-dimensional point cloud data of the target object.
[0103] The construction submodule can be used to construct a oriented bounding box for each link based on the fixed structural parameters, the geometric dimensions of each link, and the predicted angles of each joint of the spaceborne robotic arm.
[0104] The construction submodule can also be used to construct a point cloud model of the target object based on the three-dimensional point cloud data.
[0105] The calculation submodule can be used to calculate the minimum distance between each link and the target object based on the directed bounding box of each link and the point cloud model of the target object.
[0106] The calculation submodule can be used to calculate the collision probability based on the minimum distance.
[0107] In some embodiments, the construction submodule may be specifically used to calculate the adjacent link transformation matrix of the spaceborne manipulator based on the fixed structural parameters and the predicted angles of each joint of the spaceborne manipulator using a preset forward kinematics function; to accumulate and calculate the pose of each link in the world base coordinate system based on the adjacent link transformation matrix; and to construct a directed bounding box for each link based on the pose of each link in the world base coordinate system and the geometric dimensions of each link.
[0108] In some embodiments, the building unit 104 includes: a building module and a second computing module.
[0109] The construction module can be used to construct a reconstruction loss based on the training predicted state sequence and the actual state sequence.
[0110] The second calculation module can be used to calculate the divergence constraint loss based on the hidden state sequence corresponding to the training predicted state sequence.
[0111] The second calculation module can also be used to calculate the dynamic constraint loss based on the training prediction state sequence and the sample state.
[0112] The second calculation module can also be used to calculate the reward loss based on the training predicted reward and the actual reward.
[0113] The construction module can also be used to construct a total loss function based on the reconstruction loss, the divergence constraint loss, the dynamic constraint loss, and the reward loss.
[0114] In some embodiments, the second calculation module may be specifically configured to: determine the predicted pose of the end effector of the spaceborne robotic arm, and the predicted position and velocity of the target object within each training time step based on the training prediction state sequence; calculate the updated position and velocity of the target object using orbital dynamics equations based on the sample states; calculate the orbital dynamics constraint loss based on the predicted position and velocity of the target object, and the updated position and velocity of the target object; calculate the updated position and updated attitude of the end effector of the spaceborne robotic arm using position positive kinematics functions and attitude positive kinematics functions respectively based on the training prediction state sequence; calculate the multibody kinematics constraint loss based on the predicted pose of the end effector of the spaceborne robotic arm, and the updated position and updated attitude of the end effector of the spaceborne robotic arm; and calculate the dynamics constraint loss based on the orbital dynamics constraint loss and the multibody kinematics constraint loss.
[0115] In some embodiments, the prediction unit 102 may further be used to obtain the current state representation vector corresponding to both the spaceborne robotic arm and the target object, as well as multiple candidate action sequences of the spaceborne robotic arm; input the current state representation vector and each candidate action sequence from the multiple candidate action sequences into the preset state prediction model for state prediction, to obtain the predicted state sequence and predicted reward corresponding to each candidate action sequence; determine the target action sequence from the multiple candidate action sequences based on the predicted reward; solve the control parameters when the end effector of the spaceborne robotic arm contacts the target object according to the target predicted state sequence corresponding to the target action sequence; control the spaceborne robotic arm to grasp the target object based on the control parameters and the target action sequence, and collect the actual state sequence corresponding to both the spaceborne robotic arm and the target object; calculate the multidimensional execution deviation based on the target predicted state sequence and the actual state sequence, and correct it using a corresponding correction strategy based on the multidimensional execution deviation.
[0116] It should be noted that other corresponding descriptions of the functional units involved in the state prediction model training device for a spaceborne robotic arm provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 3 The corresponding description in [the document] will not be repeated here.
[0117] Based on the above, Figure 1 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 1 and Figure 3 The method for training the state prediction model of the spaceborne robotic arm is shown.
[0118] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0119] Based on the above, Figure 1 and Figure 3 The method shown, and Figure 6 To achieve the above objectives, the present application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, as shown in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 3 The method for training the state prediction model of the spaceborne robotic arm is shown.
[0120] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0121] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0122] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0123] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0124] This invention, by introducing orbital dynamics constraint loss and multibody kinematics constraint loss into the loss function, constructs a physical deduction layer between model decision-making and onboard robotic arm execution. This allows decisions to undergo dynamic consistency verification before execution. Simultaneously, this invention also performs dynamic residual correction on the predicted state sequence output by the model during the prediction phase. Through this dual physical constraint mechanism in the training and prediction phases, this invention ensures that the model's decision results conform to the physical dynamics laws of the microgravity environment in space, thereby improving the success rate of onboard robotic arm mission execution and preventing collisions that could damage the onboard robotic arm. Furthermore, the state prediction model constructed in this invention can not only deduce the execution consequences of action sequences but also predict the execution rewards. Based on this predicted reward, the optimal candidate action can be quickly identified, thereby improving the control accuracy and efficiency of the onboard robotic arm.
[0125] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0126] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A method for training a state prediction model for a spaceborne robotic arm, characterized in that, include: Collect state transition samples when the spaceborne robotic arm grasps target objects in different scenarios. The state transition samples include sample states, sample action sequences, and actual state sequences. An initial state prediction model is constructed, and the sample state and the sample action sequence are input into the initial state prediction model to perform state prediction, thereby obtaining the training prediction state sequence and training prediction reward corresponding to the sample action sequence. A multi-dimensional reward evaluation is performed on the training prediction state sequence to obtain the actual reward corresponding to the training prediction state sequence. Based on the training predicted state sequence and the actual state sequence, as well as the training predicted reward and the actual reward, a total loss function is constructed, wherein the total loss function embeds a dynamic constraint loss, which includes orbital dynamics constraint loss and multibody kinematics constraint loss. Based on the total loss function, the initial state prediction model is iteratively trained to construct a preset state prediction model. In the prediction phase, the predicted state sequence output by the preset state prediction model is obtained after dynamic residual correction.
2. The method according to claim 1, characterized in that, The step of performing multi-dimensional reward evaluation on the training predicted state sequence to obtain the actual reward corresponding to the training predicted state sequence includes: Based on the training prediction state sequence, the predicted angle and predicted torque of each joint of the spaceborne robotic arm, the predicted position of the end effector of the spaceborne robotic arm, the predicted position of the target object, and the number of task completion time steps are determined in each training time step. The task completion reward is calculated based on the predicted position of the end effector of the spaceborne robotic arm and the predicted position of the target object. Based on the predicted angles of each joint of the spaceborne robotic arm, the collision probability is calculated, and based on the collision probability, a collision safety reward is determined. The energy consumption bonus is calculated based on the predicted torque of each joint of the spaceborne robotic arm. Calculate the time efficiency bonus based on the number of time steps and time step length for completing the task. The actual reward is calculated based on the task completion reward, the collision safety reward, the energy consumption reward, and the time efficiency reward.
3. The method according to claim 2, characterized in that, The step of calculating the collision probability based on the predicted angles of each joint of the spaceborne robotic arm includes: Obtain the fixed structural parameters of the spaceborne robotic arm and the geometric dimensions of each link, as well as the three-dimensional point cloud data of the target object; Based on the fixed structural parameters, the geometric dimensions of each link, and the predicted angles of each joint of the spaceborne robotic arm, a directional bounding box is constructed for each link. Based on the three-dimensional point cloud data, construct a point cloud model of the target object; Based on the directed bounding box of each link and the point cloud model of the target object, calculate the minimum distance between each link and the target object; The collision probability is calculated based on the minimum distance.
4. The method according to claim 3, characterized in that, The step of constructing a oriented bounding box for each link based on the fixed structural parameters, the geometric dimensions of each link, and the predicted angles of each joint of the spaceborne robotic arm includes: Based on the fixed structural parameters and the predicted angles of each joint of the spaceborne robotic arm, the adjacent link transformation matrix of the spaceborne robotic arm is calculated using a preset positive kinematic function. Based on the adjacent link transformation matrix, the pose of each link in the world base coordinate system is calculated cumulatively. Based on the pose of each link in the world base coordinate system and the geometric dimensions of each link, a directed bounding box is constructed for each link.
5. The method according to claim 1, characterized in that, The construction of the total loss function based on the trained predicted state sequence and the actual state sequence, as well as the trained predicted reward and the actual reward, includes: Based on the trained predicted state sequence and the actual state sequence, a reconstruction loss is constructed; Based on the hidden state sequence corresponding to the training predicted state sequence, calculate the divergence constraint loss; The dynamic constraint loss is calculated based on the training predicted state sequence and the sample state; Calculate the reward loss based on the training predicted reward and the actual reward; Based on the reconstruction loss, the divergence constraint loss, the dynamic constraint loss, and the reward loss, a total loss function is constructed.
6. The method according to claim 5, characterized in that, The calculation of the dynamic constraint loss based on the training predicted state sequence and the sample state includes: Based on the training prediction state sequence, the predicted pose of the end effector of the spaceborne robotic arm, as well as the predicted position and velocity of the target object, are determined within each training time step. Based on the sample state, the updated position and update velocity of the target object are calculated using the orbital dynamics equations. The orbital dynamics constraint loss is calculated based on the predicted position and velocity of the target object, as well as the updated position and velocity of the target object. Based on the training predicted state sequence, the updated position and updated attitude of the end effector of the spaceborne robotic arm are calculated using the position positive kinematics function and the attitude positive kinematics function, respectively. The multibody kinematics constraint loss is calculated based on the predicted pose of the end effector of the spaceborne robotic arm, as well as the updated position and updated pose of the end effector. The dynamic constraint loss is calculated based on the orbital dynamics constraint loss and the multibody kinematics constraint loss.
7. The method according to any one of claims 1-6, characterized in that, After iteratively training the initial state prediction model according to the total loss function to construct the preset state prediction model, the method further includes: Obtain the current state representation vector that corresponds to both the spaceborne robotic arm and the target object, as well as multiple candidate action sequences of the spaceborne robotic arm; The current state representation vector and each candidate action sequence from the plurality of candidate action sequences are input into the preset state prediction model to perform state prediction, thereby obtaining the predicted state sequence and predicted reward corresponding to each candidate action sequence. Based on the predicted reward, a target action sequence is determined from the plurality of candidate action sequences; Based on the target predicted state sequence corresponding to the target action sequence, the control parameters when the end effector of the spaceborne robotic arm comes into contact with the target object are solved. Based on the control parameters and the target action sequence, the spaceborne robotic arm is controlled to grasp the target object, and the actual state sequence corresponding to both the spaceborne robotic arm and the target object is collected. Based on the target predicted state sequence and the actual state sequence, a multidimensional execution deviation is calculated, and based on the multidimensional execution deviation, a corresponding correction strategy is adopted for correction.
8. A state prediction model training device for a spaceborne robotic arm, characterized in that, include: The acquisition unit is used to acquire state transition samples when the spaceborne robotic arm grasps the target object in different scenarios. The state transition samples include sample states, sample action sequences, and actual state sequences. The prediction unit is used to construct an initial state prediction model and input the sample state and the sample action sequence into the initial state prediction model to perform state prediction, thereby obtaining the training prediction state sequence and training prediction reward corresponding to the sample action sequence. An evaluation unit is used to perform multi-dimensional reward evaluation on the training prediction state sequence to obtain the actual reward corresponding to the training prediction state sequence. A construction unit is used to construct a total loss function based on the training predicted state sequence and the actual state sequence, as well as the training predicted reward and the actual reward, wherein the total loss function embeds a dynamic constraint loss, which includes orbital dynamics constraint loss and multibody kinematics constraint loss; The training unit is used to iteratively train the initial state prediction model according to the total loss function to construct a preset state prediction model, wherein the predicted state sequence output by the preset state prediction model in the prediction stage is obtained after dynamic residual correction.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
10. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.