Space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning
By employing a reinforcement learning-based method for space robot extravehicular communication relay and dynamic tracking, and utilizing a space crawling robotic arm as a communication relay, combined with PPO-PID and LSTM-PPO algorithms, the signal transmission problem in spacecraft extravehicular communication obstruction areas was solved, achieving full coverage and stable communication, and improving communication efficiency and intelligence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING RES INST OF PRECISE MECHATRONICS CONTROLS
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies cannot effectively solve the communication problems of space-based maintenance robots in areas with external communication obstructions without disrupting the aerodynamic shape of the spacecraft, especially the signal transmission problem in areas obstructed by the hatch or metal cabin.
A reinforcement learning-based approach is adopted, using a space crawling robotic arm as a communication relay to divide the area into multiple sub-communication regions. PPO-PID and LSTM-PPO algorithms are used for dynamic target tracking and obstacle avoidance control. A multi-path wireless communication link is constructed between the in-cabin integrated controller, the space crawling robotic arm, and the space on-orbit maintenance robot to achieve signal strength-guided robotic arm pose optimization.
It achieves full coverage of the extravehicular operating area and stable communication in the complex extravehicular environment, avoiding communication blind spots, improving communication efficiency and the utilization rate of the robotic arm, ensuring the continuity and stability of the communication link, and meeting the actual space mission requirements.
Smart Images

Figure CN121900404A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of space robot technology, and in particular relates to a method for space robot extravehicular communication relay and dynamic tracking based on reinforcement learning. Background Technology
[0002] With the continuous development of space technology, on-orbit maintenance technology has gradually become an important means to ensure the long-term stable operation of spacecraft. Utilizing on-orbit maintenance robots to perform routine inspections and maintenance on spacecraft that have been stationed in orbit for extended periods has become a real need. However, due to the aerodynamic design limitations of spacecraft traveling between Earth and space, it is usually impossible to install communication antennas on the spacecraft surface. When the on-orbit maintenance robot is located behind a spacecraft hatch opening or in an area obstructed by metal hull, communication signals are severely blocked, leading to decreased communication quality or even communication interruption. Therefore, how to effectively solve the communication problem of on-orbit maintenance robots in areas with obstructed communication outside the spacecraft has become a key issue that urgently needs to be addressed in the field of on-orbit maintenance technology.
[0003] Existing technologies have proposed several methods to address space communication issues: Patent CN107682073A proposes installing multiple communication antennas at equal intervals on the spacecraft bulkhead to achieve signal coverage within a 360° range. However, this method disrupts the overall aerodynamic shape of the spacecraft and cannot meet the requirements of actual missions. Patent CN117527036A proposes using a two-degree-of-freedom antenna gimbal to achieve adaptive pointing control of the satellite remote relay telemetry and control docking antenna, but two-degree-of-freedom antennas are easily obstructed and limited by space, making it difficult to solve the communication obstruction problem for maintenance robots on the back of the hatch. Patents CN119667742A and CN119786970A propose antenna positioning methods based on 5G networks and BeiDou positioning, respectively, which can achieve precise adjustment of the antenna direction, but still do not solve the signal transmission problem in areas with communication obstruction. The robotic arm end-effector trajectory tracking algorithm based on zero-space obstacle avoidance proposed in patent CN113146610A, and the robotic arm target point online tracking method based on pose tracking system proposed in patent CN113561183B, mainly focus on the motion trajectory planning and obstacle avoidance control of the robotic arm. They usually use the target pose as a reference condition for adjusting the pose of the robotic arm, and do not consider communication quality as a key factor in guiding the movement of the robotic arm.
[0004] In summary, the existing technologies have the following shortcomings: (1) Installing multiple communication antennas on the surface of spacecraft will damage the aerodynamic shape of spacecraft and cannot meet the actual mission requirements; (2) Using two-degree-of-freedom antennas or traditional antenna positioning methods is easily subject to spatial obstruction and installation location limitations, making it difficult to achieve signal coverage in communication obstruction areas; (3) Existing robotic arm control methods do not consider communication quality factors and cannot actively adjust the robotic arm posture to improve communication performance.
[0005] It is evident that existing technologies cannot provide stable and high-quality communication guarantees for space-based maintenance robots in communication-blocked areas. There is an urgent need to develop a new method that can effectively solve the communication problems of space-based maintenance robots in communication-blocked areas outside the spacecraft without compromising the aerodynamic shape of the spacecraft. Summary of the Invention
[0006] The technical problem solved by this invention is to overcome the shortcomings of the prior art and provide a reinforcement learning-based method for space robot extravehicular communication relay and dynamic tracking, which aims to solve the problem of stable communication for space on-orbit maintenance robots in areas where communication is blocked outside the spacecraft (such as the back of the hatch or the area blocked by the metal cabin).
[0007] To address the aforementioned technical problems, this invention discloses a reinforcement learning-based method for extravehicular communication relay and dynamic tracking of space robots, comprising: Based on the operating range of the space-based on-orbit maintenance robot, the communication area of the space-based on-orbit maintenance robot is divided into multiple sub-communication areas; Based on the division of multiple sub-communication regions, the space crawling robot arm is controlled to crawl between the repeating locking mechanisms of the in-cabin robot arm and is fixed to the corresponding repeating locking mechanism of the in-cabin robot arm. The space crawling robot arm serves as a communication relay between the space on-orbit maintenance robot and the integrated control platform of the in-cabin robot. After the space crawling robotic arm crawls and is fixed in the corresponding cabin robotic arm repeated locking mechanism, the space crawling robotic arm dynamic target tracking and obstacle avoidance control based on PPO-PID is adopted to complete the accurate tracking of dynamic targets and obstacle avoidance.
[0008] The aforementioned reinforcement learning-based method for space robot extravehicular communication relay and dynamic tracking further includes: constructing an extravehicular communication system for a space-on-orbit maintenance robot, and realizing extravehicular communication for the space-on-orbit maintenance robot based on this system; wherein, the extravehicular communication system for the space-on-orbit maintenance robot includes: a space-on-orbit maintenance robot, a space crawling robotic arm, an in-cabin robot integrated control platform, and an in-cabin robotic arm repetitive locking mechanism; a wireless communication device A is installed on the space-on-orbit maintenance robot; the wireless communication device A includes: a wireless communication module A and an antenna A; the wireless communication module A is installed on the space-on-orbit maintenance robot. Inside the robot's torso, two antennas A are respectively installed at the center of the front and rear outer shells of the space-based maintenance robot's torso; the two antennas A are connected to the wireless communication module A via feed lines; the space crawling robotic arm is equipped with wireless communication devices B and C; wireless communication device B includes: wireless communication module B and antenna B; wireless communication module B and antenna B are installed at the end of the first joint of the space crawling robotic arm, and antenna B is connected to wireless communication module B via feed lines; wireless communication device C includes: wireless communication module C and antenna C; wireless communication module C and antenna C are installed at the end of the seventh joint of the space crawling robotic arm, and antenna C is connected to wireless communication module B via feed lines. The wire is connected to the wireless communication module C; the repetitive locking mechanism of the in-cabin robotic arm serves as the mechanical and electrical interface between the space crawling robotic arm and the spacecraft cabin; there are four repetitive locking mechanisms for the in-cabin robotic arm: in-cabin robotic arm repetitive locking mechanism A, in-cabin robotic arm repetitive locking mechanism B, in-cabin robotic arm repetitive locking mechanism C, and in-cabin robotic arm repetitive locking mechanism D, which are respectively arranged in the four corners near the hatch inside the spacecraft cabin; the in-cabin robot integrated control platform includes: an integrated controller and wireless communication equipment D; the wireless communication equipment D includes: a wireless communication module D and an antenna D; the integrated controller is connected via electronic cables. The antenna D is connected to the wireless communication module D via a feeder; the integrated controller is connected to the repeating locking mechanism of the four in-cabin robotic arms via electronic circuitry; the space crawling robotic arm is electrically connected to the repeating locking mechanism of the in-cabin robotic arms via connectors; the integrated controller communicates wirelessly directly with the wireless communication device A on the space-on-orbit maintenance robot via the wireless communication device D, or the integrated controller communicates wirelessly with the wireless communication device A on the space-on-orbit maintenance robot via the wireless communication device D, using the space crawling robotic arm as a communication relay, through the wireless communication devices B and C on the space crawling robotic arm.
[0009] In the aforementioned reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, the communication area of the space-on-orbit maintenance robot is divided into multiple sub-communication areas based on its operational range. This includes: determining the operational range of the space-on-orbit maintenance robot, which is an ellipsoid centered on the center of the spacecraft cabin; and, based on the ellipsoid and considering the coverage of various wireless communication devices and the signal obstruction caused by the metal cabin, dividing the communication area of the space-on-orbit maintenance robot into five sub-communication areas: Sub-communication Area A, Sub-communication Area B, Sub-communication Area C, Sub-communication Area D, Sub-communication Area E, Sub-communication Area F, Sub-communication Area C, Sub-communication Area D, Sub-communication Area E, Sub-communication Area F, Sub-communication Area D, Sub-communication Area E, Sub-communication Area F, Sub-communication Area F, Sub-communication Area E, Sub-communication Area F, Sub-communication Area F, Sub-communication Area F, Sub-communication Area D, Sub-communication Area E, Sub-communication Area F ... Sub-communication area D and sub-communication area E; wherein, the area above the spacecraft cabin that can be unobstructed by wireless communication device D is sub-communication area E; the remaining area of the ellipsoid is divided into four areas by the plane formed by the X-axis and Z-axis of the spacecraft cabin coordinate system and the plane formed by the Y-axis and Z-axis, respectively. The area containing the repeating locking mechanism A of the cabin robotic arm is sub-communication area A, the area containing the repeating locking mechanism B of the cabin robotic arm is sub-communication area B, the area containing the repeating locking mechanism C of the cabin robotic arm is sub-communication area C, and the area containing the repeating locking mechanism D of the cabin robotic arm is sub-communication area D.
[0010] In the aforementioned reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, based on multiple divided sub-communication regions, the space crawling manipulator is controlled to crawl between the repeating locking mechanisms of the in-cabin manipulator and is fixed to the corresponding repeating locking mechanism of the in-cabin manipulator. The space crawling manipulator serves as a communication relay between the space on-orbit maintenance robot and the integrated control platform of the in-cabin robot, including: When the space-based maintenance robot is performing inspection and maintenance operations in the sub-communication area E, the integrated controller communicates directly with the wireless communication device A on the space-based maintenance robot through the wireless communication device D. The space crawling robotic arm performs its handling operations normally without being affected. When the space-based maintenance robot performs inspection and maintenance operations in sub-communication area A, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism A, and electrically connects with the cabin robotic arm repeating locking mechanism A through a connector; the integrated controller communicates wirelessly with the wireless communication device A on the space-based maintenance robot through wireless communication device D, using the space crawling robotic arm as a communication relay, and through wireless communication devices B and C on the space crawling robotic arm; When the space-based maintenance robot performs inspection and maintenance operations in sub-communication area B, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding cabin robotic arm repeating locking mechanism B, and electrically connects with the cabin robotic arm repeating locking mechanism A through a connector; the integrated controller communicates wirelessly with the space crawling robotic arm as a communication relay via wireless communication device D, wireless communication device B and wireless communication device C on the space crawling robotic arm, and wireless communication device A on the space-based maintenance robot. When the space-based maintenance robot performs inspection and maintenance operations in the sub-communication area C, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism C, and electrically connects with the cabin robotic arm repeating locking mechanism A through a connector; the integrated controller communicates wirelessly with the space crawling robotic arm as a communication relay via wireless communication device D, wireless communication device B and wireless communication device C on the space crawling robotic arm, and wireless communication device A on the space-based maintenance robot. When the space-based maintenance robot performs inspection and maintenance operations within the sub-communication area D, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism D, and electrically connects with the cabin robotic arm repeating locking mechanism A via a connector; the integrated controller communicates wirelessly with the space-based maintenance robot via wireless communication device D, using the space crawling robotic arm as a communication relay, and through wireless communication devices B and C on the space crawling robotic arm.
[0011] In the aforementioned reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, a PPO-PID-based space crawling robotic arm is used for dynamic target tracking and obstacle avoidance control to achieve accurate tracking of dynamic targets and obstacle avoidance, including: Initialize the parameters of the space crawling robotic arm, including PID parameters, LSTM parameters, and PPO parameters; Obtain the pose of the space-based maintenance robot and the space-crawling robotic arm in the spacecraft coordinate system; By using PID control to adjust the pose of the space crawling robot arm, the working plane of the space crawling robot arm can be quickly brought close to the space on-orbit maintenance robot, thus achieving coarse positioning of the space crawling robot arm. Based on the occlusion conditions and signal strength, the LSTM-PPO algorithm is used to perform adaptive spatial crawling robot arm pointing precision control, so as to realize the stable tracking and obstacle avoidance of the space crawling robot arm on the space on-orbit maintenance robot.
[0012] In the above reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, coarse localization of the space crawling robotic arm is achieved through the following three constraints: Condition 1: The end effector of the space crawling robotic arm points towards the center of the space-based on-orbit maintenance robot; Condition 2: Maintain a safe distance between the space crawling robotic arm and the space on-orbit maintenance robot; Condition 3: Point the end effector of the space crawling robotic arm as far as possible toward the antenna A of the space-based maintenance robot.
[0013] In the aforementioned reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, the LSTM-PPO algorithm is employed based on occlusion conditions and signal strength to perform adaptive precise pointing control of the space crawling manipulator. This enables the space crawling manipulator to stably track and avoid obstacles for the space-based on-orbit maintenance robot, including: S1, initialize the parameters of the Actor network, Critic network, and LSTM network; S2, acquire timing state data and complete state space design; S3. Using an LSTM network, the time series state data is filtered to extract key time series state data. S4. Calculate the dense reward and sparse reward based on the key time series state data to obtain the fused reward; S5 uses the PPO planning algorithm to update the strategy network parameters, enabling the space crawling robotic arm to stably track and avoid obstacles on the space-based on-orbit maintenance robot. S6: When the space-based maintenance robot is making large displacement movements, Kalman filtering is used to predict the future pose of the space-based maintenance robot and adjust the pose of the space crawling robot arm in advance to achieve predictive tracking. S7. Repeat steps S2 to S6 until the set number of training iterations is reached or the convergence condition is met.
[0014] In the above reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, the Actor network is used to select the action of the space crawling manipulator based on the current state; the Critic network is used to evaluate the value of the selected action; and the LSTM network is used to process the time series data of the state of the space crawling manipulator and the state of the space on-orbit maintenance robot.
[0015] In the aforementioned reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, vector... To represent the state space:
[0016] in, This indicates the current pose of the space crawling robotic arm. This indicates the target pose of the space crawling robotic arm. This indicates the angles of each joint of the space crawling robotic arm. This indicates the length of each link in the space crawling robotic arm. This represents the angular velocity of each joint of the space crawling robotic arm. This indicates the step length of the space crawling robotic arm. Indicates the signal strength of the receiver in a wireless communication device. This indicates the current pose of the space-based on-orbit maintenance robot. This indicates the future pose of the space-based maintenance robot. This indicates the signal strength of receiver A in wireless communication device A.
[0017] In the above reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots, the rewards are integrated, including dense rewards and sparse rewards. Dense rewards include position rewards and orientation rewards; sparse rewards include signal strength, obstacle avoidance, acceleration guidance, and adaptive step size rewards, which guide the space crawling robot arm to move towards areas with high communication quality, safe obstacle avoidance, and efficient movement.
[0018] The present invention has the following advantages: (1) This invention discloses a space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning. The space crawling manipulator is used as a dynamic communication relay to construct a multi-path wireless communication link of "in-cabin integrated controller - space crawling manipulator - space on-orbit maintenance robot". Without changing the aerodynamic shape of the spacecraft, the space crawling manipulator pose optimization based on crawling path planning of sub-communication area and signal strength guidance effectively solves the signal transmission problem of space on-orbit maintenance robot in communication blockage areas such as the back of the spacecraft hatch. It completely eliminates the communication blind spot caused by the metal cabin blockage and realizes full coverage of the extravehicular operation area and stable communication in the complex extravehicular environment. It is significantly better than the existing limited solutions such as "multi-path fixed antennas destroy the shape and two-degree-of-freedom gimbal is easy to block".
[0019] (2) This invention discloses a space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning. According to the working range of the space on-orbit maintenance robot, the communication area of the space on-orbit maintenance robot is divided into multiple sub-communication areas. Corresponding communication strategies and space crawling manipulator deployment schemes are formulated for different sub-communication areas, which effectively improves communication efficiency and the utilization rate of space crawling manipulators and avoids resource waste.
[0020] (3) This invention discloses a space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning. It proposes a space crawling manipulator control scheme based on PID coarse localization and LSTM-PPO reinforcement learning algorithm, which enables the space crawling manipulator to track the space on-orbit maintenance robot in real time and autonomously avoid obstacles, ensuring that the link is always in the best communication state.
[0021] (4) This invention discloses a space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning. Compared with the existing technology of installing multiple communication antennas on the surface of the spacecraft, this invention does not require any modification to the spacecraft structure, avoids the problem of damaging the aerodynamic shape of the spacecraft, and is more in line with the needs of actual space missions.
[0022] (5) This invention discloses a space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning. When the space on-orbit maintenance robot performs large displacement and rapid movement, the future pose of the space on-orbit maintenance robot is predicted by Kalman filtering, and the pose of the space crawling manipulator is adjusted in advance to achieve predictive tracking and ensure the continuity and stability of the communication link.
[0023] (6) This invention discloses a space robot extravehicular communication relay and dynamic tracking method based on reinforcement learning. It designs a multi-level reward mechanism to effectively guide the space crawling robotic arm to make optimal decisions in complex environments, thereby improving the level of intelligence and task execution efficiency. Attached Figure Description
[0024] Figure 1 This is a flowchart of a reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots in an embodiment of the present invention; Figure 2 This is a schematic diagram of the layout of an external communication system for a space-based on-orbit maintenance robot according to an embodiment of the present invention; Figure 3 This is a block diagram of the external communication system for a space-based on-orbit maintenance robot according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a sub-communication area division in an embodiment of the present invention; Figure 5 This is a flowchart of a dynamic target tracking and obstacle avoidance control method for a space crawling robotic arm based on PPO-PID in an embodiment of the present invention; Figure 6 This is a schematic diagram of coarse pose localization of a spatial crawling robotic arm based on PID in an embodiment of the present invention; Figure 7 This is a schematic diagram of the dynamic target tracking and obstacle avoidance control of a space crawling robotic arm based on the LSTM-PPO algorithm in an embodiment of the present invention. Figure 8 This is a schematic diagram illustrating the principle of selecting time-series state data for a spatial crawling robotic arm based on the LSTM algorithm in an embodiment of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments disclosed in the present invention will be described in further detail below with reference to the accompanying drawings.
[0026] In existing technologies, the installation of communication antennas is limited by the aerodynamic shape of the spacecraft, making it impossible to install multiple communication antennas at equal intervals on the spacecraft bulkhead. This leads to signal attenuation or interruption, affecting the operation of the space-based on-orbit maintenance robot. Furthermore, existing robotic arm control methods do not consider using communication quality as a reference condition to adjust the robotic arm's posture to improve communication performance. Based on this, this invention discloses a reinforcement learning-based method for space robot extravehicular communication relay and dynamic tracking. On one hand, using a space crawling robotic arm as a dynamic communication relay, a multi-path wireless communication link is constructed: "In-cabin integrated controller—space crawling robotic arm—space-based on-orbit maintenance robot." Without altering the spacecraft's aerodynamic shape, through crawling path planning based on sub-communication regions and signal strength-guided posture optimization of the space crawling robotic arm, the signal transmission problem of the space-based on-orbit maintenance robot in communication-obstructed areas such as the back of the spacecraft hatch is effectively solved. This completely eliminates communication blind spots caused by metal bulkhead obstruction, achieving full coverage of the extravehicular work area and stable communication across the entire area in complex extravehicular environments. This is significantly superior to existing solutions that suffer from limitations such as "multiple fixed antennas disrupting the shape" and "two-degree-of-freedom gimbals being easily obstructed." On the other hand, a dynamic target tracking and obstacle avoidance control method for a space crawling manipulator based on PID+PPO is proposed. PID is used to complete millisecond-level coarse positioning of the manipulator's pose, LSTM is used to extract key time series state data, PPO is used to output the optimal action strategy for the space crawling manipulator, and Kalman filtering is used to achieve pose prediction and tracking. The space crawling manipulator can fuse position, signal strength and obstacle information in real time to achieve predictive tracking and autonomous obstacle avoidance in high dynamic and large displacement scenarios. This ensures the continuity and real-time performance of the communication link, while also improving the system's intelligence level and task execution efficiency in complex spatial environments.
[0027] Reference Figure 1 In this embodiment, the reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots includes: S0, Construct an external communication system for the space-based on-orbit maintenance robot, and realize external communication for the space-based on-orbit maintenance robot based on the external communication system.
[0028] In this embodiment, the reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots is implemented based on the constructed extravehicular communication system for space on-orbit maintenance robots. For example... Figures 2-3 As shown, the external communication system of the space on-orbit maintenance robot mainly includes: space on-orbit maintenance robot 1, space crawling robotic arm 2, in-cabin robot integrated control platform 3, and in-cabin robotic arm repetitive locking mechanism.
[0029] The space-based maintenance robot 1 is equipped with a wireless communication device A, which mainly includes a wireless communication module A and antennas A. The wireless communication module A is installed inside the body of the space-based maintenance robot 1, and the two antennas A are respectively installed at the center of the front and rear shells of the body of the space-based maintenance robot 1; the two antennas A are connected to the wireless communication module A through a feeder.
[0030] The space crawling robotic arm 2 is equipped with two sets of wireless communication devices: wireless communication device B and wireless communication device C. Wireless communication device B mainly includes a wireless communication module B and an antenna B; the wireless communication module B and antenna B are installed at the end of the first joint of the space crawling robotic arm 2, and antenna B is connected to the wireless communication module B via a feed line. Wireless communication device C mainly includes a wireless communication module C and an antenna C; the wireless communication module C and antenna C are installed at the end of the seventh joint of the space crawling robotic arm 2, and antenna C is connected to the wireless communication module C via a feed line.
[0031] The repetitive locking mechanism of the in-cabin robotic arm serves as the mechanical and electrical interface between the space crawling robotic arm 2 and the spacecraft cabin. There are four repetitive locking mechanisms in total: A41, B42, C43, and D44. These four mechanisms are located at the four corners near the hatch inside the spacecraft cabin.
[0032] The in-cabin robot integrated control platform 3 mainly includes: an integrated controller and wireless communication equipment D. Wireless communication equipment D mainly includes: a wireless communication module D and an antenna D. The integrated controller is connected to the wireless communication module D via electronic cables, and the antenna D is connected to the wireless communication module D via a feeder line.
[0033] The integrated controller is connected to the repeat locking mechanism of the four in-cabin robotic arms via electronic circuitry; the space crawling robotic arm 2 is electrically connected to the repeat locking mechanism of the in-cabin robotic arms via connectors.
[0034] The integrated controller can communicate wirelessly with wireless communication device A on the space-based maintenance robot directly through wireless communication device D; it can also communicate wirelessly with wireless communication device A on the space-based maintenance robot through wireless communication device D, with the space crawling robotic arm as a communication relay, via wireless communication devices B and C on the space crawling robotic arm.
[0035] S1. Based on the operating range of the space-based on-orbit maintenance robot, the communication area of the space-based on-orbit maintenance robot is divided into multiple sub-communication areas.
[0036] In this embodiment, firstly, the operating range of the space-based on-orbit maintenance robot is determined; such as... Figure 4As shown, the operational area is an ellipsoid centered on the spacecraft cabin. Further, based on the ellipsoid and considering the coverage of each wireless communication device and the signal obstruction caused by the metal cabin, the communication area of the space-based maintenance robot is divided into five sub-communication areas: Sub-communication Area A, Sub-communication Area B, Sub-communication Area C, Sub-communication Area D, and Sub-communication Area E. Sub-communication Area E is the area above the spacecraft cabin that is unobstructed by wireless communication device D. Using the planes formed by the X and Z axes and the Y and Z axes of the spacecraft cabin coordinate system, the remaining area of the ellipsoid is divided into four regions: Sub-communication Area A includes the repetitive locking mechanism A of the in-cabin robotic arm; Sub-communication Area B includes the repetitive locking mechanism B of the in-cabin robotic arm; Sub-communication Area C includes the repetitive locking mechanism C of the in-cabin robotic arm; and Sub-communication Area D includes the repetitive locking mechanism D of the in-cabin robotic arm.
[0037] S2, based on the division of multiple sub-communication areas, controls the space crawling robot arm to crawl between the repeating locking mechanisms of the in-cabin robot arm and fix it to the corresponding repeating locking mechanism of the in-cabin robot arm, using the space crawling robot arm as a communication relay between the space on-orbit maintenance robot and the integrated control platform of the in-cabin robot.
[0038] In this embodiment, the specific communication strategy for each sub-communication area is as follows: When the space-based maintenance robot is performing inspection and maintenance operations in the sub-communication area E, the integrated controller communicates directly with the wireless communication device A on the space-based maintenance robot through the wireless communication device D. The space crawling robotic arm performs its handling operations normally and is not affected.
[0039] When the space-based maintenance robot performs inspection and maintenance operations within sub-communication area A, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism A, and electrically connects with the cabin robotic arm repeating locking mechanism A via connectors. The integrated controller communicates wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D, using the space crawling robotic arm as a communication relay, through wireless communication devices B and C on the space crawling robotic arm. Of course, if communication is unobstructed, the integrated controller can also directly communicate wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D.
[0040] When the space-based maintenance robot performs inspection and maintenance operations within sub-communication area B, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding intra-cabin robotic arm repeating locking mechanism B, and electrically connects with the intra-cabin robotic arm repeating locking mechanism A via connectors. The integrated controller communicates wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D, using the space crawling robotic arm as a communication relay, through wireless communication devices B and C on the space crawling robotic arm. Of course, if communication is unobstructed, the integrated controller can also directly communicate wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D.
[0041] When the space-based maintenance robot performs inspection and maintenance operations within sub-communication area C, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding intra-cabin robotic arm repeating locking mechanism C, and electrically connects with the intra-cabin robotic arm repeating locking mechanism A via connectors. The integrated controller communicates wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D, using the space crawling robotic arm as a communication relay, through wireless communication devices B and C on the space crawling robotic arm. Of course, if communication is unobstructed, the integrated controller can also directly communicate wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D.
[0042] When the space-based maintenance robot performs inspection and maintenance operations within the sub-communication area D, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding intra-cabin robotic arm repeating locking mechanism D, and electrically connects with the intra-cabin robotic arm repeating locking mechanism A via connectors. The integrated controller communicates wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D, using the space crawling robotic arm as a communication relay, through wireless communication devices B and C on the space crawling robotic arm. Of course, if communication is unobstructed, the integrated controller can also directly communicate wirelessly with the space-based maintenance robot's wireless communication device A via wireless communication device D.
[0043] S3, after the space crawling robotic arm crawls and is fixed in the corresponding cabin robotic arm repeated locking mechanism, the space crawling robotic arm dynamic target tracking and obstacle avoidance control based on PPO-PID is adopted to complete the accurate tracking of dynamic targets and obstacle avoidance.
[0044] In this embodiment, the space crawling robotic arm is used as an intelligent agent. It autonomously learns obstacle avoidance and tracking strategies by sensing the distance deviation between the target object (the space-based maintenance robot) and obstacles. Based on the signal strength received by the transponder, it performs adaptive and precise pointing control of the robotic arm to ensure optimal communication between the space crawling robotic arm and the space-based maintenance robot at all times. This invention combines traditional PID control with the PPO reinforcement learning algorithm. First, based on the pose of the space-based maintenance robot, PID control adjusts the pose of the space crawling robotic arm, allowing its working plane to quickly approach the robot, achieving coarse localization of the robotic arm's pose. Then, the LSTM-PPO algorithm is used to enable the space crawling robotic arm to autonomously learn and track the robot's projection within the plane while avoiding obstacle projections. Among multiple solutions to the inverse kinematics of the space crawling robotic arm, possible singularities and solutions that might collide with the spacecraft's cabin or hatch are eliminated. The solution closest to the current pose of the space crawling robotic arm is selected from the remaining solutions, ultimately achieving stable tracking of the space-based maintenance robot and obstacle avoidance. Figure 5 As shown, the specific implementation process is as follows: S31, initialize the parameters of the space crawling robotic arm, PID parameters, LSTM parameters, and PPO parameters.
[0045] S32, obtain the pose of the space-based on-orbit maintenance robot and the space-crawling robotic arm in the spacecraft coordinate system.
[0046] In this embodiment, as Figure 6 As shown, the end effector pose T of the space crawling robotic arm in the spacecraft coordinate system {O} is... JXB The pose T of the space-based on-orbit maintenance robot in the spacecraft coordinate system {O} JQR The input is sent to the integrated controller as the initial pose.
[0047] S33 uses PID control to adjust the pose of the space crawling robot arm, enabling the working plane of the space crawling robot arm to quickly approach the space on-orbit maintenance robot, thus achieving coarse positioning of the space crawling robot arm.
[0048] In this embodiment, coarse positioning of the space crawling robotic arm is achieved through the following three constraints: Condition 1: The end effector of the space crawling robotic arm points towards the center of the on-orbit maintenance robot. That is, the line connecting the Z-axis of the space crawling robotic arm's end effector coordinate system {T} with the origin T of the space crawling robotic arm's end effector and the center R of the on-orbit maintenance robot. Coaxial.
[0049] Condition 2: Maintain a safe distance between the space crawling robotic arm and the space-based on-orbit maintenance robot. Set the minimum safe distance between the space crawling robotic arm and the space-based on-orbit maintenance robot to D. safe The distance between the point closest to the on-orbit maintenance robot and the on-orbit maintenance robot within the workspace of the space crawling robotic arm, which meets condition 1, is D. TR Select D safe D TR The larger value in the equation is used as the safe distance D between the space crawling robotic arm and the space on-orbit maintenance robot.
[0050] Condition 3: The end effector of the space crawling robotic arm should be pointed as far as possible towards the antenna A of the on-orbit maintenance robot. That is, the Z-axis of the end effector in the coordinate system {T} should be aligned with the direction of the antenna A. {T} The X-axis of the space-based on-orbit maintenance machine in coordinate system {R} {R} The included angle between them should be as small as possible. Therefore, the Z-position within the workspace of the crawling robot arm that satisfies conditions 1 and 2 is selected. {T} With X {R} The pose with the smallest included angle is used as the coarse positioning pose for the spatial crawling robotic arm.
[0051] Furthermore, Z can be... {T} With X {R} The included angle is taken as the error e, and the error feedback is used to minimize Z as much as possible through repeated iterations. {T} With X {R} The angle between the two sides allows the space crawling robot arm to roughly point towards the on-orbit maintenance robot, thus reducing the control problem of the space crawling robot arm in three-dimensional space to the control problem of the space crawling robot arm in two-dimensional plane.
[0052]
[0053] in, Indicates the control quantity of the robotic arm. Indicates the current error. This indicates the accumulation of error. Indicates proportional gain. Indicates integral gain. Represents differential gain. This indicates the amount of error change.
[0054] S34 utilizes the LSTM-PPO pose dynamic adjustment method based on signal strength to perform adaptive spatial crawling robot arm pointing precise control according to the signal strength received by the transponder, thereby achieving stable tracking and obstacle avoidance of the space on-orbit maintenance robot.
[0055] In this embodiment, after using PID control to adjust the pose of the space crawling manipulator so that its working plane quickly approaches the space-on-orbit maintenance robot, achieving coarse positioning of the manipulator, the LSTM-PPO algorithm can be used to dynamically adjust the manipulator's pose based on occlusion conditions and signal strength, thereby achieving stable tracking and obstacle avoidance of the space crawling manipulator on the space-on-orbit maintenance robot. Figure 7 As shown, the specific implementation process is as follows: S341, initialize the parameters of the Actor network, Critic network, and LSTM network. The Actor network is used to select actions for the space crawling robot based on the current state; the Critic network is used to evaluate the value of the selected actions; and the LSTM network is used to process the time-series data of the space crawling robot's state and the state of the space-based on-orbit maintenance robot.
[0056] S342 acquires timing state data and completes the state space design.
[0057] In this embodiment, to achieve real-time tracking of the space-crawling robotic arm against the space-on-orbit maintenance robot, it is necessary to acquire the status information of both the space-crawling robotic arm and the space-on-orbit maintenance robot, as well as environmental information, to assist the control decisions of the space-crawling robotic arm. Specifically: First, obtain the following three aspects of time-series state data: 1) The state data of the space crawling robot mainly includes: the current pose of the space crawling robot. Target pose of the space crawling robotic arm Angles of each joint of the space crawling robotic arm Length of each link in the space crawling robotic arm Angular velocity of each joint of the space crawling robotic arm Spatial crawling robotic arm step length Signal strength of wireless communication device receiver .
[0058] The state of the space crawling robotic arm is denoted as , can be represented as:
[0059] 2) Status data of the space-based on-orbit maintenance robot, mainly including: the current pose of the space-based on-orbit maintenance robot. Future pose of space-based on-orbit maintenance robot Signal strength of receiver A in wireless communication device .
[0060] The status of the space-based maintenance robot is recorded as... , can be represented as:
[0061] 3) Environmental data: Location of obstacles
[0062] Then, complete the state space design. Transform the state space using a vector. To indicate:
[0063] This state-space design incorporates all the information useful for the control decisions of the spatial crawling robotic arm, while also taking into account the dynamic changes and temporal dependencies of the environment.
[0064] S343 uses an LSTM network to filter time series state data and extract key time series state data.
[0065] In this embodiment, time-series state data can be input into an LSTM network for processing. Some state data that is not very useful for subsequent processing is discarded, and the state data at key time nodes is filtered out to form key time-series state data, which is then output. The key time-series state data output by the LSTM network will serve as the basis for subsequent reward calculation and action selection. Figure 8 As shown, the specific implementation method is as follows: S3431, Data Input: Input the timing state data into the LSTM network.
[0066] S3432, Input Gate Processing: After receiving the timing state data, the input gate, based on the current state data and the hidden state from the previous time step, determines whether the current state data and environmental information need to be written to the memory unit. A sigma activation function is used to perform a non-linear transformation on the input data to generate the activation value of the input gate. The activation value of the input gate is multiplied by the input data to obtain the update candidate value.
[0067] S3433, Forget Gate Processing: The forget gate receives the current state data and the hidden state from the previous time step, determining which state and environmental information need to be forgotten from the memory unit. An activation value for the forget gate is generated using the sigma activation function. This activation value is multiplied by the old state in the memory unit, determining which states are retained or forgotten.
[0068] S3434, Memory cell update: Add the updated candidate value processed by the input gate to the memory cell state processed by the forget gate to obtain the updated state of the memory cell. This step combines writing the current state and forgetting the state at a past time.
[0069] S3435, Output Gate Processing: The output gate receives the current input state data and the updated memory cell state, determining which states and environmental information need to be output. The sigma activation function generates the activation value of the output gate. This activation value is then multiplied by the updated state of the memory cell to obtain the current output data.
[0070] S3436, Select effective state data (i.e., state data at key time nodes): Based on the output of the LSTM network, select effective state data that can accurately reflect the motion state of the robotic arm and environmental information to form state data at key time nodes. These data will be used for the subsequent control and decision-making process of the space crawling robotic arm.
[0071] S3437, Iterative processing: Repeat steps S3431~S3436 to continuously process and select the status data of key time nodes to form key time series status data, ensuring that accurate status data can be obtained in real time in dynamic environments.
[0072] S344: Calculate dense and sparse rewards based on key time series state data to obtain a fused reward.
[0073] In this embodiment, the fusion reward mainly includes dense rewards and sparse rewards. Further, dense rewards mainly include position rewards and orientation rewards; sparse rewards mainly include signal strength, obstacle avoidance, acceleration guidance, adaptive step size rewards, etc., used to guide the space crawling robot arm to move towards areas with high communication quality, safe obstacle avoidance, and efficient movement. This can be adjusted according to the state of the space crawling robot arm. Status of space-based on-orbit maintenance robot Using the key time-series state data output by the LSTM network, we calculate the dense and sparse rewards for the space crawling robot. The dense reward, based on factors such as the position and orientation of the space crawling robot, is used to guide the robot to approach the target. The sparse reward, based on factors such as signal strength, obstacle avoidance, acceleration guidance, and adaptive step size, is used to prevent the space crawling robot from colliding with the spacecraft or the space-based maintenance robot and to improve the tracking speed. The dense and sparse rewards are then fused to obtain the final fused reward.
[0074] In the basic interaction process of reinforcement learning, the agent is the subject, the environment is the object, and the design of the reward function is a key part of the reinforcement learning task. Specifically: (a) Dense reward function design:
[0075] in, Indicates dense rewards, Indicates the position reward weight. Indicates the directional reward weight. This represents the normalized positional reward. This represents the directional reward after normalization.
[0076] Further:
[0077]
[0078] in, This indicates the current position coordinates of the space crawling robotic arm. This indicates the target position coordinates of the space crawling robotic arm. This represents the Euclidean distance between the current position and the target position of the spatial crawling robotic arm; This represents the quaternion indicating the current orientation of the crawling robotic arm. The quaternion represents the target orientation of the spatial crawling robotic arm.
[0079] (b) Sparse reward function design:
[0080] in, Indicates sparse rewards. Indicates the signal strength reward weight. Indicates the weight of obstacle avoidance rewards. This indicates that the core area is accelerating the guidance and reward weighting. This indicates the adaptive step-size reward weight. Indicates a signal strength reward. This indicates an obstacle avoidance reward. This indicates that rewards will be given to those who accelerate their efforts in the core area. This indicates an adaptive step-size reward.
[0081] Signal strength reward This is mainly used to encourage space crawling robotic arms to move in areas with high signal strength, as shown below:
[0082] when When communication is deemed to be in good condition, a substantial positive reward is given. when When communication is deemed to be in good condition, a small positive reward is given. when If the communication is deemed poor or disconnected, the task is judged as a failure and a huge negative reward is given.
[0083] Obstacle Avoidance Rewards This is primarily used to penalize collisions between space-crawling robotic arms and obstacles (such as spacecraft or space-based maintenance robots), as shown below:
[0084] when At that time, it was considered a safe zone, and no reward was given; among them, This indicates the distance between the arm or end effector of the space crawling robotic arm and the spacecraft or space-based maintenance robot. when At that time, it was considered a warning zone, and a warning was issued accordingly. Negative rewards that decrease linearly but increase; when At that time, it was considered a dangerous zone and was given a huge negative reward.
[0085] Core Area Accelerated Guidance Rewards This is primarily used to encourage space-crawling robotic arms to increase their tracking speed while remaining safe, as shown below:
[0086] when A reward will be given at that time; among them, This represents the distance between the end effector pose of the spatial crawling robotic arm and the target pose; when At that time, a compound reward will be given, which has an accelerating guiding effect; when When the tracking is successful, a large reward is added to the composite function to enhance the guidance.
[0087] Adaptive step size reward This is mainly used to encourage a space-crawling robotic arm to adjust its step size according to the current state, as shown below:
[0088] when When this happens, give positive rewards for larger strides; when At that time, give positive rewards for small steps.
[0089] (c) Design of the fusion reward function:
[0090] in, This indicates a fusion reward.
[0091] S345 uses the PPO planning algorithm to update the strategy network parameters, enabling the space crawling robotic arm to stably track and avoid obstacles for the space on-orbit maintenance robot.
[0092] In this embodiment, the PPO planning algorithm refers to: converting the state vector... Integration rewards The termination condition "Done" and other relevant state information are stored in a buffer. The Actor network selects an action based on the policy distribution. The Old Actor network stores the old policy distribution for comparison during policy updates in the PPO planning algorithm. Based on the selected action, the spatial crawling robot performs corresponding operations, interacts with the environment, and acquires new state and reward information. The Critic network then uses the state and reward information stored in the buffer... , The process involves calculating the value estimate of the current policy; comparing the Critic network's estimate with the actual reward received to calculate the error; and updating the parameters of the Actor and Critic networks using the PPO planning algorithm based on the calculated error. The PPO planning algorithm ensures the stability of the policy update by limiting the scope of the update.
[0093] Based on the structural characteristics of the spatial crawling robotic arm, the motion space design can be completed, and the motion of the spatial crawling robotic arm can be represented by a single vector. It is expressed as follows:
[0094] Among them, the angle increment of each joint of the space crawling robotic arm Range set linear velocity at the end effector of the space crawling robotic arm Range set angular velocity at the end effector of the space crawling robotic arm Range set .
[0095] S346 utilizes Kalman filtering to predict the future pose of the space-based on-orbit maintenance robot when it performs large displacement movements, and adjusts the pose of the space crawling robotic arm in advance to achieve predictive tracking.
[0096] In this embodiment, when the space-based maintenance robot needs to perform large-scale, rapid movements for its mission, the space crawling robotic arm may experience tracking delays. When the space-based maintenance robot's next action requires a large displacement, the space crawling robotic arm needs to perform predictive tracking to ensure stable communication for the space-based maintenance robot during rapid movements.
[0097] Based on the instructions issued by the in-cabin robot integrated control platform, the future pose of the space-based maintenance robot can be obtained. Then, by combining the historical pose with Kalman filtering, the future virtual pose of the space crawling robotic arm can be predicted.
[0098] in, This represents the future virtual pose of the space-crawling robotic arm. This indicates the current pose of the space crawling robotic arm. Indicates the time step. This indicates the current velocity of the end effector of the space crawling robotic arm. This represents the difference between the future pose and the current pose of the space-based on-orbit maintenance robot.
[0099] Furthermore, based on the predicted future virtual pose of the spatial crawling robot, random perturbations are added around the predicted future virtual pose to generate multiple candidate poses. Then, by evaluating the feasibility of the spatial crawling robot moving from its current position to the candidate pose and the time required to reach the candidate pose, the optimal pose is selected as the next target pose for the spatial crawling robot.
[0100] S347, Iterative Loop: Repeat steps S342~S2456 until the set number of training iterations is reached or the convergence condition is met. During each iteration, the spatial crawling robotic arm continuously interacts with the environment, learning and optimizing strategies through trial and error, thereby achieving effective tracking and obstacle avoidance control of dynamic targets.
[0101] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention by utilizing the methods and techniques disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.
[0102] The contents not described in detail in this specification are common knowledge to those skilled in the art.
Claims
1. A method for extravehicular communication relay and dynamic tracking of space robots based on reinforcement learning, characterized in that, include: Based on the operating range of the space-based on-orbit maintenance robot, the communication area of the space-based on-orbit maintenance robot is divided into multiple sub-communication areas; Based on the division of multiple sub-communication regions, the space crawling robot arm is controlled to crawl between the repeating locking mechanisms of the in-cabin robot arm and is fixed to the corresponding repeating locking mechanism of the in-cabin robot arm. The space crawling robot arm serves as a communication relay between the space on-orbit maintenance robot and the integrated control platform of the in-cabin robot. After the space crawling robotic arm crawls and is fixed in the corresponding cabin robotic arm repeated locking mechanism, the space crawling robotic arm dynamic target tracking and obstacle avoidance control based on PPO-PID is adopted to complete the accurate tracking of dynamic targets and obstacle avoidance.
2. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 1, characterized in that, Also includes: An external communication system for a space-based maintenance robot is constructed, enabling external communication for the robot. This system comprises: a space-based maintenance robot, a space-crawling robotic arm, an in-cabin robot integrated control platform, and an in-cabin robotic arm repetitive locking mechanism. The space-based maintenance robot is equipped with wireless communication device A, which includes a wireless communication module A and an antenna A. The wireless communication module A is installed inside the robot's torso, and two antennas A are respectively installed at the center of the front and rear outer shells of the robot's torso. The two antennas A are connected to the wireless communication module A via feed lines. The space-crawling robotic arm is equipped with wireless communication devices B and C. Device B includes a wireless communication module B and an antenna B, both installed at the end of the first joint of the robotic arm, with the antenna B connected to the wireless communication module B via a feed line. Device C includes a wireless communication module C and an antenna C, both installed at the end of the seventh joint of the robotic arm, with the antenna C connected to the wireless communication module C via a feed line. The in-cabin robotic arm... The repeatable locking mechanism serves as the mechanical and electrical interface between the space crawling robotic arm and the spacecraft cabin. There are four repeatable locking mechanisms for the robotic arm inside the cabin: mechanism A, mechanism B, mechanism C, and mechanism D. These four mechanisms are located at the four corners near the hatch inside the spacecraft cabin. The integrated control platform for the in-cabin robot includes an integrated controller and wireless communication equipment D. Wireless communication equipment D includes a wireless communication module D and an antenna D. The integrated controller connects to the wireless communication module D via electronic cables. The antenna D is connected to the wireless communication module D via a feeder; the integrated controller is connected to the repeating locking mechanism of the four in-cabin robotic arms via electronic circuitry; the space crawling robotic arm is electrically connected to the repeating locking mechanism of the in-cabin robotic arms via connectors; the integrated controller communicates wirelessly directly with the wireless communication device A on the space-on-orbit maintenance robot via the wireless communication device D, or the integrated controller communicates wirelessly with the wireless communication device A on the space-on-orbit maintenance robot via the wireless communication device D, using the space crawling robotic arm as a communication relay, through the wireless communication devices B and C on the space crawling robotic arm.
3. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 2, characterized in that, Based on the operational range of the space-based maintenance robot, its communication area is divided into multiple sub-communication areas. This includes: defining the operational range of the space-based maintenance robot as an ellipsoid centered on the spacecraft cabin; based on the ellipsoid, and considering the coverage of various wireless communication devices and signal obstruction by the metal cabin, dividing the space-based maintenance robot's communication area into five sub-communication areas: Sub-communication Area A, Sub-communication Area B, Sub-communication Area C, Sub-communication Area D, and Sub-communication Area E; Sub-communication Area E is the area above the spacecraft cabin that is unobstructed by wireless communication device D; Using the planes formed by the X and Z axes of the spacecraft cabin coordinate system and the planes formed by the Y and Z axes, the remaining area of the ellipsoid is divided into four areas: Sub-communication Area A includes the repetitive locking mechanism A of the in-cabin robotic arm; Sub-communication Area B includes the repetitive locking mechanism B of the in-cabin robotic arm; Sub-communication Area C includes the repetitive locking mechanism C of the in-cabin robotic arm; and Sub-communication Area D includes the repetitive locking mechanism D of the in-cabin robotic arm.
4. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 3, characterized in that, Based on multiple sub-communication regions, the space crawling robotic arm is controlled to crawl between the repeating locking mechanisms of the in-cabin robotic arm and is fixed to the corresponding repeating locking mechanism of the in-cabin robotic arm. The space crawling robotic arm serves as a communication relay between the space on-orbit maintenance robot and the integrated control platform of the in-cabin robot, including: When the space-based maintenance robot is performing inspection and maintenance operations in the sub-communication area E, the integrated controller communicates directly with the wireless communication device A on the space-based maintenance robot through the wireless communication device D. The space crawling robotic arm performs its handling operations normally without being affected. When the space-based maintenance robot performs inspection and maintenance operations in sub-communication area A, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism A, and electrically connects with the cabin robotic arm repeating locking mechanism A through a connector; the integrated controller communicates wirelessly with the wireless communication device A on the space-based maintenance robot through wireless communication device D, using the space crawling robotic arm as a communication relay, and through wireless communication devices B and C on the space crawling robotic arm; When the space-based maintenance robot performs inspection and maintenance operations in sub-communication area B, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding cabin robotic arm repeating locking mechanism B, and electrically connects with the cabin robotic arm repeating locking mechanism A through a connector; the integrated controller communicates wirelessly with the space crawling robotic arm as a communication relay via wireless communication device D, wireless communication device B and wireless communication device C on the space crawling robotic arm, and wireless communication device A on the space-based maintenance robot. When the space-based maintenance robot performs inspection and maintenance operations in the sub-communication area C, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism C, and electrically connects with the cabin robotic arm repeating locking mechanism A through a connector; the integrated controller communicates wirelessly with the space crawling robotic arm as a communication relay via wireless communication device D, wireless communication device B and wireless communication device C on the space crawling robotic arm, and wireless communication device A on the space-based maintenance robot. When the space-based maintenance robot performs inspection and maintenance operations within the sub-communication area D, it controls the space crawling robotic arm to crawl along a preset trajectory to the corresponding position of the cabin robotic arm repeating locking mechanism D, and electrically connects with the cabin robotic arm repeating locking mechanism A via a connector; the integrated controller communicates wirelessly with the space-based maintenance robot via wireless communication device D, using the space crawling robotic arm as a communication relay, and through wireless communication devices B and C on the space crawling robotic arm.
5. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 4, characterized in that, A PPO-PID-based spatial crawling robotic arm is employed for dynamic target tracking and obstacle avoidance control, enabling precise tracking of dynamic targets and obstacle avoidance, including: Initialize the parameters of the space crawling robotic arm, including PID parameters, LSTM parameters, and PPO parameters; Obtain the pose of the space-based maintenance robot and the space-crawling robotic arm in the spacecraft coordinate system; By using PID control to adjust the pose of the space crawling robot arm, the working plane of the space crawling robot arm can be quickly brought close to the space on-orbit maintenance robot, thus achieving coarse positioning of the space crawling robot arm. Based on the occlusion conditions and signal strength, the LSTM-PPO algorithm is used to perform adaptive spatial crawling robot arm pointing precision control, so as to realize the stable tracking and obstacle avoidance of the space crawling robot arm on the space on-orbit maintenance robot.
6. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 5, characterized in that, Coarse positioning of the space crawling robotic arm is achieved by using the following three constraints: Condition 1: The end effector of the space crawling robotic arm points towards the center of the space-based on-orbit maintenance robot; Condition 2: Maintain a safe distance between the space crawling robotic arm and the space on-orbit maintenance robot; Condition 3: Point the end effector of the space crawling robotic arm as far as possible toward the antenna A of the space-based maintenance robot.
7. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 5, characterized in that, Based on occlusion conditions and signal strength, the LSTM-PPO algorithm is used for adaptive, precise pointing control of the space crawling robotic arm. This enables the space crawling robotic arm to stably track and avoid obstacles from the on-orbit maintenance robot, including: S1, initialize the parameters of the Actor network, Critic network, and LSTM network; S2, acquire timing state data and complete state space design; S3. Using an LSTM network, the time series state data is filtered to extract key time series state data. S4. Calculate the dense reward and sparse reward based on the key time series state data to obtain the fused reward; S5 uses the PPO planning algorithm to update the strategy network parameters, enabling the space crawling robotic arm to stably track and avoid obstacles on the space-based on-orbit maintenance robot. S6, when the space-based on-orbit maintenance robot is making large displacement movements, Kalman filtering is used to predict the future pose of the space-based on-orbit maintenance robot, and the pose of the space crawling robotic arm is adjusted in advance to achieve predictive tracking. S7. Repeat steps S2 to S6 until the set number of training iterations is reached or the convergence condition is met.
8. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 7, characterized in that, The Actor network is used to select actions for the space crawling robot based on the current state; the Critic network is used to evaluate the value of the selected actions; and the LSTM network is used to process the time-series data of the space crawling robot's state and the space on-orbit maintenance robot's state.
9. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 7, characterized in that, Through vectors To represent the state space: in, This indicates the current pose of the space crawling robotic arm. This indicates the target pose of the space crawling robotic arm. This indicates the angles of each joint of the space crawling robotic arm. This indicates the length of each link in the space crawling robotic arm. This represents the angular velocity of each joint of the space crawling robotic arm. This indicates the step length of the space crawling robotic arm. Indicates the signal strength of the receiver in a wireless communication device. This indicates the current pose of the space-based on-orbit maintenance robot. This indicates the future pose of the space-based maintenance robot. This indicates the signal strength of receiver A in wireless communication device A.
10. The reinforcement learning-based extravehicular communication relay and dynamic tracking method for space robots according to claim 7, characterized in that, The fusion rewards include dense rewards and sparse rewards. Dense rewards include position rewards and orientation rewards. Sparse rewards include signal strength, obstacle avoidance, acceleration guidance, and adaptive step size rewards, which guide the space crawling robot arm to move towards areas with high communication quality, safe obstacle avoidance, and efficient movement.
Citation Information
Patent Citations
Communication method for spacecraft and outboard target equipment
CN107682073A
Mechanical arm tail end trajectory tracking algorithm based on null space obstacle avoidance
CN113146610A
A method and system for online target point tracking of a robotic arm based on a pose tracking system
CN113561183B
Satellite remote relay measurement and control docking antenna adaptive directional control ground equipment
CN117527036A
Matching method and system based on 5G network, Beidou positioning and non-line-of-sight microwave
CN119667742A