Robot motion control method and system in microgravity environment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-04-02
- Publication Date
- 2026-07-21
Smart Images

Figure CN121957041B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of robot motion control technology, and in particular relates to robot motion control methods and systems in microgravity environments. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Quadruped robots occupy an important position in complex environment operations due to their excellent terrain adaptability and motion stability. With the rapid development of deep space exploration technology, microgravity environments such as the moon (gravitational acceleration of about 1.62 m / s²) have become key application scenarios for quadruped robots. Their core requirements are to achieve stable gait, disturbance resistance, and high-efficiency movement control to adapt to tasks such as lunar surface exploration and scientific experiments.
[0004] However, existing model-based and reinforcement learning-based control methods are mainly designed for Earth's gravity environment, and have many key shortcomings when directly applied to microgravity environments, making it difficult to meet the needs of practical applications: First, in microgravity, the supporting reaction force of a robot is significantly reduced, making its posture highly susceptible to drifting, tilting, or even instability. Existing reinforcement learning strategies lack targeted posture constraint mechanisms, and the reward function design does not consider gravity adaptability, resulting in extremely poor robot motion stability under microgravity. Simultaneously, the robot's ground contact feedback and airborne characteristics are fundamentally altered under microgravity. Traditional ground contact thresholds and posture restrictions are incompatible with the microgravity environment, further exacerbating the difficulty of posture control and making it prone to instability phenomena such as mid-air rotation, severely impacting operational reliability.
[0005] Secondly, the lunar surface is covered with soft and dry lunar regolith, and existing foot trajectory planning methods have significant shortcomings. On the one hand, trajectories are mostly implicitly generated by the policy network or use simple linear interpolation and sine curve design, lacking explicit constraints on the starting and landing points of the swing phase. This easily leads to foot dragging and sliding, further exacerbating lunar regolith disturbance and dust storms. On the other hand, the trajectories are not optimized for the mechanical properties under microgravity, resulting in abrupt changes in lift speed and significant landing impacts. This not only induces dust storms, severely affecting visual observation accuracy and damaging equipment surfaces, but also reduces the contact stability between the feet and the ground, hindering the robot's continuous movement in unstructured terrain. Furthermore, existing reward functions lack sufficient design for specific terms such as dust prevention and soft landing, failing to effectively guide the policy to learn low-disturbance trajectories.
[0006] Third, the existing collaborative mechanism of reinforcement learning and trajectory generation lacks gravity adaptation design, and the unreasonable weight allocation of the reward function leads to low motion tracking accuracy, poor energy efficiency, and insufficient vertical stability of the body under microgravity. At the same time, the trajectory generation module and the reinforcement learning controller have a high degree of coupling, poor versatility, and are difficult to adapt to different types of quadruped robots and the complex and varied terrain conditions on the lunar surface.
[0007] In summary, fundamentally solving both the stability problem and the sand-generating problem of robot motion under microgravity is an urgent issue that needs to be addressed. Summary of the Invention
[0008] To overcome the shortcomings of the prior art, this invention provides a robot motion control method and system in a microgravity environment, which enables the robot to move stably, with low disturbance and high efficiency in a microgravity environment, while effectively suppressing sandstorms.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for robot motion control in a microgravity environment, comprising: Determine whether the robot has entered the swinging phase based on the robot's motion state. When the robot enters the swinging phase, predict the ideal foot position when the swing ends based on the current position of the robot's foot. Based on the starting point and ideal foot position of the robot's swing phase, the robot's swing phase is divided into three sub-phase motions, and the desired trajectory of the robot's foot is generated for each sub-phase motion. Robot state information, environmental information, and historical trajectory features are input into a pre-built joint control model to predict the robot's joint control commands in a microgravity environment. The joint control model is a policy network model trained by reinforcement learning. In the training of the policy network model, the reward function includes trajectory tracking reward. The trajectory tracking reward is determined by the robot's actual position and the expected position of the robot's foot.
[0010] Secondly, the present invention provides a robot motion control method in a microgravity environment, comprising: The prediction module is configured to: determine whether the robot has entered the swing phase based on the robot's motion state; when the robot enters the swing phase, predict the ideal foot position at the end of the swing based on the current foot position of the robot. The trajectory module is configured to divide the robot's swing phase into three sub-phase motions based on the starting point and ideal foot position, and generate the desired trajectory of the robot's foot for each sub-phase motion. The control module is configured to input robot state information, environmental information, and historical trajectory features into a pre-built joint control model to predict the robot's joint control commands in a microgravity environment. The joint control model is a policy network model trained by reinforcement learning. In the training of the policy network model, the reward function includes a trajectory tracking reward. The trajectory tracking reward is determined by the robot's actual position and the desired position of the robot's foot.
[0011] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0012] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0013] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0014] The above one or more technical solutions have the following beneficial effects: In this invention, when the robot enters the swinging phase, the ideal foot position at the end of the swing is predicted based on the current position of the robot's foot. The swinging phase of the robot is divided into three sub-phases, and the expected trajectory of the robot's foot is generated for each of the three sub-phases. The generated expected trajectory is directly used as the core of the reinforcement learning reward calculation, guiding the policy learning to achieve precise matching between trajectory tracking and motion control.
[0015] In this invention, by integrating reinforcement learning and trajectory planning, the robot can move stably, with low disturbance and high efficiency in a microgravity environment, while effectively suppressing sandstorms.
[0016] In this invention, a velocity-based adaptive phase update strategy is adopted to enable the gait frequency to be automatically adjusted according to the motion state. This design ensures that when the robot is stationary, there is still a basic phase growth to ensure the gait preparation state; when moving, the greater the synthetic velocity, the faster the phase growth, and the gait frequency is adaptively improved; the turning angular velocity is also taken into consideration to ensure timely gait response when turning.
[0017] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0019] Figure 1 This is an overall framework diagram of the robot motion control method in a microgravity environment according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the trajectory generation stage in Embodiment 1 of the present invention. Detailed Implementation
[0020] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0021] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0022] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0023] Example 1 This embodiment discloses a robot motion control method in a microgravity environment, including: Determine whether the robot has entered the swinging phase based on the robot's motion state. When the robot enters the swinging phase, predict the ideal foot position when the swing ends based on the current position of the robot's foot. Based on the starting point and ideal foot position of the robot's swing phase, the robot's swing phase is divided into three sub-phase motions, and the desired trajectory of the robot's foot is generated for each sub-phase motion. Robot state information, environmental information, and historical trajectory features are input into a pre-built joint control model to predict the robot's joint control commands in a microgravity environment. The joint control model is a policy network model trained by reinforcement learning. In the training of the policy network model, the reward function includes trajectory tracking reward. The trajectory tracking reward is determined by the robot's actual position and the expected position of the robot's feet.
[0024] In this embodiment, when the robot enters the swinging phase, the ideal foot position at the end of the swing is predicted based on the current position of the robot's foot. The swinging phase is divided into three sub-phases, and the desired trajectory of the robot's foot is generated for each sub-phase. The generated desired trajectory is directly used as the core of the reinforcement learning reward calculation, guiding policy learning to achieve precise matching between trajectory tracking and motion control. This embodiment achieves stable, low-disturbance, and high-efficiency movement of the robot in a microgravity environment through the integrated design of reinforcement learning and trajectory planning, while effectively suppressing sandstorms.
[0025] The robot motion control method in a microgravity environment proposed in this embodiment will be described in detail below: (a) Microgravity-adaptive reinforcement learning network 1. Network architecture setup.
[0026] It adopts an Actor-Critic dual-network architecture, implemented in PyTorch, and uses the NP3O (PPO variant) training algorithm, which incorporates a history encoder and a world model.
[0027] History encoders are used to compress observation sequences over a period of time into low-dimensional temporal features. This enhances the observability and noise resistance of states in microgravity environments; the world model is used to extract environmental dynamics from current observations and historical features, characterize microgravity contact and dynamic changes, and provide dynamic information for strategies.
[0028] The input to the Actor network consists of three parts: the robot's own state information (joint angles, joint angular velocities, body posture quaternions, body angular velocity, body linear velocity, foot contact force, terrain height, motion commands, etc.), and a low-dimensional representation of a historical trajectory. The input consists of three elements: environmental dynamics extracted from the world model, and the combination of these elements to form a complete policy input.
[0029] The Actor network backbone is a 3-layer fully connected structure with hidden layer dimensions of 512, 256, and 128 respectively, and the activation function is ELU. The output is a 12-dimensional joint motion command, which is obtained by sampling through a Gaussian policy (initial standard deviation 1.0, learnable) to obtain the control quantity. The motion scale is subject to safety constraints in the environment control module (action_scale=0.25).
[0030] The Critic network input includes the robot's own state information (joint angles, joint angular velocities, body posture quaternions, body angular velocity, body linear velocity, foot contact force, terrain height, motion commands, etc.) and environmental dynamic features extracted from the world model. The two are concatenated to form a complete evaluation input. The structure also adopts a 3-layer fully connected structure (512, 256, 128). The outputs are value estimates and cost estimates, which are used to estimate the long-term reward of the current state. It also includes a cost evaluation branch to support safety constraints.
[0031] The optimization uses Adam with a learning rate of The entropy regularization coefficient is 0.01. Training uses on-based sampling trajectories. The policy update method involves sampling a fixed step size in each iteration and repeating multiple rounds of small-batch updates, with the total number of iterations set to 10,000.
[0032] 2. Microgravity attitude constraint.
[0033] This embodiment is designed for microgravity environments, and specifically designs the robot's ground contact detection, take-off time target and attitude constraints to suppress the adverse strategies caused by "aerial attitude compensation" and improve attitude stability and safety under microgravity.
[0034] In a microgravity environment, the supporting reaction force is significantly lower than that in Earth's environment. To ensure the accuracy of ground contact detection, this embodiment scales the ground contact normal force threshold by 0.4 times that of Earth's environment. The airtime is prolonged under microgravity; therefore, the airtime target is adjusted to 0.6 seconds.
[0035] This embodiment specifies the core constraints for fuselage attitude and angular velocity: pitch angle |pitch|≤5°, roll angle |roll|≤5°, and yaw rate of change | |≤10° / s, in order to solve the key problem of air dynamics under microgravity, that is, in the microgravity environment, once the fuselage angular momentum accumulates, there is almost no natural damping. If there is a lack of constraints, the strategy can easily compensate for the speed through air pitch / roll, or even form an undesirable solution of "air roll + landing correction", which seriously affects the motion stability.
[0036] To address this, on the one hand, the attitude penalty weight (pitch / roll direction) was significantly increased by 50%, guiding the strategy to actively maintain attitude stability through continuous reward feedback; on the other hand, the angular velocity constraint strength was further increased, reducing the upper limit of the fuselage rotation angular velocity from 15° / s in the Earth's environment to 10° / s, strictly limiting the fuselage rotation speed and preventing instability in mid-air from the source; at the same time, a roll / pitch threshold termination mechanism was added, triggering training termination immediately when |pitch| or |roll| exceeds 60°, thus promptly avoiding serious safety risks.
[0037] Within each control cycle (Δt=0.005s), the projected gravity direction is calculated using fuselage quaternions and angular velocity, and attitude stability is assessed in real time accordingly. The fuselage height is estimated using terrain sampling points and used for the altitude-keeping penalty. The target height is set to 0.35m, and the altitude-keeping penalty coefficient is set to 6.0 to avoid floating or sinking in microgravity environments.
[0038] 3. Design of microgravity-adapted reward function.
[0039] This embodiment uses a linear weighted summation form to construct the total reward function:
[0040] in, Assign weights to each sub-item. These correspond to sub-rewards or penalties. The total reward consists of multiple basic reinforcement learning rewards and microgravity-specific stability rewards, while also incorporating specialized rewards adapted to foot trajectory, precisely matching the control requirements of the microgravity environment.
[0041] The basic reinforcement learning reward component focuses on optimizing weight allocation for microgravity characteristics: (1) Speed tracking reward: Contact is unstable in microgravity environment, and the strategy tends to "save effort" and give up speed tracking.
[0042] Therefore, the weight of speed tracking rewards is significantly increased, making it the primary driver of strategy optimization. The tracking rewards for linear velocity and yaw rate are as follows:
[0043]
[0044] in, This represents the exponential decay coefficient in speed tracking rewards, used to control the sensitivity of speed errors to rewards; Let x be the components of the desired linear velocity of the machine body in the x and y directions of the machine body coordinate system. These are the components of the actual linear velocity of the machine body in the x and y directions of the machine body coordinate system; The desired yaw rate; This is the actual yaw rate. The weighting was increased from 1.5 for the Earth's environment to 2.0. The weighting was moderately increased from 0.5 to 0.8; This represents the square of the L2 norm.
[0045] (2) Constraints on energy consumption and motion smoothness: Motion availability is greater under microgravity, but excessive shaking will lead to energy waste and unstable control.
[0046] Therefore, this embodiment adopts the design principle of "weakening but not zeroing" for energy consumption and smoothing terms to suppress meaningless oscillations, while retaining sufficient "degrees of freedom of motion" to adapt to low-gravity dynamics.
[0047] The corresponding penalties include: Action change rate penalty:
[0048] The weight of the action change rate penalty term was reduced from -0.03 for the Earth environment to -0.01.
[0049] Joint acceleration penalty:
[0050] The weight of the joint acceleration penalty term is reduced from -5e-7 for the Earth environment to -7e-8.
[0051] Torque energy consumption penalty:
[0052] The weight of the torque energy consumption penalty term was reduced from -0.0001 for the Earth environment to -0.00005.
[0053] in, and The control actions at the current time t and the previous time, i.e., time t-1; and The joint velocities at the current time t and the previous time, i.e., time t-1; Joint torque at this moment; This represents the square of the L2 norm.
[0054] In actual debugging, practical judgment criteria can be followed: if the machine body is stable and the speed tracking is good but the joints are shaking, the penalty weight can be increased appropriately; if the speed tracking is lagging and the feet cannot reset in time, the penalty weight should be reduced to retain a certain "stickiness" so that the strategy can choose the necessary actions autonomously. (3) Vertical stability: Under microgravity conditions, the risk of floating and sinking increases significantly, which is a precursor to motion instability.
[0055] This embodiment strengthens the vertical stability constraint, forming a control logic of "more aggressive horizontally, but must be stable vertically", which effectively prevents floating under microgravity.
[0056] The corresponding penalty items are: Vertical velocity stability term :
[0057] fuselage height retention item :
[0058] in, Let h be the vertical velocity of the aircraft, and h be the height of the aircraft. For target altitude.
[0059] The weights for the vertical speed stability term and the fuselage altitude maintenance term are -1.5 and -6.0, respectively, to improve the correction capability.
[0060] Building upon this foundation, additional specialized rewards (anti-slip reward, translational stability reward, soft landing reward, and trajectory tracking reward) adapted to the third-order smooth trajectory generation module are added. The specific calculation logic for these rewards is supported by the trajectory generation module. By evaluating the fit between the actual foot trajectory and the desired trajectory, the strategy is guided to output joint control commands that meet the requirements for preventing sand blowing, thereby achieving deep synergy between reinforcement learning and trajectory planning.
[0061] The overall design embodies the principles of ensuring controllability through basic rewards, optimizing weights to adapt to microgravity, and enhancing synergy through specific rewards. Simultaneously, the "positive-only reward" mechanism performs non-negative pruning after summing, preventing excessive negative feedback from impacting early training stability. This hierarchical structure effectively suppresses undesirable strategies involving mid-air rolls followed by landing corrections, while simultaneously improving motion continuity and safety under microgravity.
[0062] (II) Generation of three-order smooth foot trajectory for sand blowing prevention In a microgravity environment, the third-order smooth foot trajectory generation and reinforcement learning network work together through a closed loop of trajectory generation-reward guidance-policy optimization. The trajectory generation module provides the reinforcement learning with an accurate anti-dust reference trajectory, while the reinforcement learning policy outputs joint control commands that conform to the trajectory constraints based on a microgravity-adapted reward function.
[0063] Step 1: Gait phase update and motion state determination.
[0064] In this embodiment, the reinforcement learning controller outputs a continuous motion command vector. ,in, Forward speed command (unit: m / s). This is a lateral speed command (unit: m / s). This is the yaw rate command (unit: rad / s). The command update frequency is the same as the control frequency, which is 200Hz. = 0.005s). In actual deployment, this command comes from remote control commands.
[0065] To distinguish whether the robot is in motion or standing still, this implementation sets two judgment thresholds. and The decision logic is as follows:
[0066] in, Forward speed command, For lateral speed command, This is the yaw rate command.
[0067] If moving is true, the robot enters a moving state and performs subsequent gait phase updates and trajectory generation; if it is false, the robot enters a stationary state, keeping its feet in the default support position to avoid unnecessary energy consumption and joint vibration under microgravity.
[0068] This implementation employs a velocity-based adaptive phase update strategy, enabling gait frequency to automatically adjust according to the motion state. Specifically, the gait phase of each of the four legs... (i = 0, 1, 2, 3 correspond to left front leg FL, right front leg FR, left hind leg RL, and right hind leg RR respectively) Update according to the following rules:
[0069] Wherein, the phase growth function Designed as follows:
[0070]
[0071] Where mod 1 means modulo 1, It is a yaw rate command; Forward speed command, For lateral speed command, For time intervals.
[0072] This design ensures that when the robot is stationary, The basic phase growth is 0, but there is still a basic phase growth (1.2) to ensure the gait preparation state; during movement, the greater the synthetic velocity, the faster the phase growth, and the gait frequency is adaptively improved; the steering angular velocity is also taken into consideration to ensure timely gait response during steering.
[0073] To achieve diagonal gait (Trot), this embodiment sets the initial phase offset vectors for the four legs. for:
[0074] That is, the left front leg and the right hind leg are in phase, and the right front leg and the left hind leg are in phase, forming a diagonal synchronous swing pattern, which avoids gait disorder caused by phase deviation under microgravity.
[0075] Step 2: Calculate the swing start point buffer and target landing point.
[0076] This implementation defines the phase interval of the swing phase as follows: The support phase is When the phase of a certain leg is detected to be from... Jump to At that time, it is determined that the leg has entered the swing phase.
[0077] At the instant of entering the swing phase, the precise position of the current foot in the body coordinate system is obtained through the simulation engine or robot sensors. The starting point is then stored in a cache variable. Throughout the subsequent swing phase (approximately 0.4 phase cycles), the starting point remains unchanged, avoiding trajectory starting point drift caused by state estimation noise, sensor latency, or jitter in the output of the reinforcement learning policy.
[0078] At the start of the swing phase, this implementation predicts the ideal landing point at the end of the swing based on the current body movement commands and the duration of the swing:
[0079] in, The initial position at the instant of entering the swing phase. This represents the displacement.
[0080] Displacement The calculation comprehensively considers forward motion, lateral motion, and turning motion, including the forward displacement component. :
[0081] in, The estimated duration of the oscillation phase is inversely coupled to the phase growth function, meaning that as the phase growth coefficient increases... The corresponding decreases, and vice versa; This is the forward step size scaling factor; The speed-adaptive gain allows for a moderately increased stride at high speeds. This is a forward velocity command. Ultimately... Restricted to Between these steps, prevent excessive stride length from causing instability.
[0082] Lateral displacement components :
[0083] in, This is an approximate turning radius, defaulting to the distance between the front and rear of the foot; The turning contribution coefficient, The symbol indicates the direction of the leg's side swing; This is a lateral speed command; It is a yaw rate command.
[0084] height component During the initial calculation, it is temporarily set to AND. Similarly, due to the undulations in the ground, subsequent adjustments were made based on the terrain elevation.
[0085] target landing point The calculation is performed once at the beginning of the swing phase and remains locked throughout the entire swing phase to prevent the landing point from drifting continuously during the swing due to continuous changes in motion commands, thereby ensuring the certainty of the trajectory's endpoint.
[0086] Step 3: Generating a smooth third-order foot trajectory.
[0087] The core innovation of this implementation method lies in defining the initial position and target landing point of the swing phase. The process is further refined into three functionally defined sub-stages, and a smooth trajectory function is designed for each stage.
[0088] The total length of the swing phase is set to the phase interval [0, 0.4), which is then divided into the lifting phase. ∈[0,0.125), translation stage ∈[0.125,0.275), landing stage ∈[0.275,4), which can be finely adjusted according to the robot structure and terrain characteristics.
[0089] The lifting phase trajectory lifts the foot vertically off the ground, avoiding horizontal dragging. (Horizontal coordinates) Keep the starting point The vertical coordinate z remains unchanged; a cosine acceleration curve is used from the starting point. Smoothly raise the preset swing height .
[0090]
[0091] This function is in When the speed is zero, The speed is also zero, achieving continuous acceleration and effectively avoiding the "sudden acceleration" phenomenon during start-up; The preset swing height; This is the boundary between the lifting phase and the translation phase.
[0092] The translational phase trajectory enables the feet to move horizontally above the target landing point while maintaining height stability, based on the elevation height. Within the horizontal plane... To the target point Perform linear interpolation; height remains constant. .
[0093]
[0094] in, This refers to the phase at the boundary between the lifting and translation phases; This refers to the phase at the boundary between the translation phase and the landing phase. The preset swing height; x ( ), y ( ), z ( () is phase The position at the foot.
[0095] The landing phase trajectory shows the foot descending vertically from the swing height to the target height. Achieve a soft landing. The horizontal coordinates are locked at... The vertical coordinate is smoothly reduced using a cosine deceleration curve.
[0096]
[0097] in, The phase marking the end of the landing phase; It is the proportion of the time already elapsed in the landing phase out of the total landing phase time; It is the proportion of the time elapsed in the translation phase out of the total translation phase time. This refers to the phase at the boundary between the lifting and translation phases; This is the phase at the boundary between the translation phase and the landing phase.
[0098] This function also ensures that the velocity is zero at the moment of landing, greatly reducing the impact force between the feet and the ground.
[0099] With the above design, the foot trajectory at the stage boundary (i.e. , and (At the point) the position, velocity, and acceleration are continuous.
[0100] Step 4: Terrain adaptation and support locking.
[0101] The trajectories described above were all generated in the machine's coordinate system. To incorporate terrain information, they need to be transformed to the world coordinate system.
[0102] in, The quaternion posture of the organism. This is the rotation matrix corresponding to the quaternion. Let be the foot position vector in the body coordinate system. This represents the position of the machine in the world coordinate system.
[0103] In this embodiment, a pre-generated terrain height map is used to query target points in real time. Corresponding ground height During the landing phase, the target altitude It was directly corrected to:
[0104] in, To ensure a safety margin, the landing point always conforms to the actual terrain.
[0105] When detected At this point, the foot transitions from the swing phase to the support phase. The current foot position (already corrected for terrain) is locked and saved as the support position. This position remains unchanged throughout the support phase until the next swing phase begins.
[0106] Step 5: Decoupling and cooperating with the reinforcement learning controller.
[0107] This embodiment guides the reinforcement learning policy to learn and track the generated foot trajectory by designing a dedicated reward function, without requiring the trajectory information as an observation input to the policy. The trajectory generation module operates independently, determining the stage the robot is in based on its current phase during the swing phase and calculating the expected foot position in the volume coordinate system in real time. These expected positions are used only for reward calculation and are not directly input into the policy network.
[0108] In this embodiment, four reward items directly related to trajectory quality are designed to complement the microgravity adaptation reward function of the reinforcement learning network, jointly guiding the strategy to output joint torques and making the actual foot position approach the desired trajectory. In this embodiment, the anti-scratching reward coefficient is 5.0, the translational phase stability reward coefficient is 3.0, the soft landing reward coefficient is 3.0, and the trajectory tracking reward coefficient is 5.0. m=4 represents the four feet of the robot, ensuring that the reward weights at each stage are adapted to the control requirements under microgravity.
[0109] Anti-mopping reward :
[0110] in, Let i be the horizontal velocity vector of foot i. Let i be the vertical velocity of the foot tip. During the lifting phase, the feet are encouraged to generate upward vertical speed to avoid dragging with the ground.
[0111] Stable rewards during the translation phase :
[0112] in, The actual height of foot tip i For the desired height, This causes the oscillation to remain highly constant during the middle stage.
[0113] Soft landing rewards :
[0114] in, and Let i be the perpendicular contact force of foot i at this moment t and the previous moment (t-1). The sudden change in contact force during the landing phase is punished.
[0115] Tracking rewards :
[0116] Directly drive the actual foot position To the desired location near, .
[0117] These rewards, combined with traditional motion rewards (speed tracking, energy efficiency, posture stability, etc.), constitute a multi-objective reward function. During training, the reinforcement learning strategy gradually learns a joint torque control strategy capable of accurately tracking a third-order smooth trajectory by maximizing cumulative rewards. The expected trajectory provided by the trajectory generator serves as a high-quality reference signal, significantly reducing the randomness of strategy exploration, accelerating training convergence, and ensuring the standardization and safety of the ultimately learned gait, thus achieving effective synergy between trajectory planning and reinforcement learning under microgravity.
[0118] Example 2 The purpose of this embodiment is to provide a robot motion control method in a microgravity environment, including: The prediction module is configured to: determine whether the robot has entered the swing phase based on the robot's motion state; when the robot enters the swing phase, predict the ideal foot position at the end of the swing based on the current foot position of the robot. The trajectory module is configured to divide the robot's swing phase into three sub-phase motions based on the starting point and ideal foot position, and generate the desired trajectory of the robot's foot for each sub-phase motion. The control module is configured to input robot state information, historical trajectory, and environmental information into a pre-built joint control model to predict joint control commands for the robot in a microgravity environment. The joint control model is a policy network model trained by reinforcement learning. In the training of the policy network model, the reward function includes a trajectory tracking reward. The trajectory tracking reward is determined by the robot's actual position and the desired position of the robot's foot.
[0119] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0120] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0121] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0122] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0123] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0124] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0125] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0126] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0127] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0128] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0129] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for controlling robot motion in a microgravity environment, characterized in that, include: Determine whether the robot has entered the swinging phase based on the robot's motion state. When the robot enters the swinging phase, predict the ideal foot position when the swing ends based on the current position of the robot's foot. Based on the starting point and ideal foot position of the robot's swing phase, the swing phase is divided into three sub-phases, and the desired trajectory of the robot's foot is generated for each sub-phase; specifically: The robot's swinging phase is divided into a lifting phase, a translation phase, and a landing phase. During the lifting phase, the robot's horizontal coordinates remain unchanged, while its vertical coordinates are smoothly lifted to the preset swing height using a cosine acceleration curve. During the translation phase, the robot's horizontal coordinates are linearly interpolated while the vertical coordinates remain unchanged. During the landing phase, the robot's horizontal coordinates represent the ideal foot position, while its vertical coordinates use a cosine deceleration curve for a smooth descent. By inputting robot state information, historical trajectory and environmental information into a pre-built joint control model, the joint control commands of the robot in a microgravity environment can be predicted. The joint control model is a policy network model trained through reinforcement learning; in the training of the policy network model, the reward function includes a trajectory tracking reward; the trajectory tracking reward is determined by the robot's actual position and the desired position of the robot's foot. The reward function also includes an anti-mopping reward. for: ; in, Let i be the horizontal velocity vector of foot i. Let i be the vertical velocity of the foot tip. and All are coefficients, where m is the number of robot feet.
2. The robot motion control method in a microgravity environment as described in claim 1, characterized in that, The reward function also includes speed tracking reward, motion change rate penalty, joint acceleration penalty, torque energy consumption penalty, vertical speed stability term and fuselage altitude maintenance term; the speed tracking reward includes linear speed tracking reward and yaw rate tracking reward.
3. The robot motion control method in a microgravity environment as described in claim 1, characterized in that, Define the phase intervals for the swing phase and the support phase respectively, and determine whether the robot leg has entered the swing phase based on the phase change of the robot leg.
4. The robot motion control method in a microgravity environment as described in claim 1, characterized in that, The reward function also includes translation phase stability rewards and soft landing rewards; The stable reward during the translation phase for: ; in, The actual height of foot tip i For the desired height, For coefficients; The soft landing reward for: ; in, and These represent the vertical contact forces at foot i at time t and time t-1, respectively. is a coefficient; m is the number of robot feet; The trajectory tracking reward for: ; in, This refers to the actual position of the foot. m represents the desired position; m is the number of robot feet. is a coefficient.
5. The robot motion control method in a microgravity environment as described in claim 1, characterized in that, It also includes: employing a velocity-based adaptive phase update strategy to enable the gait frequency to automatically adjust according to the robot's motion state, specifically: ; ; ; Where mod 1 means modulo 1, It is a yaw rate command; Forward speed command, For lateral speed command, For time intervals; The gait phase of the i-th leg is indicated by the superscript t, which indicates time t.
6. A method for controlling robot motion in a microgravity environment, characterized in that, include: The prediction module is configured to: determine whether the robot has entered the swing phase based on the robot's motion state; when the robot enters the swing phase, predict the ideal foot position at the end of the swing based on the current foot position of the robot. The trajectory module is configured to: divide the robot's swing phase into three sub-phases based on the starting point and ideal foot position, and generate the desired trajectory for the robot's foot for each sub-phase; specifically: The robot's swinging phase is divided into a lifting phase, a translation phase, and a landing phase. During the lifting phase, the robot's horizontal coordinates remain unchanged, while its vertical coordinates are smoothly lifted to the preset swing height using a cosine acceleration curve. During the translation phase, the robot's horizontal coordinates are linearly interpolated while the vertical coordinates remain unchanged. During the landing phase, the robot's horizontal coordinates represent the ideal foot position, while its vertical coordinates use a cosine deceleration curve for a smooth descent. The control module is configured to: input robot state information, historical trajectory, and environmental information into a pre-built joint control model to predict joint control commands for the robot in a microgravity environment; the joint control model is a policy network model trained by reinforcement learning; in the training of the policy network model, the reward function includes trajectory tracking reward; the trajectory tracking reward is determined by the robot's actual position and the desired position of the robot's foot. The reward function also includes an anti-mopping reward. for: ; in, Let i be the horizontal velocity vector of foot i. Let i be the vertical velocity of the foot tip. and All are coefficients, where m is the number of robot feet.
7. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-5.
9. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-5.
Citation Information
Patent Citations
Traction type blind guiding leg and arm robot
CN115944509A
Detection robot movement control method and device under conditions of small sky body and weak gravitation
CN116755432A