A pipeline robot motion control method and system

Through multimodal sensor data fusion and strategy model driving, pipeline robots realize autonomous motion control, solving the problems of motion inflexibility and signal interference in complex pipeline environments, and improving stability and safety.

CN120170753BActive Publication Date: 2025-08-08PEKING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510647478.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-08
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing pipeline robots are inflexible in complex pipeline environments, prone to collisions or lags, and difficult to achieve high-precision control, especially in dynamic obstacles and signal interference scenarios, which affect their efficiency and safety throughput.

Method used

Multimodal sensor fusion is used to obtain pipeline environment and robot status data, and a pre-trained strategy model is used to generate control instructions to drive the motor to realize autonomous motion control, including safe docking and exit path planning when communication is interrupted.

Benefits of technology

It significantly improves the motion stability and load capacity of the pipeline robot, reduces mechanical wear, enhances the motion coherence and safety in complex environments, and reduces the dependence on artificial remote control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120170753B_ABST
    Figure CN120170753B_ABST
Patent Text Reader

Abstract

This invention discloses a pipeline robot motion control method and system, relating to the field of robot control technology. The method comprises: responding to a direction command triggered by an operator via a remote terminal, acquiring status data of the pipeline robot body collected by an internal sensor module, as well as pipeline environment data collected by an external sensor module; generating control commands for each drive motor using a pre-trained strategy model based on the direction command, status data, and pipeline environment data; and adjusting the operating state of each drive motor according to the control command. This invention can achieve autonomous motion control of the pipeline robot in complex pipeline environments, reducing manual dependence and improving the robot's motion stability, load capacity, and throughput efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control technology, and in particular to a pipeline robot motion control method and system. Background Art

[0002] Currently, pipeline robots rely primarily on manual remote control, requiring operators to manually adjust the rotation direction and speed of each motor. However, in complex pipeline environments, such as those with dynamic obstacles inside the pipeline or when rapid obstacle avoidance is required, manual operation often results in inflexible robot movement due to delayed responses or operational errors, even leading to collisions or stalls, severely limiting their efficiency and safety. Especially in emergency situations, the operator's real-time decision-making capabilities struggle to meet the demands of high-precision control.

[0003] Existing cable-free pipeline robots typically use wireless signal transmission for remote control, but practical applications often face issues with signal interference and transmission delays. When the robot navigates complex pipeline structures (such as bends and reducers) or needs to execute multiple coordinated commands, wireless signal instability can easily lead to command loss or response delays, making precise motion control difficult. For example, if critical commands are not transmitted in a timely manner during obstacle avoidance, the robot may become stuck due to inconsistent movements, or even damage the pipeline or its own structure, seriously compromising operational reliability.

[0004] In addition, the internal environment of the pipeline is complex and changeable, with impurities such as oil and condensed water often adhering to the inner wall, and the friction coefficient in different areas varies significantly due to differences in material, temperature, and humidity. Existing manual control methods make it difficult to perceive environmental conditions in real time and dynamically adjust the drive strategy, resulting in unreasonable torque distribution of the drive wheels and easy wheel slippage. This not only reduces the robot's longitudinal driving force and load capacity, but also increases mechanical wear and shortens the service life of the equipment. Especially in long-distance or high-load scenarios, existing technologies have difficulty balancing motion efficiency and stability, limiting the widespread application of pipeline robots in industrial inspection, maintenance and other fields. Summary of the Invention

[0005] In order to solve the problems existing in the prior art, the present invention provides a pipeline robot motion control method and system, which can realize autonomous motion control of the pipeline robot in a complex pipeline environment, reduce manual dependence, and improve the robot's motion stability, load capacity and passing efficiency.

[0006] To achieve the above objectives, the present invention provides a pipeline robot motion control method, which is applied to a pipeline robot processor; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of movable joint units; each of the joint units includes one or more drive motors; the method includes:

[0007] In response to a direction instruction triggered by an operator through a remote terminal, the state data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module are obtained;

[0008] Based on the direction instruction, the state data and the pipeline environment data, a pre-trained strategy model is used to generate a control instruction for each of the drive motors;

[0009] According to the control instruction, the operating state of each driving motor is adjusted.

[0010] Optionally, the method further includes:

[0011] When a communication interruption with the remote end is detected, a safe area in the pipeline is determined by the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each of the drive motors;

[0012] Control each of the driving motors to execute a corresponding docking instruction, so that the pipeline robot moves to the safe area.

[0013] Optionally, the method further includes:

[0014] When it is detected that the duration of the communication interruption exceeds a preset threshold, a driving path pointing to a target pipeline exit is determined based on a pipeline route map; the pipeline route map is pre-stored or generated in real time based on the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-designated pipeline exit;

[0015] Based on the state data, pipeline environment data and driving path, generating an exit instruction for each of the drive motors through the strategy model;

[0016] Each driving motor is controlled to execute a corresponding exit instruction, so that the pipeline robot moves to the target pipeline outlet.

[0017] Optionally, the method further includes:

[0018] When it is detected based on the pipeline environmental data that there is a pending area in the driving direction and no new direction instruction is received within a preset time, a prompt message is sent to the remote end; the pending area is a pipeline section with branch channels, turning channels and / or inaccessible areas.

[0019] Optionally, the method further includes:

[0020] If no new direction instruction is received within a preset period of time after the prompt message is sent, a safe area in the pipeline is determined by the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each of the drive motors;

[0021] Control each of the driving motors to execute a corresponding docking instruction, so that the pipeline robot moves to the safe area.

[0022] Optionally, the training method of the policy model includes:

[0023] Constructing a dynamic simulation model of the pipeline robot and a corresponding pipeline simulation training environment;

[0024] According to the preset motion task objectives, a composite reward function is constructed that integrates posture stability reward, slippage suppression reward, and task execution reward.

[0025] In the pipeline simulation training environment, the control strategy of the drive motor in the dynamic simulation model is iteratively updated through a proximal strategy optimization algorithm. When the composite reward function converges to a preset value, the trained strategy model is output.

[0026] Optionally, the method further includes:

[0027] In the simulation-to-physical migration stage, the trained policy model is calibrated for dynamic environmental parameters using a domain adaptation algorithm; the dynamic environmental parameters include at least sensor measurement error parameters.

[0028] The present invention also provides a pipeline robot motion control system, which is applied to a pipeline robot processor; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of movable joint units; each of the joint units includes one or more drive motors; the system includes:

[0029] A data acquisition unit, configured to acquire, in response to a direction instruction triggered by an operator via a remote terminal, the state data of the pipeline robot body acquired by the internal sensor module and the pipeline environment data acquired by the external sensor module;

[0030] an instruction generation unit, configured to generate control instructions for each of the drive motors based on the direction instruction, the state data, and the pipeline environment data using a pre-trained strategy model;

[0031] A motion control unit is used to adjust the operating state of each of the drive motors according to the control instructions.

[0032] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0033] The pipeline robot motion control method provided by this invention acquires real-time pipeline environmental information, robot body posture information, and drive motor operating information, and inputs this into a pre-trained strategy model to autonomously generate control commands for each drive motor, significantly reducing reliance on manual remote control. In complex pipeline networks, the robot can autonomously adjust the drive wheel torque based on the pipeline environment (such as obstacle distribution, pipeline structural changes, and differences in pipe wall friction coefficient). This effectively prevents wheel slippage, improves longitudinal driving force and load capacity, reduces mechanical wear, and extends the equipment's service life.

[0034] During manual control, the operator only needs to issue simple motion direction commands (such as forward, backward, or turn), and the strategy model will combine the input data to generate refined motor control instructions. This mechanism significantly reduces reliance on high-frequency, multi-dimensional remote control commands, thereby reducing the amount of signal transmission data, alleviating the problem of command packet loss caused by wireless signal interference or delay, and enhancing the robot's movement consistency and reliability in complex pipe structures (such as bends and reducers).

[0035] In addition, based on the autonomous decision-making capabilities of the strategy model, the robot can quickly respond to dynamic obstacles and sudden working conditions, achieve flexible obstacle avoidance and stable cornering, overcome problems such as collisions and jams caused by reaction delays or operational errors in traditional manual control, and significantly improve operational efficiency and safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present invention.

[0037] Figure 1 This is a schematic diagram of a method flow chart of a pipeline robot motion control method according to an embodiment of the present invention;

[0038] Figure 2 This is a schematic diagram of active control of the pipeline robot during pipeline movement according to an embodiment of the present invention;

[0039] Figure 3 This is a schematic diagram of the driving structure of a pipeline robot according to an embodiment of the present invention;

[0040] Figure 4 A schematic diagram of the strategy model training process shown in an embodiment of the present invention;

[0041] Figure 5 This is a schematic diagram illustrating the principle of a training method based on reinforcement learning according to an embodiment of the present invention;

[0042] Figure 6A schematic diagram of a training test scenario and a simulation test algorithm migration diagram illustrating an embodiment of the present invention;

[0043] Figure 7 This is a schematic diagram of the module structure of the pipeline robot motion control system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] The purpose of the present invention is to provide a pipeline robot motion control method and system, which can realize autonomous closed-loop motion control in complex pipeline environments based on the fusion of strategy models and multimodal sensors, reduce dependence on manual remote control operations, and improve the robot's motion stability, load capacity and passing efficiency in scenarios with dynamic obstacles, sudden changes in friction coefficients and signal interference.

[0046] See Figure 1 , Figure 1 The figure is a flow chart of the pipeline robot motion control method.

[0047] The above pipeline robot motion control method is applied to the processor of the pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of movable joint units; each joint unit includes one or more drive motors; the method includes:

[0048] Step 101: In response to a direction instruction triggered by an operator through a remote terminal, status data of the pipeline robot body collected by an internal sensor module and pipeline environment data collected by an external sensor module are obtained.

[0049] In the application, after the operator sends a directional command via a remote terminal (such as the cloud or a host computer), the pipeline robot enters the command response phase. The remote terminal, acting as the human-machine interface, receives directional commands such as forward, reverse, pause, or turn input from the operator. After receiving the directional command, the pipeline robot acquires data collected by internal and external sensor modules.

[0050] For example, the internal sensor module may include an inertial measurement unit (IMU), an angle encoder, and a torque sensor. The IMU monitors the robot's pitch, roll, and yaw angles in real time, while the angle encoder records the rotation angle and angular velocity of each drive motor. The torque sensor provides feedback on the torque output of each drive wheel and / or drive motor. These data collectively constitute the in-pipe robot's state data, representing the robot's current posture, motion state, and the real-time load on the power system.

[0051] At the same time, the external sensor module activates environmental perception, acquiring pipeline environmental data through multimodal sensors. For example, these might include lidar or vision sensors to scan the three-dimensional geometric features of the pipeline's inner wall (such as diameter changes, bend curvature, and the location of the reducer). Proximity sensors detect the distance and position of obstacles, and friction coefficient sensors analyze the effects of pipe wall material and surface deposits (such as oil and condensate) on motion resistance. This pipeline environmental data encompasses not only static structural information but also the real-time location and trajectory of dynamic obstacles, ensuring the robot's comprehensive awareness of the complex working conditions.

[0052] Status and environmental data are transmitted to the processor via the communication module for integration. The processor performs preprocessing on this raw data, such as noise filtering, coordinate alignment, and data fusion. Furthermore, during data acquisition, sensor validity is verified in real time. If a sensor anomaly is detected, a redundant data switching mechanism is triggered to ensure the continuity and reliability of information collection.

[0053] See also Figure 2 and Figure 3 , Figure 2 This is a schematic diagram of the active control of the pipeline robot during pipeline movement. Figure 3 This is a schematic diagram of the pipeline robot drive structure. and Stands for active drive, 、 and Stands for transmission drive. Figure 2 The numbers 1 and 2 in the figure represent the tension spring and the passive rotation respectively.

[0054] In applications, operators can remotely control the pipeline robot through a remote terminal. The pipeline robot has five active degrees of freedom and seven passive degrees of freedom, for a total of 12 degrees of freedom for an intelligent robot. Furthermore, due to the use of omnidirectional wheels for local motion, each equipped with 12 small rollers, there are six omnidirectional wheels totaling 72 local degrees of freedom. Regarding the pipeline robot's drive structure, the front and rear spherical wheels are driven by the drive motors of the first and fourth joint units of the pipeline robot, while the center rollers are driven by the drive motors of the second and third joint units through a bevel gear transmission. Each drive motor is equipped with an angle encoder and a torque sensor, and each segment is equipped with an inertial measurement unit (IMU) to sense the posture angles (pitch, roll, and yaw) of each segment. Furthermore, angle sensors are installed between the segments. In addition to the shaft connection between the first and second segments and the third and fourth segments, tension springs are also used to connect the segments. The front and rear spherical wheels can rotate independently around the cross axis, with four rotational degrees of freedom. The second and third trunk sections are connected by an axis, with one passive rotational degree of freedom.

[0055] Step 102: Based on the direction instruction, the state data and the pipeline environment data, a pre-trained strategy model is used to generate control instructions for each drive motor.

[0056] In the application, direction commands, state data, and pipeline environment data are fed into a pre-trained policy model as input. The direction command, as a high-level control goal, specifies the robot's overall motion intention, such as forward, backward, or turning.

[0057] The strategy model can be formed through reinforcement learning training optimization. The model input performs multimodal fusion of direction instructions, state data and pipeline environment data, and forms a unified state representation vector after preprocessing. This vector is input into the strategy network, and the network analyzes the optimal motion strategy under the current working conditions based on the pre-learned dynamic relationship and environmental interaction experience. Specifically, the model dynamically balances the torque distribution and rotation direction of each drive motor by calculating composite reward functions such as posture stability, slip suppression and task execution, ensuring that in complex pipeline scenarios, the movement direction required by the direction instruction can be achieved, and the wheel slip can be actively suppressed to enhance the longitudinal driving force. For example, in cornering conditions, the model will dynamically adjust the torque difference between the inner and outer drive wheels according to the friction coefficient of the pipe wall and the robot posture angle to avoid skidding or understeer caused by uneven torque distribution.

[0058] The model outputs refined control instructions for each drive motor, including target torque, rotation direction, and force-position hybrid control parameters. These control instructions are converted into pulse-width modulated signals by the underlying motor controller to drive the actuators of each joint unit. For example, to address the reduced friction coefficient caused by oil contamination on the inner wall of a pipe, the model reduces the torque output of the drive wheel to prevent slippage. At the same time, force-position hybrid control ensures that the contact force between the wheel and the pipe wall remains within a reasonable range, thereby improving motion stability.

[0059] The target torque is used to adjust the output torque of the drive motor. The longitudinal driving force when the drive wheel contacts the pipe wall can be set by adjusting the current or voltage signal. The rotational direction parameter determines the forward / reverse rotation of the motor rotor by changing the winding power-on timing or control logic of the drive motor, thereby determining the robot's overall motion direction (forward, backward, or turning). The force-position hybrid control parameter coordinates the position closed-loop and torque closed-loop control of the drive motor, combining the real-time displacement feedback of the angle encoder with the dynamic force feedback of the torque sensor. This achieves dual constraints on the drive wheel's displacement accuracy and contact force stability under complex friction conditions. For example, when the pipe wall friction coefficient suddenly changes, the force-position hybrid control can dynamically adjust the output characteristics of the drive motor, ensuring that the wheel moves along the preset trajectory while maintaining the wheel-wall contact force within a safe threshold, thereby avoiding slippage caused by insufficient contact force or mechanical damage caused by contact force overload.

[0060] Furthermore, during the simulation-to-physical migration phase, the policy model undergoes domain adaptation algorithm calibration, dynamically compensating for sensor measurement errors and deviations from real-world environmental parameters to ensure the reliability and environmental adaptability of control commands. Ultimately, the generated control commands are transmitted via a bus to each drive motor, enabling autonomous closed-loop motion control of the pipeline robot under complex working conditions.

[0061] Step 103: Adjust the operating state of each drive motor according to the control instruction.

[0062] During the execution of instructions, control instructions are transmitted via the bus to the underlying controller of each drive motor. The controller interprets the instructions as specific pulse-width modulation signals or current control signals, driving the motor to perform the corresponding action. For example, for a certain drive wheel, if the instruction requires increasing the torque to overcome the friction reduction caused by oil contamination on the pipe wall, the controller will increase the duty cycle of the input current and increase the output torque of the motor; if the direction of rotation needs to be adjusted to achieve steering, the direction of rotor rotation is changed by switching the power-on sequence of the motor windings. At the same time, the force-position hybrid control mode dynamically adjusts the output characteristics of the motor by collecting position data fed back by the angle encoder and force feedback data from the torque sensor in real time, ensuring a balance between the target position and the contact force, and avoiding wheel slippage or motion stagnation caused by overload or underload.

[0063] During the adjustment process, the actual operating status of each drive motor is continuously monitored, including real-time torque, speed, and position deviation. This data is fed back to the processor via an internal sensor module, forming a closed-loop control circuit. For example, if the actual torque of a drive wheel deviates from the target value, a dynamic compensation mechanism is triggered. This mechanism, combining the posture data from the inertial measurement unit with the changes in the friction coefficient detected by the environmental sensors, fine-tunes the control signal to eliminate the error. This mechanism allows the operating status of each drive motor to be precisely controlled, ultimately achieving stable motion for the pipeline robot under complex working conditions.

[0064] In one embodiment, the pipeline robot motion control method further includes:

[0065] When a communication interruption with the remote end is detected, the safe area in the pipeline is determined through the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each drive motor;

[0066] Control each drive motor to execute the corresponding docking command to move the pipeline robot to a safe area.

[0067] In the application, upon detecting a communication interruption with the remote end, the robot can immediately initiate an autonomous emergency response mechanism. Communication interruptions can be caused by signal interference, transmission equipment failure, or environmental obstruction. In these cases, the remote end cannot continue to send directional commands, forcing the robot to rely on its local policy model for autonomous decision-making. Pipeline environmental data is continuously collected by external sensor modules, while status data is provided by internal sensor modules.

[0068] The strategy model uses a pre-trained reinforcement learning strategy to assess safe zones within the current pipeline environment. Safe zones must meet the following criteria: be located within straight pipe sections or low-curvature bends to avoid structural instability; be located away from dynamic obstacles or areas where static obstacles are concentrated to reduce collision probability; and have a high pipe wall friction coefficient and minimal surface debris to ensure effective grip for the drive wheels. Furthermore, the model uses a multi-objective optimization algorithm, combining prior knowledge of the pipeline roadmap (such as branch layout and outlet locations) with real-time sensor data to calculate the priority and accessibility of safe zones and ultimately select the optimal target location.

[0069] When generating a docking instruction, the strategy model plans an obstacle avoidance path based on the location of the safe area and the current state of the robot, and decomposes it into the control parameters of each drive motor. For example, if the safe area is located in the rear straight pipe section, the model will calculate the torque distribution and rotation direction combination of the drive wheels to make the robot retreat with the minimum turning radius, while adjusting the force-position hybrid control parameters of the intermediate rollers to balance the body posture. The docking instruction contains the target torque, rotation direction and displacement increment of each motor to ensure that the robot maintains longitudinal driving force and lateral stability during movement. During the execution phase, the underlying controller converts the instruction into a motor drive signal, and accurately controls the movement of each drive wheel by dynamically adjusting the pulse width modulation duty cycle or switching the power-on timing of the motor windings.

[0070] During the movement, the changes in the environment and the robot's state can be continuously monitored. If a dynamic obstacle is detected invading the planned path or a sudden change in the friction coefficient of the pipe wall is detected, the strategy model will update the safety area assessment results in real time and regenerate the docking instructions. For example, when a new obstacle appears in the backward path, the model may switch to the side bend section as a temporary safety area and achieve lateral translation by adjusting the torque difference between the ball wheel and the roller. All actions are achieved through closed-loop control. Internal sensors provide real-time feedback on the motor execution status and body posture, and external sensors update environmental data, forming an autonomous cycle of "perception-decision-execution-verification" until the robot successfully reaches the safe area and enters the standby state, waiting for communication to be restored or subsequent instructions to be triggered.

[0071] In one embodiment, the pipeline robot motion control method further includes:

[0072] When it is detected that the communication interruption duration exceeds a preset threshold, a driving path pointing to the target pipeline exit is determined based on the pipeline route map; the pipeline route map is pre-stored or generated in real time based on pipeline environmental data; the target pipeline exit is the nearest pipeline exit or a pre-designated pipeline exit;

[0073] Based on the status data, pipeline environment data and driving path, the strategy model generates the exit command of each drive motor;

[0074] Control each drive motor to execute the corresponding exit command to move the pipeline robot to the target pipeline outlet.

[0075] In the application, if a communication interruption is detected lasting longer than a preset threshold, the system determines that the remote end cannot restore the connection in the short term and initiates an autonomous exit mechanism. This preset threshold can be flexibly configured based on the risk level of the pipeline operation scenario. For example, a shorter threshold can be set in flammable, explosive, or high-humidity environments to prioritize equipment safety. Pipeline route maps, serving as the basis for path planning, come in two types: a static map pre-built using 3D modeling or historical inspection data, containing information about pipeline branches, exit coordinates, and key nodes; and a topological map dynamically generated based on real-time pipeline environmental data. This map uses simultaneous localization and mapping (SLAM) technology, a fusion of lidar, vision sensors, and odometry, to update pipeline structural features and obstacle distribution in real time. Pipeline route maps can be generated by either a policy model or other functional modules, such as the route construction module. The target pipeline exit is selected based on a priority rule. If multiple exits exist, the closest one is selected to shorten the exit path. If a specific exit (such as a maintenance hatch or a safety compartment) is preset, that exit is the final destination.

[0076] The strategy model generates an exit command based on the current robot state data (including drive motor load, body posture, and remaining battery power), pipeline environment data (such as obstacle location and pipe wall friction coefficient), and the geometric constraints of the driving path. The driving path must meet the requirements of obstacle avoidance, low energy consumption, and motion stability: the model plans the global optimal path through a path search algorithm (such as A or RRT), and dynamically adjusts the local obstacle avoidance trajectory based on real-time sensor data. For example, when there are dynamic obstacles in the path, the model calculates the detour trajectory and decomposes it into the torque distribution and rotation direction parameters of each drive motor to ensure that the robot avoids the obstacle with the minimum turning radius. The exit command specifically includes the target torque gradient of the drive motor (such as slow acceleration to avoid wheel slippage), the segmented rotation direction sequence (such as turning left to adjust the heading before going straight), and the force-position hybrid control parameters (such as reducing torque in smooth pipe sections and enhancing the displacement closed-loop accuracy).

[0077] During the execution phase, the drive motors adjust their operating states based on the exit instructions. For example, in narrow curved sections, the torque of the middle roller is limited to prevent excessive squeezing of the pipe wall, while the front and rear ball wheels achieve precise steering through differential rotation. In straight pipe sections, all drive motors simultaneously increase torque to accelerate movement. Furthermore, the execution effect can be monitored in real time through internal sensors. If the actual motion trajectory deviates from the planned path (e.g., wheel slippage due to pipe wall deformation), the strategy model will recalculate the path and update the instructions. At the same time, sensor errors (such as accumulated odometer deviation) can be compensated online through domain adaptive algorithms to ensure positioning accuracy in real-world environments.

[0078] In one embodiment, the pipeline robot motion control method further includes:

[0079] When it is detected based on the pipeline environmental data that there is a pending area in the driving direction and no new direction instructions are received within the preset time, a prompt message is sent to the remote end; the pending area is a pipeline section with branch channels, turning channels and / or inaccessible areas.

[0080] While the pipeline robot is traveling along a preset path, the external sensor module continuously collects pipeline environment data, including the geometric structure scanned by the lidar, the characteristics of pipeline branches captured by the visual sensor, and the distribution of obstacles detected by the proximity sensor. When a pending area is detected ahead in the direction of travel, the path decision interruption mechanism is triggered. The pending area refers to the pipeline section where the passage path needs to be determined, such as the intersection of branch channels (left, right or upward branch needs to be selected), sharp turn channels (curvature exceeds the preset safety threshold), or areas that are impassable due to collapse, foreign object blockage, etc. The basis for determining such areas includes sudden changes in pipe diameter, excessive obstacle density, or new structures not marked in the topological map.

[0081] In the application, the location, type, and environmental parameters of the pending area (such as the azimuth angle, turning radius, or obstruction size of a branch passage) can be integrated into prompt information and sent to the remote user interface via a wireless communication module. The prompt information can include the highlighted 3D coordinates of the pending area, suggested travel options (such as a list of possible directions for a branch passage), and environmental risk assessment results (such as a warning of insufficient friction coefficient in a turning passage).

[0082] After sending the prompt message, a command waiting timer of a preset length can be started. If the direction command returned by the remote end is not received within the preset time, it is determined that the operator cannot respond in time and the robot is immediately switched to the safe decision-making mode. The policy model can re-evaluate the passability of the pending area based on real-time pipeline environmental data and robot body status data. For example, for branch channels, the policy model can select the path with the highest probability based on the connectivity analysis of the pipeline roadmap; for inaccessible areas, a detour trajectory is generated based on historical obstacle avoidance strategies; if autonomous decision-making is still not feasible, the pending area can be marked as a special area, and the reverse backtracking mechanism is triggered to control the robot to retreat along the original path to the nearest safe node, while continuously sending environmental status updates to the remote end until communication is restored or manual intervention completes path replanning.

[0083] Throughout the entire process, internal sensor modules synchronously monitor the robot's real-time status, including drive motor load, body posture stability, and energy consumption rate, to ensure that no secondary failures occur during the waiting or reverse movement phases due to power overload or posture imbalance. For example, when pausing before a sharp turn, the strategy model dynamically adjusts the torque distribution between the ball wheel and roller to offset the body roll torque caused by the tilt of the pipe wall. In addition, the generation and transmission of prompt information can both use redundant communication protocols to ensure that information exchange can be completed through multi-band transmission or data compression technology in weak signal environments, avoiding decision delays due to communication packet loss.

[0084] In one embodiment, the pipeline robot motion control method further includes:

[0085] If no new direction command is received within a preset period of time after the prompt message is sent, the safe area in the pipeline is determined through the strategy model based on the pipeline environment data and status data, and the docking command for each drive motor is generated;

[0086] Control each drive motor to execute the corresponding docking command to move the pipeline robot to a safe area.

[0087] Among them, the determination of the safe area, the generation and execution of the docking instruction, etc. can refer to the relevant introduction in the above embodiments and will not be repeated here.

[0088] In one embodiment, the training method of the policy model includes:

[0089] Build a dynamic simulation model of the pipeline robot and the corresponding pipeline simulation training environment;

[0090] According to the preset motion task objectives, a composite reward function is constructed that integrates posture stability reward, slippage suppression reward, and task execution reward.

[0091] In the pipeline simulation training environment, the control strategy of the drive motor in the dynamic simulation model is iteratively updated through the proximal policy optimization algorithm. When the composite reward function converges to the preset value, the trained policy model is output.

[0092] The application builds a dynamic simulation model of the pipeline robot. This model, based on the robot's physical structural parameters (such as mass distribution, joint connection stiffness, and drive motor characteristics) and kinematic constraints (such as the contact geometry of the omnidirectional wheels and the rotational degrees of freedom of the spherical wheels), accurately replicates the motion response characteristics of a real robot through multi-body dynamics software. The pipeline simulation training environment simulates a variety of pipeline scenarios, including straight pipes, curved pipes, reducers, tees, and dynamic obstacle distribution. It also sets different environmental parameters such as pipe wall friction coefficient, temperature, and humidity to cover the complex working conditions likely to be encountered in actual operations.

[0093] In a simulation environment, training objectives can be quantified using a composite reward function. For example, the attitude stability reward calculates the penalty for deviations from the robot's equilibrium state based on the pitch, roll, and angular velocity data fed back by the inertial measurement unit, encouraging the model to maintain stability during motion. The slip suppression reward assesses the degree of slippage by comparing the actual and theoretical linear speeds of the drive wheels, dynamically adjusting the penalty weight based on the friction coefficient of the pipe wall to optimize the torque distribution strategy. The task execution reward is directly linked to pre-set motion goals, such as progress rewards for reaching a specified location, successful obstacle avoidance rewards, or energy efficiency rewards, ensuring that the model completes its tasks efficiently while meeting stability and anti-slip requirements.

[0094] During training, the control strategy is iteratively and interactively optimized using a proximal policy optimization algorithm. In each round of training, the policy model inputs the current state (including robot data and environmental perception data), and the model outputs control commands for the drive motors (such as torque and rotation direction). The environment then calculates the next state based on the dynamics model and provides immediate rewards. The policy network parameters are updated through gradient ascent or descent to maximize the expected value of the cumulative reward. For example, in a cornering task, improper torque distribution may initially cause the vehicle to tilt or the wheels to slip. This significantly reduces the attitude stability reward and the slip suppression reward, forcing the policy network to adjust the motor output parameters. After several iterations, the model gradually learns an optimized strategy that balances these rewards, such as applying higher torque to the drive wheel on the inside of the curve to offset centrifugal force while reducing the speed of the outer wheel to suppress slip.

[0095] The training termination condition can be set to converge to a preset value for the composite reward function, meaning that the fluctuation range of the reward value over multiple consecutive iterations is less than the allowable error. At this point, the policy model can autonomously generate control instructions adapted to different pipeline scenarios based on the input state. For example, it can proactively reduce torque output and enhance displacement closed-loop control in oil-contaminated pipeline sections, or rapidly plan obstacle avoidance trajectories in the presence of dynamic obstacles.

[0096] See also Figure 4 , Figure 4 This is a diagram of the policy model training process.

[0097] In applications, SolidWorks software can be used to define the connection relationships between the pipeline robot's 3D connections and export the pipeline robot's 3D urdf (Uniform Robot Description) format file. Next, a motion library consisting of three rollers and two spherical wheels is created based on its motion characteristics, generating different motion control strategies. Reward functions such as forward and backward rewards, cornering rewards, and in-place rotation rewards are designed to meet different task requirements. The reward function R(t) is then obtained through weighted summation. The urdf file can then be imported into a reinforcement learning simulation platform, such as Isaac Gym, to build a training environment and define parameters such as the collision relationship and friction coefficient between the robot and the simulated environment. Policy parameters are then initialized, the current policy interacts with the environment, and trajectories are collected. The loss function is optimized using gradient ascent, and the policy iterations are repeated until convergence and termination criteria are met. After training, an optimized and mature policy model is obtained, which can autonomously generate optimal motion control commands based on the environment. In applications, the policy model can be deployed on a real-world pipeline robot through a sim-to-sim and then sim-to-real process.

[0098] See also Figure 5 , Figure 5 Schematic diagram of the training method based on reinforcement learning.

[0099] During reinforcement learning training, the policy model input includes robot proprioception information, specifically the angles, angular velocities, and motor torques of the five active degrees of freedom (DOFs), the angles and angular velocities of the seven passive degrees of freedom (DOFs), as well as the robot's posture (pitch, roll, and yaw, acquired by the IMU) and angular velocity. It also includes environmental information, such as the pipe's diameter, material, friction coefficient, type (straight, curved, etc.), and the three-dimensional conditions within the pipe (static and dynamic obstacles). This information is fed into the reinforcement learning model architecture, which utilizes Proximal Policy Optimization (PPO). This architecture comprises a hierarchical policy network and imitation learning pre-training (for accelerated convergence). The bottom layer is a proportional-differential (PD) control layer for motor motion, and the upper layer is a motion planning layer. This architecture runs on a simulation platform. The output of reinforcement learning training is motor control commands, including motion control strategies such as position control, torque control, and force-position hybrid control. Finally, the trained policy model is transferred to a real-world pipeline robot for virtual migration and algorithm deployment, enabling adaptive optimization control of the pipeline robot.

[0100] See also Figure 6 , Figure 6 Schematic diagram of training test scenarios and simulation test algorithm migration.

[0101] Figure 6The training and testing process for the pipeline robot is demonstrated in the following. The basic elements of the training and testing scenarios include various pipe types, such as straight pipes, reducers, tees, and elbows. Based on these elements, a simulation environment is built to simulate the pipeline robot's operation in various pipeline structures. After optimization through simulation testing, the pipeline robot is then put into real-world testing. Through repeated iterations, the reinforcement learning optimization strategy is continuously refined to improve the pipeline robot's maneuverability and environmental adaptability.

[0102] In one embodiment, the pipeline robot motion control method further includes:

[0103] In the simulation-to-physical migration stage, the trained policy model is calibrated for dynamic environmental parameters through a domain adaptation algorithm; the dynamic environmental parameters include at least sensor measurement error parameters.

[0104] In applications, during the simulation-to-physical migration phase, a domain adaptation algorithm can be used to dynamically calibrate the trained policy model's environmental parameters to eliminate discrepancies between the simulated and real-world environments. This algorithm dynamically adjusts the policy model's input data and decision logic by analyzing real-world deviations not accounted for in the simulation model, such as sensor measurement errors, mechanical transmission backlash, and environmental noise. Dynamic environmental parameters primarily include sensor measurement errors, such as the zero offset and sensitivity errors of the inertial measurement unit (IMU), the angular resolution deviation of the angle encoder, the nonlinear response characteristics of the torque sensor, and the ranging noise of the lidar or vision sensor. These errors are typically idealized in simulation, but in real-world scenarios they can significantly affect the accuracy of control command generation. Dynamic environmental parameters can also include parameters affecting ambient temperature and humidity, power supply fluctuations, and thermal effects on the drive motor.

[0105] Taking sensor measurement error parameters as an example, the calibration process can be divided into two phases: online data acquisition and parameter optimization. During the online data acquisition phase, the real-world pipeline robot performs a pre-set standard motion (such as uniform linear motion or fixed-point turning), synchronously recording the actual output data from both the internal and external sensor modules. This data is then compared with the expected data for the same motion in a simulated environment. For example, the deviation between the actual pitch angle data of the inertial measurement unit and the simulated value can be used to calibrate the zero offset, while the difference between the actual displacement increment of the angle encoder and the theoretical value can reflect its resolution error. By statistically analyzing the error distribution of multiple data sets, a sensor error correction model is established.

[0106] During the parameter optimization phase, the domain adaptation algorithm embeds the sensor error correction model into the input processing layer of the policy model, performing real-time corrections to the raw sensor data. Specifically, the angle encoder's measured angle is linearly fitted to determine the proportional coefficient and offset, mapping the raw measurement value to the corrected physical quantity. The torque sensor's output is filtered to suppress noise interference, and the nonlinear response curve is dynamically adjusted based on actual load feedback. Furthermore, the policy model incorporates a dynamic compensation mechanism when generating control commands. For example, it incorporates corrected actual torque data into the torque closed-loop control to offset deviations from the idealized model during simulation training.

[0107] The calibrated policy model significantly improves its adaptability to real-world conditions. For example, in low-friction pipe sections, torque commands generated based on idealized friction coefficients during simulation training may cause actual drive wheel slip. However, the calibrated model combines the corrected friction coefficient estimate with actual slip feedback to dynamically reduce torque output and increase the displacement closed-loop control weight, thereby suppressing slip.

[0108] Furthermore, an online continuous learning function can be configured. When a gradual change in environmental parameters is detected (e.g., a continuous decrease in the friction coefficient due to oil accumulation on the pipe wall), a parameter recalibration process is automatically triggered to ensure that the policy model maintains high robustness throughout the entire operation cycle. Ultimately, the dynamically calibrated policy model can generate control instructions that closely match real-world physical constraints, enabling a seamless transition from virtual training to real-world application, and ensuring the pipeline robot's motion stability and mission reliability in complex working conditions.

[0109] Corresponding to the aforementioned application function implementation method embodiment, the present invention also provides a pipeline robot motion control system and corresponding embodiments.

[0110] See Figure 7 , Figure 7 This is a schematic diagram of the modular structure of the pipeline robot motion control system. The pipeline robot motion control system is applied to the pipeline robot's processor. The pipeline robot includes an internal sensor module, an external sensor module, and multiple articulated joint units. Each joint unit includes one or more drive motors. The system includes:

[0111] The data acquisition unit 71 is used to obtain the status data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module in response to the direction command triggered by the operator through the remote terminal;

[0112] An instruction generation unit 72 is used to generate control instructions for each drive motor based on the direction instruction, state data and pipeline environment data using a pre-trained strategy model;

[0113] The motion control unit 73 is used to adjust the operating state of each drive motor according to the control instruction.

[0114] In one embodiment, the instruction generation unit 72 is further configured to:

[0115] When a communication interruption with the remote end is detected, the safe area in the pipeline is determined through the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each drive motor;

[0116] The motion control unit 73 is further configured to control each drive motor to execute a corresponding docking instruction, so that the pipeline robot moves to a safe area.

[0117] In one embodiment, the instruction generation unit 72 is further configured to:

[0118] When it is detected that the communication interruption duration exceeds a preset threshold, a driving path pointing to the target pipeline exit is determined based on the pipeline route map; the pipeline route map is pre-stored or generated in real time based on pipeline environmental data; the target pipeline exit is the nearest pipeline exit or a pre-designated pipeline exit;

[0119] Based on the status data, pipeline environment data and driving path, the strategy model generates the exit command of each drive motor;

[0120] The motion control unit 73 is further configured to control each drive motor to execute a corresponding exit instruction, so that the pipeline robot moves to the target pipeline outlet.

[0121] In one embodiment, the pipeline robot motion control system further includes:

[0122] The prompt unit is used to send a prompt message to the remote end when it is detected that there is a pending area in the driving direction based on the pipeline environmental data and no new direction instruction is received within a preset time; the pending area is a pipeline section with branch channels, turning channels and / or inaccessible areas.

[0123] In one embodiment, the instruction generation unit 72 is further configured to:

[0124] If no new direction command is received within a preset period of time after the prompt message is sent, the safe area in the pipeline is determined through the strategy model based on the pipeline environment data and status data, and the docking command for each drive motor is generated;

[0125] The motion control unit 73 is further configured to control each drive motor to execute a corresponding docking instruction, so that the pipeline robot moves to a safe area.

[0126] Regarding the system in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method. Relevant content about the policy model can also be found in the above content and will not be elaborated here.

[0127] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A pipeline robot motion control method, characterized in that: A processor for a pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of movable joint units; each of the joint units includes one or more drive motors; the method includes: In response to a direction command triggered by an operator via a remote terminal, the system acquires status data of the pipeline robot body collected by the internal sensor module and pipeline environment data collected by the external sensor module; the status data includes the rotation angle and angular velocity of each drive motor, the torque output state of each drive wheel and / or the drive motor; and the pipeline environment data includes the friction coefficient of the pipe wall. Based on the direction instruction, the state data and the pipeline environment data, a pre-trained strategy model is used to generate a control instruction for each of the drive motors; Adjusting the operating state of each of the drive motors according to the control instructions; the control instructions include target torque, rotation direction and force-position hybrid control parameters; When a communication interruption with the remote end is detected, a safe area in the pipeline is determined by the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each of the drive motors; Controlling each of the driving motors to execute a corresponding docking instruction to move the pipeline robot to the safe area; When it is detected that the duration of the communication interruption exceeds a preset threshold, a driving path pointing to a target pipeline exit is determined based on a pipeline route map; the pipeline route map is pre-stored or generated in real time based on the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-designated pipeline exit; Based on the state data, pipeline environment data and driving path, generating an exit instruction for each of the drive motors through the strategy model; Each driving motor is controlled to execute a corresponding exit instruction, so that the pipeline robot moves to the target pipeline outlet.

2. The pipeline robot motion control method according to claim 1, characterized in that: The method further comprises: When it is detected based on the pipeline environmental data that there is a pending area in the driving direction and no new direction instruction is received within a preset time, a prompt message is sent to the remote end; the pending area is a pipeline section with branch channels, turning channels and / or inaccessible areas.

3. The pipeline robot motion control method according to claim 2, characterized in that: The method further comprises: If no new direction instruction is received within a preset period of time after the prompt message is sent, a safe area in the pipeline is determined by the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each of the drive motors; Control each of the driving motors to execute a corresponding docking instruction, so that the pipeline robot moves to the safe area.

4. The pipeline robot motion control method according to claim 1, characterized in that: The training method of the strategy model includes: Constructing a dynamic simulation model of the pipeline robot and a corresponding pipeline simulation training environment; According to the preset motion task objectives, a composite reward function is constructed that integrates posture stability reward, slippage suppression reward, and task execution reward. In the pipeline simulation training environment, the control strategy of the drive motor in the dynamic simulation model is iteratively updated through a proximal strategy optimization algorithm. When the composite reward function converges to a preset value, the trained strategy model is output.

5. The pipeline robot motion control method according to claim 4, characterized in that: The method further comprises: In the simulation-to-physical migration stage, the trained policy model is calibrated for dynamic environmental parameters using a domain adaptation algorithm; the dynamic environmental parameters include at least sensor measurement error parameters.

6. A pipeline robot motion control system, characterized in that: A processor for a pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of movable joint units; each of the joint units includes one or more drive motors; the system includes: a data acquisition unit, configured to acquire, in response to a direction command triggered by an operator via a remote terminal, state data of the pipeline robot body acquired by the internal sensor module and pipeline environment data acquired by the external sensor module; the state data including the rotation angle and angular velocity of each of the drive motors, the torque output state of each drive wheel and / or the drive motor; and the pipeline environment data including the pipe wall friction coefficient; An instruction generation unit, configured to generate control instructions for each of the drive motors based on the direction instruction, the state data, and the pipeline environment data using a pre-trained strategy model; the control instructions include a target torque, a rotation direction, and a force-position hybrid control parameter; A motion control unit, configured to adjust the operating state of each of the drive motors according to the control instructions; The instruction generation unit is further configured to: When a communication interruption with the remote end is detected, a safe area in the pipeline is determined by the strategy model based on the pipeline environment data and status data, and a docking instruction is generated for each of the drive motors; The motion control unit is also used for: Controlling each of the driving motors to execute a corresponding docking instruction to move the pipeline robot to the safe area; The instruction generation unit is further configured to: When it is detected that the duration of the communication interruption exceeds a preset threshold, a driving path pointing to a target pipeline exit is determined based on a pipeline route map; the pipeline route map is pre-stored or generated in real time based on the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-designated pipeline exit; Based on the state data, pipeline environment data and driving path, generating an exit instruction for each of the drive motors through the strategy model; The motion control unit is also used for: Each driving motor is controlled to execute a corresponding exit instruction, so that the pipeline robot moves to the target pipeline outlet.

Citation Information

Patent Citations

  • Underwater robot automatic course reversal control method, computer and storage medium

    CN107957729A

  • Control method of long-distance water supply pipeline detection robot

    CN114110303A

  • Robot motion control method and system and electronic equipment

    CN116619382A