Pipeline robot motion control method and system

By a method of obtaining the state and environment data of pipeline robots in real time and generating control instructions using the strategy model, the problems of inflexible movement of pipeline robots and wireless signal interference in the prior art are solved, and higher motion stability and pass efficiency are achieved.

CN120170753AActive Publication Date: 2025-06-20PEKING UNIV

Patent Information

Application Number
CN202510647478.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing pipeline robots are not flexible in complex pipeline environments, and are prone to collision or lag due to delays or errors in manual operation, and wireless signal interference and transmission delay problems seriously affect the reliability and efficiency of motion control.

Method used

A pipeline robot motion control method is adopted to obtain robot status and environment data in real time through internal and external sensor modules, and combine pre-trained strategy models to generate control instructions for driving the motor to realize autonomous motion control. This method can determine the safe area and dock when the communication is interrupted, and can independently lead to the target pipeline exit when the communication is interrupted exceeds the threshold.

Benefits of technology

It significantly reduces the dependence on manual remote control operation, improves the motion stability, load capacity and throughput efficiency of the robot in complex pipeline environments, reduces mechanical wear and extends the service life of the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120170753A_ABST
    Figure CN120170753A_ABST
Patent Text Reader

Abstract

The invention discloses a pipeline robot motion control method and system, and relates to the technical field of robot control. The method comprises the steps that in response to a direction instruction triggered by an operator through a remote end, state data, collected by an internal sensor module, of a pipeline robot body and pipeline environment data collected by an external sensor module are obtained; based on the direction instruction, the state data and the pipeline environment data, utilizing a pre-trained strategy model to generate a control instruction of each driving motor; and adjusting the running state of each driving motor according to the control instruction. Autonomous motion control over the pipeline robot can be achieved in the complex pipeline environment, manual dependence is reduced, and the motion stability, the load capacity and the passing efficiency of the robot are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot control, and particularly to a method and system for controlling the movement of a pipeline robot. Background Art

[0002] At present, in practical applications, pipeline robots mainly rely on manual remote control operations, and the operator needs to manually adjust the rotation direction and speed of each motor. However, in a complex pipe network environment, such as when there are dynamic obstacles inside the pipeline or rapid obstacle avoidance is required, manual operations often lead to inflexible robot movement due to reaction delays or operation errors, and even collisions or jams may occur, severely limiting its passing efficiency and safety. Especially in emergency situations, the real-time decision-making ability of the operator is difficult to meet the high-precision control requirements.

[0003] Existing cable-free pipeline robots usually use wireless signal transmission to achieve remote control operations, but in practical applications, they often face problems of signal interference and transmission delays. When the robot is in a complex pipeline structure (such as a bent pipe or a variable-diameter pipe) or needs to execute multi-instruction collaborative actions, the instability of the wireless signal easily causes packet loss or response lag of the instructions, making it difficult to achieve refined motion control. For example, during the obstacle avoidance process, if a key instruction fails to be conveyed in time, the robot may get stuck due to inconsistent actions, and even damage the pipeline or its own structure, seriously affecting the operation reliability.

[0004] In addition, the internal environment of the pipeline is complex and changeable, and the inner wall is often attached with impurities such as oil stains and condensate water, and the friction coefficients in different areas vary significantly due to differences in materials, temperature, and humidity. Existing manual control methods are difficult to perceive the environmental state in real time and dynamically adjust the driving strategy, resulting in unreasonable torque distribution of the driving wheels and easy occurrence of wheel slip. This not only reduces the longitudinal driving force and load capacity of the robot, but also aggravates mechanical wear and shortens the service life of the equipment. Especially in long-distance or high-load scenarios, the existing technology is difficult to balance the motion efficiency and stability, restricting the wide application of pipeline robots in industrial inspection, maintenance and other fields. Summary of the Invention

[0005] To solve the problems existing in the prior art, the present invention provides a method and system for controlling the movement of a pipeline robot, which can realize the autonomous movement control of the pipeline robot in a complex pipeline environment, reduce the dependence on manual labor, and improve the movement stability, load capacity and passing efficiency of the robot.

[0006] To achieve the above object, the present invention provides a method for controlling the movement of a pipeline robot, which is applied to the processor of the pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of articulation units connected movably; each of the articulation units includes one or more driving motors; the method includes: In response to a direction instruction triggered by an operator through a remote terminal, obtain the status data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module; Based on the direction instruction, status data, and pipeline environment data, use a pre-trained policy model to generate control instructions for each of the drive motors; According to the control instructions, adjust the operating states of each of the drive motors.

[0007] Optionally, the method further includes: When a communication interruption with the remote terminal is detected, based on the pipeline environment data and status data, determine a safe area in the pipeline through the policy model, and generate docking instructions for each of the drive motors; Control each of the drive motors to execute the corresponding docking instruction, so that the pipeline robot moves to the safe area.

[0008] Optionally, the method further includes: When it is detected that the communication interruption duration exceeds a preset threshold, according to the pipeline route map, determine a driving path pointing to the target pipeline exit; the pipeline route map is pre-stored or is generated in real time according to the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-specified pipeline exit; Based on the status data, pipeline environment data, and driving path, generate exit instructions for each of the drive motors through the policy model; Control each drive motor to execute the corresponding exit instruction, so that the pipeline robot moves to the target pipeline exit.

[0009] Optionally, the method further includes: When it is detected according to the pipeline environment data that there is a pending area in the driving direction and no new direction instruction is received within a preset time, send a prompt message to the remote terminal; the pending area is a pipeline section with branch channels, turning channels, and / or impassable areas.

[0010] Optionally, the method further includes: If no new direction instruction is received within a preset period after sending the prompt message, then based on the pipeline environment data and status data, determine a safe area in the pipeline through the policy model, and generate docking instructions for each of the drive motors; Control each of the drive motors to execute the corresponding docking instruction, so that the pipeline robot moves to the safe area.

[0011] Optionally, the training method of the policy model includes: Construct a dynamic simulation model of the pipeline robot and a corresponding pipeline simulation training environment; Construct a composite reward function that integrates attitude stability reward, slip suppression reward, and task execution reward according to a preset motion task objective; In the pipeline simulation training environment, iteratively update the control strategy of the drive motors in the dynamic simulation model through the proximal policy optimization algorithm. When the composite reward function converges to a preset value, output the trained policy model.

[0012] Optionally, the method further includes: In the simulation-to-physical migration stage, calibrate the dynamic environment parameters of the trained policy model through a domain adaptation algorithm; the dynamic environment parameters at least include sensor measurement error parameters.

[0013] The present invention also provides a pipeline robot motion control system, which is applied to the processor of the pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of articulating joint units; each of the joint units includes one or more drive motors; the system includes: A data acquisition unit, configured to obtain the state data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module in response to a direction instruction triggered by an operator through a remote terminal; An instruction generation unit, configured to generate control instructions for each of the drive motors based on the direction instruction, the state data, and the pipeline environment data by using a pre-trained policy model; A motion control unit, configured to adjust the operating states of the drive motors according to the control instructions.

[0014] According to the specific embodiments provided by the present invention, the following technical effects are disclosed: The pipeline robot motion control method provided by the present invention can significantly reduce the dependence on manual remote control operations by obtaining pipeline environment information, robot body attitude information, and drive motor operation information in real time and inputting them into a pre-trained policy model. In a complex pipeline network environment, the robot can autonomously adjust the torque of the drive wheels according to the pipeline environment (such as obstacle distribution, pipeline structure changes, and differences in wall friction coefficients), effectively avoid wheel slippage, improve longitudinal driving force and load capacity, while reducing mechanical wear and extending the service life of the equipment.

[0015] During manual operation, the operator only needs to issue simple motion direction instructions (such as forward, backward, or turning), and the strategy model can generate refined motor control instructions by combining the input data. This mechanism significantly reduces the dependence on high-frequency and multi-dimensional remote control instructions, thereby reducing the amount of signal transmission data, alleviating the problem of instruction packet loss caused by wireless signal interference or delay, and enhancing the motion coherence and reliability of the robot in complex pipeline structures (such as elbow pipes and variable-diameter pipes).

[0016] In addition, based on the autonomous decision-making ability of the strategy model, the robot can quickly respond to dynamic obstacles and sudden working conditions, achieve flexible obstacle avoidance and stable cornering, overcome problems such as collisions and jams caused by reaction delays or operation errors in traditional manual control, and significantly improve the operation efficiency and safety. Brief Description of the Drawings

[0017] The above and other objects, features, and advantages of the present invention will become more apparent by describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, wherein, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.

[0018] Figure 1 It is a schematic flowchart of the method for controlling the movement of the pipeline robot shown in the embodiments of the present invention; Figure 2 It is a schematic diagram of active control during the movement of the pipeline robot shown in the embodiments of the present invention; Figure 3 It is a schematic diagram of the drive structure of the pipeline robot shown in the embodiments of the present invention; Figure 4 It is a schematic flowchart of the training process of the strategy model shown in the embodiments of the present invention; Figure 5 It is a schematic diagram of the principle of the training method based on reinforcement learning shown in the embodiments of the present invention; Figure 6 It is a schematic diagram of the migration of the training test scenario and the simulation test algorithm shown in the embodiments of the present invention; Figure 7 It is a schematic diagram of the module structure of the pipeline robot motion control system shown in the embodiments of the present invention. Detailed Description of the Embodiments

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] The object of the present invention is to provide a method and system for motion control of a pipeline robot, which can realize autonomous closed-loop motion control based on the fusion of a policy model and multi-modal sensors in a complex pipeline environment, reduce the dependence on manual remote control operations, and improve the motion stability, load capacity and passing efficiency of the robot in scenarios of dynamic obstacles, sudden changes in friction coefficient and signal interference.

[0021] Please refer to Figure 1 , Figure 1 which is a schematic flow diagram of the method for motion control of a pipeline robot.

[0022] The above-mentioned method for motion control of a pipeline robot is applied to the processor of the pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of articulation units that are movably connected; each articulation unit includes one or more drive motors; the method includes: Step 101: In response to a direction instruction triggered by an operator through a remote end, obtain the state data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module.

[0023] In the application, after the operator sends a direction instruction through a remote end (such as the cloud, host computer), the pipeline robot enters the instruction response stage. The remote end serves as a human-computer interaction interface and receives direction instructions such as forward, backward, pause or turn input by the operator. After receiving the direction instruction, the pipeline robot obtains the data collected by the internal sensor module and the external sensor module.

[0024] Exemplarily, the internal sensor module may include an inertial measurement unit, an angle encoder and a torque sensor; among them, the inertial measurement unit real-time monitors the pitch angle, roll angle and yaw angle of the robot torso, the angle encoder records the rotation angle and angular velocity of each drive motor, and the torque sensor feeds back the torque output state of each drive wheel and / or drive motor. These data together constitute the state data of the pipeline robot body, which is used to characterize the current posture, motion state of the robot and the real-time load condition of the power system.

[0025] At the same time, the external sensor module activates the environment perception function and obtains pipeline environment data through multi-modal sensors. For example, it may include a lidar or a vision sensor, which is used to scan the three-dimensional geometric features of the inner wall of the pipeline (such as diameter change, elbow curvature, position of the reduced diameter section), a proximity sensor to detect the distance and azimuth of obstacles, and a friction coefficient sensor to analyze the influence of the pipe wall material and surface attachments (such as oil stains, condensate) on the motion resistance. The pipeline environment data not only covers static structure information, but also may include the real-time position and motion trajectory of dynamic obstacles, ensuring that the robot can comprehensively perceive the current complex working conditions.

[0026] The status data and environmental data are transmitted to the processor through the communication module for integration. The processor can preprocess these raw data, such as noise filtering, coordinate alignment, and data fusion. In addition, during the data acquisition process, the effectiveness of the sensors can be verified in real time. If a sensor anomaly is detected, a redundant data switching mechanism is triggered to ensure the continuity and reliability of information collection.

[0027] See Figure 2 and Figure 3 , Figure 2 is a schematic diagram of the active control of the pipeline robot during pipeline movement, Figure 3 is a schematic diagram of the drive structure of the pipeline robot. Among them, and represent active drives, , and represent transmission drives. Figure 2 The labels 1 and 2 in

[0028] In the application, the operator can remotely control the pipeline robot through the remote end. The pipeline robot has 5 active degrees of freedom and 7 passive degrees of freedom, and is an intelligent robot with 12 degrees of freedom as a whole. In addition, in local movement, since omnidirectional wheels are used, each omnidirectional wheel is equipped with 12 small rollers, and there are 6 omnidirectional wheels in total, with 72 local degrees of freedom. Regarding the drive structure of the pipeline robot, the front and rear ball wheels are driven by the drive motors of the first joint unit and the fourth joint unit of the pipeline robot. The middle rollers are driven by the drive motors of the second joint unit and the third joint unit of the pipeline robot through bevel gear transmission. Each drive motor is equipped with an angle encoder and a torque sensor, and each section is equipped with an Inertial Measurement Unit (IMU) that can sense the body attitude angles (pitch angle, roll angle, and yaw angle) of each section of the torso. In addition, angle sensors are installed between the torso and the torso. Between the first and the second sections, and between the third and the fourth sections, in addition to being connected by an axis, they are also connected by a tension spring. The front and rear ball wheels can rotate independently around the cross axis and have four rotational degrees of freedom. The second and third torso sections are connected by an axis and have one passive rotational degree of freedom.

[0029] Step 102: Based on the direction instruction, status data, and pipeline environment data, use the pre-trained policy model to generate control instructions for each drive motor.

[0030] In the application, the direction instruction, status data, and pipeline environment data can be used as inputs and transmitted to the pre-trained policy model. The direction instruction, as a high-level control target, indicates the overall movement intention of the robot, such as moving forward, backward, or turning.

[0031] The policy model can be optimized through reinforcement learning training. At the input end of the model, the direction instruction, state data, and pipeline environment data are multi-modally fused, and after preprocessing, a unified state representation vector is formed. This vector is input into the policy network, which, based on the pre-learned dynamic relationship and environmental interaction experience, analyzes the optimal motion policy under the current working conditions. Specifically, the model dynamically balances the torque distribution and rotation direction of each drive motor by calculating composite reward functions such as attitude stability, slip suppression, and task execution, ensuring that in complex pipeline scenarios, it can not only achieve the motion direction required by the direction instruction but also actively suppress wheel slippage and enhance the longitudinal driving force. For example, in a cornering working condition, the model will dynamically adjust the torque difference between the inner and outer drive wheels according to the pipe wall friction coefficient and the robot's attitude angle, avoiding side slip or understeering caused by uneven torque distribution.

[0032] The output of the model is the refined control instructions for each drive motor, which can include target torque, rotation direction, and force-position hybrid control parameters. These control instructions can be converted into pulse width modulation signals through the underlying motor controller to drive the actuators of each joint unit. For example, in response to a decrease in the friction coefficient caused by oil stains on the inner wall of the pipeline, the model will reduce the torque output of the drive wheels to suppress slippage, and at the same time, ensure that the contact force between the wheels and the pipe wall is maintained within a reasonable range through force-position hybrid control, thereby improving the motion stability.

[0033] Among them, the target torque is used to adjust the output torque of the drive motor, and the longitudinal driving force when the drive wheel contacts the pipe wall can be set by adjusting the current or voltage signal; the rotation direction parameter determines the forward and reverse states of the motor rotor by changing the winding power-on sequence or control logic of the drive motor, thereby determining the overall motion direction of the robot (forward, backward, or turning); the force-position hybrid control parameter combines the real-time displacement feedback of the angle encoder and the dynamic force feedback of the torque sensor by coordinating the position closed-loop and torque closed-loop control of the drive motor, achieving double constraints on the displacement accuracy and contact force stability of the drive wheel under complex friction conditions. For example, when the pipe wall friction coefficient suddenly changes, the force-position hybrid control can dynamically adjust the output characteristics of the drive motor, ensuring that the wheel moves along the preset trajectory while maintaining the wheel-wall contact force within the safety threshold, avoiding slippage caused by insufficient contact force or mechanical damage caused by contact force overload.

[0034] In addition, the policy model can be calibrated through a domain adaptation algorithm during the simulation-to-physical migration stage, dynamically compensating for sensor measurement errors and real environment parameter deviations to ensure the reliability and environmental adaptability of the control instructions. Finally, the generated control instructions are transmitted to each drive motor through the bus, realizing the autonomous closed-loop motion control of the pipeline robot under complex working conditions.

[0035] Step 103: Adjust the operating states of each drive motor according to the control instruction.

[0036] During the execution of the instruction, the control instruction is transmitted to the underlying controller of each drive motor through the bus. The controller parses the instruction into specific pulse-width modulation signals or current control signals, and the drive motor executes the corresponding actions. For example, for a certain drive wheel, if the instruction requires increasing the torque to overcome the friction reduction caused by the oil stain on the pipe wall, the controller will increase the duty cycle of the input current to increase the output torque of the motor; if it is necessary to adjust the rotation direction to achieve steering, the rotation direction of the rotor is changed by switching the energization sequence of the motor windings. At the same time, the force-position hybrid control mode dynamically adjusts the output characteristics of the motor by real-time collecting the position data feedback by the angle encoder and the force feedback data of the torque sensor, ensuring the balance between the target position and the contact force, and avoiding wheel slippage or movement stagnation caused by overloading or underloading.

[0037] During the adjustment process, the actual operating states of each drive motor can be continuously monitored, including real-time torque, speed, and position deviation. These data are fed back to the processor through the internal sensor module to form a closed-loop control loop. For example, if there is a deviation between the actual torque of a certain drive wheel and the target value, a dynamic compensation mechanism will be triggered, and the control signal will be fine-tuned by combining the attitude data of the inertial measurement unit and the change in the friction coefficient sensed by the environmental sensor to eliminate the error. Through the above mechanism, the operating states of each drive motor can be precisely regulated, and ultimately the stable movement of the pipeline robot under complex working conditions can be achieved.

[0038] In one embodiment, the above pipeline robot motion control method further includes: When a communication interruption with the remote end is detected, based on the pipeline environment data and status data, determine the safe area in the pipeline through the policy model, and generate the docking instructions for each drive motor; Control each drive motor to execute the corresponding docking instruction to move the pipeline robot to the safe area.

[0039] In the application, when a communication connection interruption with the remote end is detected, the autonomous emergency response mechanism can be immediately started. The communication interruption may be caused by signal interference, transmission equipment failure, or environmental occlusion. At this time, the remote end cannot continue to send direction instructions, and the robot needs to rely on the local policy model to make autonomous decisions. The pipeline environment data is continuously collected by the external sensor module, and the status data is provided by the internal sensor module.

[0040] The policy model evaluates the safe areas in the current pipeline environment through pre-trained reinforcement learning policies. The safe areas need to meet the following conditions: being in a straight pipe section or a low-curvature elbow section to avoid the risk of structural instability; being far from dynamic obstacles or static obstacle aggregation areas to reduce the collision probability; having a relatively high pipe wall friction coefficient and few surface attachments to ensure effective grip of the drive wheels. In addition, the model can also calculate the priority and reachability of the safe areas through a multi-objective optimization algorithm, combining the prior knowledge of the pipeline route map (such as pipeline branch layout, exit position) with real-time perception data, and finally select the optimal target position.

[0041] When generating the docking instruction, the policy model plans an obstacle avoidance path according to the position of the safe area and the current state of the robot, and decomposes it into the control parameters of each drive motor. For example, if the safe area is located in the rear straight pipe section, the model will calculate the torque distribution and rotation direction combination of the drive wheels to make the robot retreat with the minimum turning radius, and at the same time adjust the force-position hybrid control parameters of the middle rollers to balance the body attitude. The docking instruction includes the target torque, rotation direction and displacement increment of each motor to ensure that the robot maintains longitudinal driving force and lateral stability during movement. During the execution stage, the underlying controller converts the instruction into a motor drive signal, and precisely controls the movement of each drive wheel by dynamically adjusting the pulse width modulation duty cycle or switching the power-on sequence of the motor windings.

[0042] During the movement, the changes in the environment and the state of the robot body can be continuously monitored. If it is detected that a dynamic obstacle invades the planned path or the pipe wall friction coefficient suddenly changes, the policy model will update the evaluation result of the safe area in real time and regenerate the docking instruction. For example, when a new obstacle appears in the retreat path, the model may switch to the side elbow section as a temporary safe area and achieve lateral translation by adjusting the torque difference between the ball wheels and the rollers. All actions are realized through closed-loop control. The internal sensors feedback the motor execution state and the body attitude in real time, and the external sensors update the environmental data, forming an autonomous cycle of "perception - decision - execution - verification" until the robot successfully reaches the safe area and enters the standby state, waiting for the communication to resume or subsequent instruction triggers.

[0043] In one embodiment, the above pipeline robot motion control method further includes: When it is detected that the communication interruption duration exceeds the preset threshold, determine the driving path pointing to the target pipeline exit according to the pipeline route map; the pipeline route map is pre-stored or is generated in real time according to the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-specified pipeline exit; Based on the state data, pipeline environment data and driving path, generate the exit instructions for each drive motor through the policy model; Control each drive motor to execute the corresponding exit instruction to make the pipeline robot move to the target pipeline exit.

[0044] In an application, when it is detected that the communication interruption duration exceeds a preset threshold, it can be determined that the remote end cannot resume connection in the short term, and then the autonomous exit mechanism is activated. The preset threshold can be flexibly set according to the risk level of the pipeline operation scenario. For example, a shorter threshold is set in flammable, explosive or high-humidity environments to give priority to ensuring equipment safety. The pipeline route map, as the basis for path planning, is divided into two types: one is a static map pre-constructed through 3D modeling or historical inspection data, which contains pipeline branches, outlet coordinates and key node information; the other is a topological map dynamically generated based on real-time pipeline environment data, which realizes simultaneous localization and mapping (SLAM) through the fusion of lidar, visual sensors and odometers, and updates the pipeline structure features and obstacle distribution in real time. Among them, the generation of the pipeline route map can be generated by a policy model or by other functional modules such as a route construction module. The selection of the target pipeline outlet follows the priority rule. If there are multiple outlets, the nearest outlet is preferentially selected to shorten the exit path; if a specific outlet (such as a maintenance port or a safety cabin) is preset, then this outlet is taken as the final target.

[0045] The policy model generates an exit instruction based on the current robot's state data (including drive motor load, body attitude and remaining battery power), pipeline environment data (such as obstacle position, pipe wall friction coefficient) and geometric constraints of the driving path. The driving path needs to meet the requirements of obstacle avoidance, low energy consumption and motion stability: the model plans the globally optimal path through a path search algorithm (such as A* or RRT), and dynamically adjusts the local obstacle avoidance trajectory in combination with real-time sensor data. For example, when there are dynamic obstacles in the path, the model will calculate the detour trajectory and decompose it into torque distribution and rotation direction parameters of each drive motor to ensure that the robot avoids obstacles with the minimum turning radius. The exit instruction specifically includes the target torque gradient of the drive motor (such as slow acceleration to avoid wheel slippage), the segmented rotation direction sequence (such as turning left first to adjust the heading and then going straight), and the force-position hybrid control parameters (such as reducing the torque and enhancing the displacement closed-loop accuracy in smooth pipe sections).

[0046] In the execution stage, the drive motors adjust their operating states according to the exit instruction. For example, in a narrow elbow section, the torque of the middle roller is restricted to prevent excessive extrusion of the pipe wall, while the front and rear spherical wheels achieve precise steering through differential rotation; in a straight pipe section, all drive motors synchronously increase the torque to accelerate the movement. In addition, the execution effect can be monitored in real time through internal sensors. If it is detected that the actual movement trajectory deviates from the planned path (such as wheel slippage caused by pipe wall deformation), the policy model will recalculate the path and update the instruction. At the same time, the sensor error (such as the cumulative deviation of the odometer) can also be compensated online through a domain adaptation algorithm to ensure the positioning accuracy in the real environment.

[0047] In one embodiment, the above-mentioned pipeline robot motion control method further includes: When a pending area is detected in the driving direction based on pipeline environment data and no new direction instruction is received within the preset time, a prompt message is sent to the remote end; the pending area is a pipeline section with branch channels, turning channels, and / or impassable areas.

[0048] During the process of the pipeline robot driving along the preset path, the external sensor module continuously collects pipeline environment data, including the geometric structure scanned by the lidar, the pipeline branch features captured by the vision sensor, and the obstacle distribution detected by the proximity sensor. When a pending area is detected ahead in the driving direction, the path decision interruption mechanism is triggered. The pending area refers to a pipeline section where the passing path needs to be determined, such as the intersection of branch channels (where the left, right, or upward branch needs to be selected), a sharp turning channel (with a curvature exceeding the preset safety threshold), or an impassable area with collapses, foreign object blockages, etc. The determination basis for such areas includes sudden changes in pipe diameter, excessive obstacle density, or newly added structures not marked in the topological map.

[0049] In the application, the position, type, and environmental parameters of the pending area (such as the azimuth angle of the branch channel, the turning radius, or the size of the blockage) can be integrated into the prompt message and sent to the operation interface of the remote end through the wireless communication module. The prompt message can include the three-dimensional coordinates of the pending area with highlighted markings, the recommended passing options (such as a list of feasible directions for the branch channel), and the environmental risk assessment results (such as a warning of insufficient friction coefficient for the turning channel).

[0050] After sending the prompt message, a command waiting timer with a preset duration can be started. If no direction instruction returned by the remote end is received within the preset time, it is determined that the operator cannot respond in time, and then the safety decision mode is switched. The policy model can re-evaluate the passability of the pending area based on the real-time pipeline environment data and the robot body state data. For example, for a branch channel, the policy model can select the path with the highest probability according to the connectivity analysis of the pipeline route map; for an impassable area, a detour trajectory can be generated by combining historical obstacle avoidance strategies; if the autonomous decision is still not feasible, the pending area can be marked as a special area, and the reverse backtracking mechanism can be triggered to control the robot to retreat along the original path to the nearest safe node, while continuously sending the environmental state update to the remote end until the communication is restored or manual intervention completes the path replanning.

[0051] Throughout the process, the internal sensor module synchronously monitors the real-time status of the robot, including the load of the drive motors, the attitude stability of the fuselage, and the energy consumption rate, ensuring that no secondary failures are caused by power overload or attitude imbalance during the waiting or reverse movement phases. For example, when pausing in front of a sharp-turn passage, the policy model dynamically adjusts the torque distribution between the ball wheels and the rollers to counteract the side-tilting moment of the fuselage caused by the inclined pipe wall. In addition, the generation and transmission of prompt messages can both adopt redundant communication protocols to ensure that information interaction can still be completed through multi-band transmission or data compression technology in a weak signal environment, avoiding decision-making delays caused by communication packet loss.

[0052] In one embodiment, the above pipeline robot motion control method further includes: If no new direction instruction is received within a preset time period after sending the prompt message, based on the pipeline environment data and status data, determine the safe areas in the pipeline through the policy model, and generate docking instructions for each drive motor; Control each drive motor to execute the corresponding docking instruction to move the pipeline robot to the safe area.

[0053] Among them, the determination of the safe area, the generation and execution of the docking instruction, etc. can refer to the relevant introduction in the foregoing embodiments, and will not be elaborated here.

[0054] In one embodiment, the training method of the policy model includes: Construct a dynamic simulation model of the pipeline robot and the corresponding pipeline simulation training environment; According to the preset motion task objective, construct a composite reward function that combines attitude stability reward, slip suppression reward, and task execution reward; In the pipeline simulation training environment, iteratively update the control strategy of the drive motors in the dynamic simulation model through the proximal policy optimization algorithm. When the composite reward function converges to a preset value, output the trained policy model.

[0055] In application, a dynamic simulation model of the pipeline robot can be constructed. This model is based on the physical structure parameters of the robot (such as mass distribution, joint connection stiffness, drive motor characteristics) and kinematic constraints (such as the contact geometric relationship of omnidirectional wheels, the rotational degrees of freedom of ball wheels), and accurately reproduces the motion response characteristics of the real robot through multi-body dynamics software. The pipeline simulation training environment simulates diverse pipeline scenarios, including straight pipes, bent pipes, variable-diameter pipes, tee pipes, and dynamic obstacle distributions, and sets different environmental parameters such as pipe wall friction coefficients, temperature, and humidity to cover the complex working conditions that may be encountered in actual operations.

[0056] In a simulation environment, the training objective can be quantified by a composite reward function. Exemplarily, the attitude stability reward term is based on the pitch angle, roll angle, and angular velocity data fed back by an inertial measurement unit, calculates the penalty value for the deviation of the robot's attitude from the equilibrium state, and encourages the model to maintain the stability of the fuselage during movement; the slip suppression reward term evaluates the degree of slippage through the difference between the actual linear velocity and the theoretical linear velocity of the driving wheels, and dynamically adjusts the penalty weight in combination with the pipe wall friction coefficient to optimize the torque distribution strategy; the task execution reward term is directly related to the preset motion target, such as the progress reward for reaching a specified position, the reward for successful obstacle avoidance, or the energy consumption efficiency reward, to ensure that the model efficiently completes the task on the premise of meeting stability and anti-slip requirements.

[0057] During the training process, the proximal policy optimization algorithm is used to iteratively interact and optimize the control strategy. In each round of training, the current state (including the robot's body data and environmental perception data) is input into the policy model, and the model outputs the control commands for the driving motors (such as torque magnitude and rotation direction). Subsequently, the environment calculates the next state according to the dynamic model and feedbacks the immediate reward. The parameters of the policy network are updated through gradient ascent or descent to maximize the expected value of the cumulative reward. For example, in a cornering task, in the initial stage of the model, improper torque distribution may cause the fuselage to tilt or the wheels to slip. At this time, the attitude stability reward and the slip suppression reward are significantly reduced, forcing the policy network to adjust the motor output parameters; after several iterations, the model gradually learns the optimization strategy for balancing various reward terms, such as applying higher torque to the driving wheels on the inner side of the curve to counteract the centrifugal force, and at the same time reducing the rotation speed of the outer wheels to suppress slippage.

[0058] The training termination condition can be set as the convergence of the composite reward function to a preset value, that is, the fluctuation range of the reward value in consecutive multiple iterations is less than the allowable error. At this time, the policy model can autonomously generate control commands adapted to different pipeline scenarios according to the input state, such as actively reducing the torque output and enhancing the displacement closed-loop control in the oil-contaminated pipe section, or quickly planning an obstacle avoidance trajectory under the interference of dynamic obstacles.

[0059] See Figure 4 , Figure 4 for the schematic diagram of the training process of the policy model.

[0060] In the application, according to the three-dimensional connection method of the pipeline robot, the connection relationship between each connection module can be defined by SolidWorks software and a three-dimensional urdf (Unified Robot Description) format file of the pipeline robot can be exported. Then, based on its motion characteristics, a motion library containing 3 rollers and 2 ball wheels is established to generate different motion control strategies. Then, for different task requirements, reward functions such as forward and backward rewards, cornering rewards, and in-place rotation rewards are designed, and the reward function R(t) is obtained through weighted summation. After that, the urdf file can be imported into a reinforcement learning simulation platform such as Issac gym to construct a training environment, and parameters such as the collision relationship and friction coefficient between the robot and the simulation environment are defined. Then, the policy parameters are initialized, the current policy interacts with the environment and collects trajectories, and the loss function is optimized using the gradient ascent method. The policy iteration is continuously repeated until convergence meets the termination conditions. After the training is completed, an optimized and mature policy model is obtained, and this policy model can autonomously generate optimal motion control instructions according to the environmental state. In the application, sim to sim (simulation to simulation) can be carried out first, and then sim to real (simulation to real), and finally the deployment of the policy model of the real pipeline robot is completed.

[0061] See Figure 5 , Figure 5 is a schematic diagram of the principle of the training method based on reinforcement learning.

[0062] During the process of reinforcement learning training, the input part of the policy model includes the perception information of the robot body, specifically the angles, angular velocities, and motor torques of 5 active degrees of freedom, the angles and angular velocities of 7 passive degrees of freedom, as well as the robot's attitude (pitch angle, roll angle, and yaw angle, obtained by the IMU), angular velocity, etc., and environmental perception information, such as the diameter, material, friction coefficient, type (straight pipe, elbow pipe, etc.) of the pipeline and the three-dimensional situation inside the pipeline (static and dynamic obstacles). These information are input into the reinforcement learning model architecture, which adopts the Proximal Policy Optimization (PPO) strategy, including a hierarchical policy network and imitation learning pre-training (to accelerate convergence). The bottom layer is the PD (Proportional-Differential) control layer of motor motion, and the upper layer is the motion planning layer, and it runs based on the simulation platform. The output after reinforcement learning training is the control instructions of the motor, including motion control strategies such as position control, torque control, and force-position hybrid control. Finally, the trained policy model is migrated to the real pipeline robot for virtual migration and algorithm deployment to achieve adaptive optimization control of the pipeline robot.

[0063] See Figure 6 , Figure 6 is a schematic diagram of the training test scenario and the migration of the simulation test algorithm.

[0064] Figure 6The process of training and testing the pipeline robot is shown. The basic elements of the training and testing scenario include various pipeline types such as straight pipes, reducing pipes, tee pipes, and elbow pipes. Based on these elements, a simulation environment is built to simulate the operation of the pipeline robot in different pipeline structures. After being tested and optimized in the simulation environment, the pipeline robot is then put into real-scenario testing, iterated repeatedly, and the reinforcement learning optimization strategy is continuously optimized to improve the passing ability and environmental adaptability of the pipeline robot.

[0065] In one embodiment, the above pipeline robot motion control method further includes: In the simulation-to-physical migration stage, the dynamic environment parameters of the trained policy model are calibrated through a domain adaptation algorithm; the dynamic environment parameters at least include sensor measurement error parameters.

[0066] In application, in the simulation-to-physical migration stage, the dynamic environment parameters of the trained policy model can be calibrated through a domain adaptation algorithm to eliminate the differences between the simulation environment and the real physical environment. The domain adaptation algorithm dynamically adjusts the input data and decision logic of the policy model by analyzing the actual physical deviations not covered by the simulation model, such as sensor measurement errors, mechanical transmission clearances, and environmental noise interference. The dynamic environment parameters mainly include sensor measurement error parameters, such as the zero-position offset and sensitivity error of the inertial measurement unit, the angle resolution deviation of the angle encoder, the non-linear response characteristics of the torque sensor, and the ranging noise of the lidar or vision sensor. These errors are usually idealized in the simulation environment, but may significantly affect the generation accuracy of control commands in the real scenario. In addition, the dynamic environment parameters may also include environmental temperature and humidity influence parameters, power supply fluctuation parameters, drive motor thermal effect parameters, etc.

[0067] Taking the sensor measurement error parameters as an example, the calibration process can be divided into two stages: online data acquisition and parameter optimization. In the online data acquisition stage, the real pipeline robot executes preset standard actions (such as uniform linear motion or fixed-point turning), synchronously records the actual output data of the internal sensor module and the external sensor module, and compares it with the expected data of the same action in the simulation environment. For example, the deviation between the actual pitch angle data of the inertial measurement unit and the simulation value can be used to calibrate the zero-position offset, and the difference between the actual displacement increment of the angle encoder and the theoretical value can reflect its resolution error. By statistically analyzing the error distribution of multiple groups of data, a correction model for sensor errors is established.

[0068] During the parameter optimization phase, the domain adaptation algorithm embeds the sensor error correction model into the input processing layer of the policy model to perform real-time correction on the original sensor data. Specifically, for the angle encoder, the measured angle is used to determine the proportional coefficient and offset through linear fitting, mapping the original measurement value to the corrected physical quantity; for the torque sensor, the output is processed by a filtering algorithm to suppress noise interference, and the nonlinear response curve is dynamically adjusted in combination with the actual load feedback. At the same time, a dynamic compensation mechanism is introduced when the policy model generates control instructions. For example, in torque closed-loop control, the corrected actual torque data is fused to offset the deviation of the idealized model in simulation training.

[0069] The calibrated policy model can significantly improve the adaptability to the real environment. For example, in a low-friction pipe section, the torque command generated based on the ideal friction coefficient in simulation training may cause the real drive wheel to slip. The calibrated model can combine the corrected friction coefficient estimate and the actual slip feedback to dynamically reduce the torque output and enhance the weight of the displacement closed-loop control, thereby suppressing slip.

[0070] In addition, an online continuous learning function can be set. When it detects a gradual change in environmental parameters (such as the friction coefficient continuously decreasing due to oil accumulation on the pipe wall), it can automatically trigger the parameter recalibration process to ensure that the policy model maintains high robustness throughout the operation cycle. Finally, the dynamically calibrated policy model can generate control instructions that highly match the real physical constraints, realizing seamless migration from virtual training to physical application, and ensuring the motion stability and task reliability of the pipeline robot under complex working conditions.

[0071] Corresponding to the foregoing method embodiments for realizing application functions, the present invention also provides a pipeline robot motion control system and corresponding embodiments.

[0072] Please refer to Figure 7 , Figure 7 which is a schematic diagram of the module structure of the pipeline robot motion control system. The pipeline robot motion control system is applied to the processor of the pipeline robot; the pipeline robot includes an internal sensor module, an external sensor module, and a plurality of articulation units that are movably connected; each articulation unit includes one or more drive motors; the system includes: A data acquisition unit 71, configured to obtain the state data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module in response to a direction instruction triggered by an operator through a remote terminal; An instruction generation unit 72, configured to generate control instructions for each drive motor based on the direction instruction, the state data, and the pipeline environment data by using a pre-trained policy model; A motion control unit 73, configured to adjust the operating states of the drive motors according to the control instructions.

[0073] In one embodiment, the above instruction generation unit 72 is further configured to: When detecting a communication interruption with the remote end, based on the pipeline environment data and status data, determine the safe areas in the pipeline through the policy model, and generate docking instructions for each drive motor; The above motion control unit 73 is further configured to: control each drive motor to execute the corresponding docking instruction, so that the pipeline robot moves to the safe area.

[0074] In one embodiment, the above instruction generation unit 72 is further configured to: When detecting that the communication interruption duration exceeds a preset threshold, determine the driving path pointing to the target pipeline exit according to the pipeline route map; the pipeline route map is pre-stored or generated in real time according to the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-specified pipeline exit; Based on the status data, pipeline environment data, and driving path, generate exit instructions for each drive motor through the policy model; The above motion control unit 73 is further configured to: control each drive motor to execute the corresponding exit instruction, so that the pipeline robot moves to the target pipeline exit.

[0075] In one embodiment, the above pipeline robot motion control system further includes: A prompt unit, configured to send a prompt message to the remote end when it is detected according to the pipeline environment data that there is a pending area in the driving direction and no new direction instruction is received within a preset time; the pending area is a pipeline section with a branch channel, a turning channel, and / or an impassable area.

[0076] In one embodiment, the above instruction generation unit 72 is further configured to: If no new direction instruction is received within a preset period after sending the prompt message, based on the pipeline environment data and status data, determine the safe areas in the pipeline through the policy model, and generate docking instructions for each drive motor; The above motion control unit 73 is further configured to: control each drive motor to execute the corresponding docking instruction, so that the pipeline robot moves to the safe area.

[0077] Regarding the system in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method. The relevant content regarding the policy model can also be referred to the foregoing content, and will not be elaborated in detail here.

[0078] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A pipeline robot motion control method, characterized in that: A processor applied to a pipeline robot; the pipeline robot comprises an internal sensor module, an external sensor module, and a plurality of movable joint units; each of the joint units comprises one or more drive motors; the method comprises: In response to a direction command triggered by an operator through a remote terminal, the state data of the pipeline robot body collected by the internal sensor module and the pipeline environment data collected by the external sensor module are obtained; Based on the direction command, the state data and the pipeline environment data, a control command for each of the drive motors is generated using a pre-trained strategy model; According to the control instruction, the operating state of each of the driving motors is adjusted.

2. The pipeline robot motion control method according to claim 1, characterized in that: The method further comprises: When it is detected that the communication with the remote end is interrupted, based on the pipeline environment data and status data, a safe area in the pipeline is determined by the strategy model, and a parking instruction for each of the drive motors is generated; Each of the driving motors is controlled to execute a corresponding docking instruction, so that the pipeline robot moves to the safe area.

3. The pipeline robot motion control method according to claim 2, characterized in that: The method further comprises: When it is detected that the communication interruption duration exceeds a preset threshold, a driving path pointing to a target pipeline exit is determined according to a pipeline route map; the pipeline route map is pre-stored or generated in real time according to the pipeline environment data; the target pipeline exit is the nearest pipeline exit or a pre-designated pipeline exit; Based on the state data, pipeline environment data and driving path, generating an exit instruction for each of the drive motors through the strategy model; Each driving motor is controlled to execute a corresponding exit instruction, so that the pipeline robot moves to the target pipeline outlet.

4. The pipeline robot motion control method according to claim 1, characterized in that: The method further comprises: When it is detected according to the pipeline environment data that there is a pending area in the driving direction and no new direction instruction is received within a preset time, a prompt message is sent to the remote end; the pending area is a pipeline section with branch channels, turning channels and / or inaccessible areas.

5. The pipeline robot motion control method according to claim 4, characterized in that: The method further comprises: If no new direction instruction is received within a preset period of time after the prompt information is sent, a safe area in the pipeline is determined through the strategy model based on the pipeline environment data and status data, and a parking instruction for each of the drive motors is generated; Each of the driving motors is controlled to execute a corresponding docking instruction, so that the pipeline robot moves to the safe area.

6. The pipeline robot motion control method according to claim 1, characterized in that: The training method of the strategy model includes: Constructing a dynamic simulation model of the pipeline robot and a corresponding pipeline simulation training environment; According to the preset motion task objectives, a composite reward function is constructed that integrates posture stability reward, slippage suppression reward, and task execution reward. In the pipeline simulation training environment, the control strategy of the drive motor in the dynamic simulation model is iteratively updated through a proximal strategy optimization algorithm, and when the composite reward function converges to a preset value, the trained strategy model is output.

7. The pipeline robot motion control method according to claim 6, characterized in that: The method further comprises: In the simulation-to-physical migration stage, the trained strategy model is calibrated for dynamic environmental parameters through a domain adaptation algorithm; the dynamic environmental parameters at least include sensor measurement error parameters.

8. A pipeline robot motion control system, characterized in that: A processor applied to a pipeline robot; the pipeline robot comprises an internal sensor module, an external sensor module, and a plurality of movable joint units; each of the joint units comprises one or more drive motors; the system comprises: A data acquisition unit, configured to acquire the state data of the pipeline robot body acquired by the internal sensor module and the pipeline environment data acquired by the external sensor module in response to a direction instruction triggered by the operator through the remote end; An instruction generation unit, configured to generate control instructions for each of the drive motors based on the direction instruction, the state data and the pipeline environment data using a pre-trained strategy model; A motion control unit is used to adjust the operating state of each of the drive motors according to the control instruction.

Citation Information

Patent Citations

  • Underwater robot automatic course reversal control method, computer and storage medium

    CN107957729A

  • Control method of long-distance water supply pipeline detection robot

    CN114110303A

  • Robot motion control method and system and electronic equipment

    CN116619382A

  • Self-adaptive multi-mode sensor fusion method and system for robot navigation and obstacle avoidance

    CN118443000A

  • Flexible supporting wheel type driving method and system for pipeline inspection robot

    CN118884836A

Cited By

  • Humanoid robot interaction system

    CN120715912A

  • Large shield slurry pipeline detection robot obstacle crossing algorithm system based on AI calculation

    CN121433312A

  • Intelligent control method for buried pipeline detection robot

    CN122044176A

  • An intelligent control method for a buried pipeline detection robot

    CN122044176B