Self-growing robot based on tendon driving and steering control method thereof
By combining a rope-driven guidance mechanism and a CQL offline reinforcement learning model, the problems of insufficient steering control accuracy and path following error in self-growing robots during long-distance growth are solved, achieving accurate bending and position correction at the end of the growing part and improving the steering capability of the self-growing robot.
Patent Information
- Application Number
- CN202610074332.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-02-27
AI Technical Summary
Existing self-growing robots suffer from insufficient precision in steering control and difficulty in correcting path following errors, especially during long-distance growth. The tendon-driven tension torque is difficult to guarantee the accuracy of bending, and traditional motion planning strategies cannot fully utilize their compliance and adaptability.
A novel rope-driven guiding mechanism is adopted. Through rope-driven guiding components symmetrically distributed on both sides of the inner wall of the cylindrical film, the frictional torque is increased by the toothed arrangement of multiple guide tubes and driving ropes. Combined with the CQL offline reinforcement learning model, precise steering control is achieved, and the tension of the driving rope is adjusted to realize the active bending of the growth site.
Even during long-distance growth, the rope-driven guidance mechanism ensures accurate bending of the end of the growth section, and the CQL offline reinforcement learning model effectively corrects positional deviations, improving the steering accuracy and adaptability of the self-growing robot.
Smart Images

Figure CN121572280A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control, and in particular to a tendon-driven self-growing robot and its steering control method. Background Technology
[0002] Self-growing robots are a novel type of soft robot that can extend their length indefinitely by outwardly rotating materials. Thanks to their strong compliance, they can navigate by interacting with environmental obstacles, thus exhibiting strong adaptability to confined spaces. Effective steering actuation and control mechanisms are needed for navigation in confined spaces. Existing steering methods mainly include active and passive steering. Passive steering primarily uses pre-bending as a technical means. This method is simple to manufacture and allows movement in restricted environments through environmental interaction, maximizing the compliance of the self-growing robot. However, pre-bending is irreversible, making it difficult to correct path-following errors. Active steering mainly includes pneumatic artificial muscles and tendon actuation. Due to the compressibility of gas, the precision of pneumatic artificial muscles is limited, and they cannot grow in tandem with the self-growing robot, resulting in insufficient control over the overall shape of the robot. Tendon actuation relies on tendon displacement generated by pulling the tendon to achieve steering control. Specifically, one end of each of the two tendons is fixed to the end of the self-growing robot. By applying different tensions to the two tendons, they are displaced to different degrees, ultimately causing the self-growing robot to bend towards the side where the tendons are tightened. However, when the self-growing robot is long, the tension torque generated by pulling the tendons cannot guarantee that the bending occurs accurately at the robot's end, thus limiting the robot's ability to control its shape.
[0003] In the motion control of self-growing robots, traditional obstacle avoidance-based motion planning strategies cannot fully leverage the adaptability of self-growing robots, exhibiting limited generalization ability across different scenarios. Learning-based planning strategies can effectively address this issue, but they also face the challenge of constructing suitable datasets. Summary of the Invention
[0004] The purpose of this application is to provide a tendon-driven self-growing robot and its steering control method, which can ensure that the end of the growing site accurately bends.
[0005] To achieve the above objectives, this application provides the following solution: In one aspect, this application provides a tendon-driven self-growing robot, including a drive storage mechanism, a cylindrical film, and a rope-driven guidance mechanism; One end of the cylindrical film is sealed and fixed at the air outlet of the driving storage mechanism, and the other end is turned inward and wound around the material shaft inside the driving storage mechanism through the air outlet. After the cylindrical film is inflated by the driving storage mechanism, it forms the growth part of the self-growing robot. The inward turn of the cylindrical film forms the end of the growth part. The rope-driven guiding mechanism includes two sets of rope-driven guiding components symmetrically distributed on both sides of the inner wall of the cylindrical film. Each set of rope-driven guiding components includes multiple guide tubes and a driving rope. The multiple guide tubes are arranged in a toothed pattern along the axial direction of the cylindrical film. The driving rope passes through all the guide tubes and is fixed at one end at the air outlet and wound around the winding shaft inside the driving storage mechanism through the air outlet. The active bending of the growth site is achieved by adjusting the tension of the two drive ropes.
[0006] Secondly, this application provides a steering control method for a self-growing robot, used to control the active steering of the tendon-driven self-growing robot described in the first aspect. The steering control method includes: The actual end position of the growing part of the self-growing robot is input into the pre-trained CQL offline reinforcement learning model. The CQL offline reinforcement learning model outputs a target tendon pre-bending command. The target tendon pre-bending command is used to control and adjust the tension of the two drive ropes in the self-growing robot so that the actual end position of the growing part reaches the ideal end position. The training steps of the CQL offline reinforcement learning model include: Step 110: Obtain the ideal growth action sequence of the growth site, and apply noise to the ideal growth action sequence to obtain a noisy growth action sequence; Step 120: Perform growth simulation on the growth site based on the current action in the ideal growth action sequence and the noisy growth action sequence to obtain the ideal end position and the actual end position of the growth site; Step 130: If the distance between the ideal end position and the actual end position is greater than a threshold, then apply a tendon pre-bending command to the growth site using the noise growth action sequence to adjust the actual end position, so that the distance between the ideal end position and the actual end position is less than the threshold, and record the ideal end position, the actual end position, and the tendon pre-bending command as a set of training data, and update the current action in the ideal growth action sequence and the noise growth action sequence, wherein the growth actions in the ideal growth action sequence and the noise growth action sequence are used as the current action in order from front to back; Repeating steps 120 and 130 can obtain multiple sets of training data to form a training dataset. Step S140: Train the CQL offline reinforcement learning model using the training dataset.
[0007] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the steering control method for the self-growing robot described in any one of the above.
[0008] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the steering control method for the self-growing robot described in any one of the above descriptions.
[0009] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the steering control method for the self-growing robot described above.
[0010] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a tendon-driven self-growing robot and its steering control method. Compared with existing tendon-driven self-growing robots, the core difference lies in the use of a novel rope-driven guiding mechanism to achieve active bending of the end of the growth part in the self-growing robot. Specifically, the rope-driven guiding mechanism includes two sets of rope-driven guiding components symmetrically distributed on both sides of the inner wall of a cylindrical film. Each set of rope-driven guiding components includes multiple guide tubes and a driving rope. The multiple guide tubes are arranged in a toothed pattern along the axial direction of the cylindrical film. The driving rope passes through all the guide tubes and has one end fixed at the air outlet, while the other end is wound around a winding shaft inside the driving storage mechanism through the air outlet. Because the multiple guide tubes in the same set of rope-driven guiding mechanisms are arranged in a toothed pattern along the axial direction of the cylindrical film, the friction between the driving rope and the guide tubes can be increased, thereby generating a sufficiently large frictional torque to control the morphology of the growth part. Even when the length of the growth part is long, the frictional torque can still ensure that the end of the growth part bends accurately. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1This is a schematic diagram of the structure of a tendon-driven self-growing robot according to one embodiment of this application; Figure 2 This is a schematic diagram of a material roll drive for a tendon-driven self-growing robot according to an embodiment of this application; Figure 3 This is a schematic diagram of the winding shaft drive of a tendon-driven self-growing robot according to an embodiment of this application; Figure 4 This is a schematic diagram of the bending of the growth site under the action of the tendon in one embodiment of this application; Figure 5 This is a schematic diagram showing the dimensions of the tendon and conduit in one embodiment of this application; Figure 6 This is a schematic diagram of segmented modeling and error of the growth site in one embodiment of this application; Figure 7 This is a schematic diagram of the morphological simulation of the growth site in a virtual scene in one embodiment of this application; Figure 8 This is a simulation diagram of the tendon drive of the growth site in a virtual scene in one embodiment of this application; Figure 9 This is a flowchart illustrating the construction process of the training dataset in one embodiment of this application; Figure 10 This is a schematic diagram of the training process of a CQL offline reinforcement learning model in one embodiment of this application; Figure 11 This is a schematic diagram of the state vector of a CQL offline reinforcement learning model in one embodiment of this application; Figure 12 This is a schematic diagram of a sparse action sequence of a CQL offline reinforcement learning model in one embodiment of this application; Figure 13 This is a structural diagram of the Actor network in one embodiment of this application; Figure 14 This is a structural diagram of the Critic network in one embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0015] Reference Figure 1 In one exemplary embodiment, a tendon-driven self-growing robot is provided, which includes a drive storage mechanism, a cylindrical film (plastic cylindrical film body), a rope-driven guidance mechanism, and an end-effector sensing mechanism.
[0016] One end of the cylindrical film is sealed and fixed at the air outlet of the drive storage mechanism, and the other end is turned inward and wrapped around the material shaft inside the drive storage mechanism through the air outlet. After the cylindrical film is inflated by the drive storage mechanism, it forms the growth part of the self-growing robot. The inward turn of the cylindrical film forms the end of the growth part.
[0017] The end-sensing mechanism is located at the end of the growth section. For example, the end-sensing mechanism consists of an end cap and a gyroscope chip installed at the end of the growth section. The gyroscope chip is responsible for sensing the attitude of the end of the growth section in real time, and the end cap is used to drive the gyroscope chip to grow along with the growth section.
[0018] For example, the tubular film can be made of polyethylene plastic.
[0019] As described above, the cylindrical film can be divided into a first part and a second part, which are integrally formed. The end of the first part is sealed and connected to the air outlet of the driving storage mechanism. The driving storage mechanism can inflate the first part to make it expand, thereby forming the growth part of the self-growing robot. The second part is located inside the first part, and its end is wound around a material roll inside the driving storage mechanism through the air outlet. The end of the growth part is the connection point between the first part and the second part, that is, the inward turning position of the cylindrical film. During the continuous inflation and expansion of the first part, the second part continuously turns outward to transform into the first part, and the expansion length of the first part continuously increases, thereby forming the growth effect of the growth part. At the same time, the material roll continuously releases the second part.
[0020] Reference Figure 2 Specifically, each material reel is driven by a drive motor, which is connected to the material reel via a coupling. An opening is provided on one side of the drive storage mechanism, and an outlet plug is sealed at the opening. An air outlet is located on the outlet plug, and the end of the cylindrical film can be glued to the outlet plug for a better sealing connection.
[0021] It is understandable that if no steering mechanism is provided on the tubular film, the growth site will grow into a straight tubular structure. The steering of the self-generated site includes active steering and passive steering. Passive steering involves pre-setting a pre-bending structure on the tubular film, such as folding and fixing one side of the tubular film, causing the two sides of the tubular film to have different lengths at that position. The growth site then bends towards the shorter side. Since this pre-bending needs to be pre-set and the bending angle is fixed, it cannot be adjusted during self-growth, hence it is considered passive bending. The specific implementation steps are as follows: First, the tubular film is laid flat and folded to a certain length, forming Z-shaped folds. Then, the folded side of the film is fixed with tape. Subsequently, when the tubular film is inflated, the growth site will turn towards the side fixed with tape, thereby achieving directional control. It should be noted that the pre-bending structure is existing technology and will not be specifically described in this embodiment. In this embodiment, the focus is on providing a novel tendon-driven active bending mechanism, which is achieved by a rope-driven guiding mechanism.
[0022] The rope-driven guiding mechanism includes two sets of rope-driven guiding assemblies symmetrically distributed on both sides of the inner wall of the cylindrical film. Each set of rope-driven guiding assemblies includes multiple conduits and a driving rope (as a tendon). The multiple conduits are arranged in a toothed pattern along the axial direction of the cylindrical film. The driving rope passes through all the conduits and is fixed at one end to the air outlet, while the other end is wound around a winding shaft inside the driving storage mechanism through the air outlet. Active bending of the growth site is achieved by adjusting the tension of the two driving ropes.
[0023] In the same set of rope-driven guide components, multiple conduits are arranged in a toothed pattern along the axial direction of the cylindrical film. Specifically, this means that the conduits are not arranged in a straight line, but rather adjacent conduits are staggered. For example, an installation reference line extending along the axial direction can be defined on the cylindrical film, with adjacent conduits located on opposite sides of the reference line. Each conduit is parallel to the axial direction of the cylindrical film, and multiple conduits on the same side of the reference line are on the same straight line, with equal spacing between them along the axial direction of the cylindrical film. This unique arrangement effectively increases the friction between the drive rope and the conduits, thereby generating a sufficiently large frictional torque to control the morphology of the growth site.
[0024] For the drive rope, one end is fixed at the air outlet, passes through all the conduits, and the other end wraps back to the air outlet and is wound around the winding shaft inside the drive storage mechanism. Two drive ropes correspond to two different winding shafts. The mechanism of active bending is that one drive rope can be tightened by one winding shaft to increase its tension (increased tension, shortened length), while the other drive rope can be released by the other winding shaft to decrease its tension (decreased tension, increased length). Since the two sets of rope-driven guiding components are located on opposite sides inside the cylindrical film, the growth site can bend towards the side where the drive rope tension increases, and the bending angle is related to the magnitude of the tension change in the drive rope, thus achieving active steering control of the growth site.
[0025] Specifically, refer to Figure 3 Each winding spool is also equipped with a drive assembly, which includes a stepper motor, a coupling, and a gear reducer. The stepper motor is connected to the input shaft of the gear reducer via the coupling, and the output shaft of the gear reducer is connected to the winding spool. The stepper motor can drive the winding spool to rotate through the gear reducer, tightening or releasing the drive rope (drive tendon). Accordingly, to facilitate the movement of the drive rope, a guide pulley is also provided inside the drive storage mechanism to guide the movement of the drive rope. Simultaneously, a force sensor is installed on the drive rope to detect its tension.
[0026] like Figure 4 As shown, when the drive rope on the upper side of the growth section is tensioned while the drive rope on the lower side is released, the frictional torque generated between the drive rope and the guide tube causes the growth section to bend upwards at the position of the last guide tube. It should be noted that the tension in the drive rope is not uniformly distributed along the rope; the bending occurs where the frictional torque is greatest. This is derived from the winch force formula: in, Let be the friction angle. μ Let be the coefficient of friction. Assuming that the friction angle is the same for each segment of the duct, the end of the duct is where the tensile force attenuates the most, thus the frictional torque is the greatest, and therefore the growth site will bend upward at the position of the last duct.
[0027] Assuming the angle of the bending at the growth site is... θ The displacement of the driving rope is l The bending radius of the growth site is r The relationship between the displacement of the driving rope and the end bending of the growth site can be established based on the isocurvature model: l = θr The displacement of the driving rope refers to the displacement of any point on the rope before and after it is pulled.
[0028] If there is a reference point on a tendon, the displacement of the reference point after pulling the tendon relative to the position before pulling can be defined as positive displacement during contraction and negative displacement during release. As described above, this application provides a novel tendon-driven self-growing robot. Compared to existing tendon-driven self-growing robots, the core difference lies in the use of a novel rope-driven guiding mechanism to achieve active bending of the end of the growth section in the self-growing robot. Specifically, the rope-driven guiding mechanism includes two sets of rope-driven guiding components symmetrically distributed on both sides of the inner wall of a cylindrical film. Each set of rope-driven guiding components includes multiple conduits and a driving rope. The multiple conduits are arranged in a toothed pattern along the axial direction of the cylindrical film. The driving rope passes through all the conduits and has one end fixed at the air outlet, while the other end is wound around a winding shaft inside the driving storage mechanism through the air outlet. Because the multiple conduits in the same set of rope-driven guiding mechanisms are arranged in a toothed pattern along the axial direction of the cylindrical film, the friction between the driving rope and the conduits is increased, thereby generating a sufficiently large frictional torque to control the morphology of the growth section. Even when the growth section is long, the frictional torque can still ensure that the end of the growth section accurately bends.
[0029] Preferably, the conduit is a silicone conduit. The silicone conduit can further increase the friction between the conduit and the drive rope. Furthermore, the silicone conduit can maximize the flexibility of the growth site.
[0030] Reference Figure 5 Let D be the inner diameter of the silicone conduit, d be the diameter of the drive rope, n be the length of a single silicone conduit, and m and s be the axial and perpendicular distances between two adjacent silicone conduits, respectively. The angle between the portion of the drive rope between two adjacent silicone conduits and the silicone conduit itself is α.
[0031] Considering that the growth site is driven by the friction between the silicone conduit and the drive rope, α cannot be too small. Conversely, if α is too large, the drive rope will not be able to slide smoothly. Therefore, in this embodiment, α is preferably 20°. Alternatively, any reasonable angle around 20° can be selected. For example, α can be between 15° and 25°.
[0032] Similarly, if the ratio between the diameter of the drive rope and the inner diameter of the silicone conduit is large, it will result in greater friction, preventing the drive rope from sliding smoothly. Conversely, if the ratio is small, sufficient friction will not be generated. Therefore, in this embodiment, d / D is preferably between 0.4 and 0.5.
[0033] The above is an embodiment of a tendon-driven self-growing robot. During the use of the self-growing robot, the growth site grows, with the goal of reaching a preset target position at its end. To achieve this goal, pre-bending structures are pre-set on a cylindrical film, causing the growth site to bend at preset angles at certain preset position nodes. Both the preset position nodes and the preset bending angles are calculated in reverse from the target position. However, considering deviations during the actual growth process (mainly the potential discrepancy between the actual bending angle and the preset bending angle), the actual position of the growth site's end may deviate from the target position. To eliminate this deviation, an active bending mechanism is needed to ensure the actual position of the growth site's end considers the target position, thus eliminating or reducing the positional deviation. This requires very precise calculations to determine the tightening and releasing lengths of the two drive ropes. Therefore, a precise and efficient steering control method is also needed to achieve active steering control of the growth site.
[0034] In one exemplary embodiment, a steering control method for a self-growing robot is also provided, for controlling the active steering of the tendon-driven self-growing robot in the above embodiment.
[0035] The steering control method includes: inputting the actual end-effector position of the self-growing robot's growth site into a pre-trained CQL offline reinforcement learning model; and outputting a target tendon pre-bending command through the CQL offline reinforcement learning model. This target tendon pre-bending command is used to control and adjust the tension of the two drive ropes in the self-growing robot so that the actual end-effector position of the growth site reaches the ideal end-effector position. Specifically, the target tendon pre-bending command includes the tightening or loosening amount of the two drive ropes, i.e., the displacement of the drive ropes.
[0036] Reinforcement learning-based control methods have the advantage of strong generalization ability when facing different scenarios, but they have the problem of difficulty in constructing training datasets. In order to solve this problem, this embodiment also provides a training step for an offline CQL reinforcement learning model, which specifically includes the steps of constructing training datasets.
[0037] The training steps of the CQL offline reinforcement learning model include steps S110, S120 and S130.
[0038] Step 110: Obtain the ideal growth action sequence of the growth site, and add noise to the ideal growth action sequence to obtain the noisy growth action sequence.
[0039] Reference Figure 6Due to the presence of the pre-bending structure, it can be used as a segmentation point for the growth section. A pre-bending structure exists between the preceding and following growth sections, thus creating a predetermined angle between them. Therefore, the growth section can be modeled in segments, and the ideal overall shape of the growth section can be described based on the pre-bending position and angle. This is also known as the "pre-bending action sequence": in, For the first The length of the segment growth section (the length of each segment can be determined based on each pre-bending position). For the first One pre-bending angle.
[0040] However, considering the manufacturing errors in the actual pre-bending process, the actual overall shape of the growth area may vary. for: in, For the first Length error of the segment growth part For the first One pre-bending angle error.
[0041] Considering the outward turning characteristics of the tubular film material, it is difficult to simulate the outward turning behavior of the material in real time in the virtual scene in order to ensure the real-time performance of the simulation.
[0042] like Figure 7 As shown, this embodiment simulates a growth region by connecting multiple growth segments (simulating a growth section) in series. First, the initial growth segment is fixed as the base of the growth region, and a global coordinate system is established here. The length increase of the growth region is simulated by adding new end growth segments at the end.
[0043] The growth segments take the shape of capsules, with the center of the upper part of the previous capsule coinciding with the center of the lower part of the next capsule. Configurable joints are added here, and appropriate springs and damping are added at the joints to simulate the flexible characteristics of the self-growing robot.
[0044] When the growth unit receives a growth command, if it is a straight growth command, a new growth segment is added at the end, and a straight growth connection is added at this point. If it is a curved growth command, a new growth segment is also added at the end, but the configurable joint is offset to a certain extent to add a pre-bending connection.
[0045] like Figure 7As shown, the growth part in the environment will grow along the edge e of the obstacle due to the guidance of the environmental obstacle, until the corner point A of the obstacle.
[0046] like Figure 8 As shown, the active steering behavior driven by tendons is simulated in a virtual scene. The growth segment at the location of the terminal duct of the growth site is found, and a drive is added to the configurable joint corresponding to the growth segment. At this time, other growth segments after this growth segment will follow this joint and rotate to one side.
[0047] Step 120: Perform growth simulation on the growth site based on the current action in the ideal growth action sequence and the noisy growth action sequence to obtain the ideal end position and the actual end position of the growth site.
[0048] Step 130: If the distance between the ideal end position and the actual end position is greater than a threshold, then apply a tendon pre-bending command to the growth site using the noise growth action sequence to adjust the actual end position so that the distance between the ideal end position and the actual end position is less than the threshold. Record the ideal end position, the actual end position, and the tendon pre-bending command as a set of training data. Update the current action in the ideal growth action sequence and the noise growth action sequence. The growth actions in the ideal growth action sequence and the noise growth action sequence are used as the current actions in order from front to back.
[0049] Repeating steps 120 and 130 can yield multiple sets of training data to form a training dataset.
[0050] The following example illustrates steps 120 and 130 in detail. In this example, the virtual scene contains two scenes. The first scene's growth action sequence has no added noise and is considered an ideal growth action sequence. The second scene's growth action sequence has added noise and is considered a noisy growth action sequence. The growth action sequence is the morphological sequence described above, consisting of the length of a growth segment and the joint angle between that growth segment and the previous growth segment. If the joint angle of a growth action is 0, meaning the corresponding growth segment is linearly related to the previous growth segment, then the growth action is called a 0 action; otherwise, it is called a non-zero action.
[0051] Reference Figure 9 In one example, the specific process for constructing the training dataset is as follows: S1: Receive a randomly generated sequence of growth actions for the growth site, including an ideal growth action sequence and a noisy growth action sequence corresponding to two scenarios respectively. The growth actions in the two sequences correspond one-to-one in order from front to back. Iterate through the actions in the two sequences according to the order and use them as the current action in turn.
[0052] S2: Determine if the current action is 0. If yes, proceed to step S4; otherwise, proceed to step S3.
[0053] S3: Add linear joints and their growth segments to the ends of the growth sites in both the first and second scenes.
[0054] S4: Add bending joints and their growth segments to the ends of the growth sites in the first and second scenes, respectively.
[0055] The bending angle of the flexing joint and the length of the growth segment are determined based on the current growth action.
[0056] S5: Determine whether the distance between the end positions of the growth parts in the two scenes (the ideal end position and the actual end position, respectively) is greater than the threshold. If yes, proceed to step S6; otherwise, proceed to step S7.
[0057] S6: Execute a tendon pre-bending command on the growth site in the second scene, causing the terminal growth segment to rotate at a certain angle toward the corresponding position in the first scene, so as to eliminate or reduce the distance between the terminal positions of the growth sites in the two scenes.
[0058] S7: Record the end position of the growth site in the two scenes and the specific instruction value of the corresponding tendon pre-bending command.
[0059] S8: Determine whether the end flag has been reached (determine whether the growth action sequence has been completed). If yes, reset the scene to complete this simulation and obtain the training dataset consisting of the end trajectory of the growth site and the tendon pre-bending instruction set in the two scenes. Otherwise, return to step S2 to continue adding growth segments to complete this simulation.
[0060] It should be noted that the end-effector trajectories in the two scenarios are the ideal end-effector trajectory and the actual end-effector trajectory, respectively. Both end-effector trajectories contain a series of end-effector positions, and there is a one-to-one correspondence between the end-effector positions in the two trajectories. For example, the nth ideal end-effector position in the ideal end-effector trajectory corresponds to the nth actual end-effector position in the actual end-effector trajectory. The corresponding ideal end-effector positions, actual end-effector positions, and tendon pre-bending commands constitute a set of training data, and multiple sets of training data constitute a training dataset.
[0061] Step S130: Train the CQL (Conservative Q-Learning) offline reinforcement learning model using the training dataset. Specifically, this includes: determining experience tuples based on the training dataset. The experience tuples include the current state, action, reward, and next state. The state includes the total length of the growth site, the distance from the terminal duct to the end of the growth site, the angle between the terminal segment and the initial segment of the growth site, the ideal end position, and the actual end position. The action is the tendon pre-flexion command.
[0062] It should be noted that a set of training data constitutes the data basis of an experience tuple. Based on a pair of corresponding ideal end positions and actual end positions, a state can be obtained. The corresponding tendon pre-flexion command is an action in that state, and the reward for this action can be pre-determined and set by the human.
[0063] The reward calculation function is: in, In the state The following actions are given The reward The current distance between the ideal end position and the actual end position. For about The weight function, Action threshold, action Less than Time-based action recognition Zero action, action Not less than Time-based action recognition For non-zero actions, The first reward value, As the second reward value, For the first reward function, The second reward function is the function that provides the reward for the action; both the first and second reward functions are related to the action. and The function, To perform the action front and back The change in quantity.
[0064] The expression for the weighting function is: The expression for the first reward function is: The expression for the second reward function is: Training loss function of CQL offline reinforcement learning model for: in, and These are the conservative regularization term and its coefficient. For standard timing differential loss, For state The expected value function for network, For state The following actions are given value, For action The expected value function This is a collection of action sequences. For the goal Value, i.e., the target network's Values used for supervision Learning through the internet.
[0065] The above describes the training process for the CQL offline reinforcement learning model, defining its empirical tuples, reward function, loss function, etc. See below for reference. Figure 10 This paper provides an example to illustrate the training process of an offline reinforcement learning model using CQL.
[0066] S-A1: Loads an offline reinforcement learning dataset, which may include training datasets corresponding to multiple pairs of terminal trajectories (ideal terminal trajectory and actual terminal trajectory). The training dataset for each pair of terminal trajectories includes multiple empirical tuples. (One empirical tuple corresponds to a pair of ideal end positions and actual end positions).
[0067] in, , and , These include the current state, action, reward, and next state. And references... Figure 11 The state vector is: The total length of the growth site, The distance from the end of the terminal duct to the end of the growth site. The angle between the terminal segment of the growth site and the positive Y-axis. ), ( ( ) represent the ideal end position and the actual end position, respectively.
[0068] S-A2: Initializes the Critic and Actor networks and sets hyperparameters such as learning rate, discount factor, and CQL regularization coefficient.
[0069] S-A3: Sampled Batch Training. A batch of empirical data is randomly sampled from the offline dataset for computation in the current iteration.
[0070] S-A4: Calculate the current reward function. When the distance between the ideal terminal position and the actual terminal position is small, the policy is constrained to output zero actions as much as possible to avoid unnecessary control intervention; when the deviation between the two is large, non-zero actions are encouraged to drive the growth site to actively correct its course. Based on the above principles, the above-mentioned reward function was designed in this embodiment.
[0071] S-A5: Calculate the CQL loss, including the conservative regularization term and the standard time-series difference loss: For details, please refer to the above. Explanation.
[0072] S-A6: Gradient Descent Network Update. The gradient of the loss function with respect to the network parameters is calculated using the backpropagation algorithm, and the Critic and Actor network parameters are updated using gradient descent.
[0073] S-A7: Training round determination. Determine whether the preset training round or convergence condition has been met; if not, return to step S-A3 to continue training; if met, proceed to step S-A8.
[0074] Step S-A8: Output Policy. Output the trained policy network as the final decision policy.
[0075] For example, sparse action sequences in reinforcement learning are as follows: Figure 12 As shown, the robot determines whether to perform a zero action or a non-zero action based on the current state at regular intervals. The robot's actions are sparse action sequences, with most values being zero, indicating that the tendons are not performing any action. A small portion of the values are non-zero, indicating that the tendons are being controlled to move, thus eliminating path-following errors. Figure 12 middle, The value of the current action is: when the current action is greater than 0, the terminal segment of the growth site is deflected to the left; when the current action is less than 0, the terminal segment of the growth site is deflected to the right.
[0076] For example, the structure diagram of the Actor network is as follows: Figure 13As shown, a bimodal structure is adopted to model two decision modes, "zero action" and "non-zero action," for the sparse action decision problem. The state input is first processed through two fully connected layers to extract high-dimensional feature representations, and then enters two parallel output branches.
[0077] The zero-action output branch is composed of an FC(128)–ELU–Sigmoid structure, which is used to predict the probability of outputting zero actions; the non-zero-action output branch is composed of an FC(128)–ELU–Tanh structure, which is used to generate continuous non-zero action values.
[0078] The output of the Actor network is shown in the following formula: in, For action, Values for non-zero actions. As a gating variable, it controls whether to output zero action, and its behavior follows a probability. Bernoulli distribution , The current distance between the ideal end position and the actual end position. The sharpness parameter controls the steepness of the decision boundary. and These represent the center and range of the distance threshold, respectively. The standard deviation is denoted as .
[0079] The structure diagram of the Critic network is as follows: Figure 14 As shown, to evaluate the value of state-action pairs, the network consists of three fully connected layers and residual connected layers to enhance feature representation and alleviate the gradient vanishing problem in deep networks. The Q-network output incorporates a CQL penalty term to conservatively estimate out-of-distribution actions, effectively suppressing overestimation caused by distribution shifts during offline training and improving the stability and security of the policy. In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0080] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0081] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0082] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0083] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0084] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0085] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0086] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A tendon-driven self-growing robot, characterized in that, Includes a drive storage mechanism, a cylindrical film, and a rope-driven guide mechanism; One end of the cylindrical film is sealed and fixed at the air outlet of the driving storage mechanism, and the other end is turned inward and wound around the material shaft inside the driving storage mechanism through the air outlet. After the cylindrical film is inflated by the driving storage mechanism, it forms the growth part of the self-growing robot. The inward turn of the cylindrical film forms the end of the growth part. The rope-driven guiding mechanism includes two sets of rope-driven guiding components symmetrically distributed on both sides of the inner wall of the cylindrical film. Each set of rope-driven guiding components includes multiple guide tubes and a driving rope. The multiple guide tubes are arranged in a toothed pattern along the axial direction of the cylindrical film. The driving rope passes through all the guide tubes and is fixed at one end at the air outlet and wound around the winding shaft inside the driving storage mechanism through the air outlet. The active bending of the growth site is achieved by adjusting the displacement of the two drive ropes.
2. The tendon-driven self-growing robot according to claim 1, characterized in that, The catheter is a silicone catheter.
3. The tendon-driven self-growing robot according to claim 2, characterized in that, The ratio of the diameter of the drive rope to the inner diameter of the silicone conduit is between 0.4 and 0.
5.
4. The tendon-driven self-growing robot according to claim 2 or 3, characterized in that, The portion of the drive rope located between two adjacent silicone conduits forms an angle of 20° with the silicone conduit.
5. A steering control method for a self-growing robot, characterized in that, Used to control the active steering of the tendon-driven self-growing robot according to any one of claims 1-4; The steering control method includes: The actual end position of the growing part of the self-growing robot is input into the pre-trained CQL offline reinforcement learning model. The CQL offline reinforcement learning model outputs a target tendon pre-bending command. The target tendon pre-bending command is used to control and adjust the displacement of the two drive ropes in the self-growing robot so that the actual end position of the growing part reaches the ideal end position. The training steps of the CQL offline reinforcement learning model include: Step 110: Obtain the ideal growth action sequence of the growth site, and apply noise to the ideal growth action sequence to obtain a noisy growth action sequence; Step 120: Perform growth simulation on the growth site based on the current action in the ideal growth action sequence and the noisy growth action sequence to obtain the ideal end position and the actual end position of the growth site; Step 130: If the distance between the ideal end position and the actual end position is greater than a threshold, then apply a tendon pre-bending command to the growth site using the noise growth action sequence to adjust the actual end position, so that the distance between the ideal end position and the actual end position is less than the threshold, and record the ideal end position, the actual end position, and the tendon pre-bending command as a set of training data, and update the current action in the ideal growth action sequence and the noise growth action sequence, wherein the growth actions in the ideal growth action sequence and the noise growth action sequence are used as the current action in order from front to back; Repeating steps 120 and 130 can obtain multiple sets of training data to form a training dataset. Step S140: Train the CQL offline reinforcement learning model using the training dataset.
6. The steering control method for a self-growing robot according to claim 5, characterized in that, The offline reinforcement learning model for CQL is trained using the training dataset, specifically including: Based on the training dataset, experience tuples are determined. Each experience tuple includes the current state, action, reward, and next state. The state includes the total length of the growth site, the distance from the distal duct to the end of the growth site, the angle between the distal segment of the growth site and the positive Y-axis, the ideal end position, and the actual end position. The action is the tendon pre-flexion command. The reward calculation function is: in, In the state The following actions are given The reward The current distance between the ideal end position and the actual end position. For about The weight function, Action threshold, action Less than Time-based action recognition Zero action, action Not less than Time-based action recognition For non-zero actions, The first reward value, As the second reward value, For the first reward function, The second reward function is the function that rewards the first reward function and the second reward function, both of which are related to the action. and The function, To perform the action front and back The change in quantity.
7. The steering control method for a self-growing robot according to claim 6, characterized in that, The expression for the weighting function is: The expression for the first reward function is: The expression for the second reward function is: in, To perform the action front and back The change in quantity.
8. The steering control method for a self-growing robot according to claim 6, characterized in that, The training loss function of the CQL offline reinforcement learning model for: in, and These are the conservative regularization term and its coefficient. For standard timing differential loss, For state The expected value function for network, For state The following actions are given value, For action The expected value function This is a collection of action sequences. For the goal value.
9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steering control method for the self-growing robot according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steering control method for the self-growing robot as described in any one of claims 1-8.