Self-growing robot and motion planning method
By using a stepper motor to drive flexible materials in self-growth robots, combining pre-bending structure and pneumatic artificial muscles, the problem of insufficient motion accuracy and flexibility of self-growth robots in narrow spaces is solved, and higher motion accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510504517.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-06-10
AI Technical Summary
When self-growth robots move in a narrow space, there are problems of insufficient accuracy and flexibility.
The stepper motor drives flexible materials to achieve self-growth, combining pre-bending structures and pneumatic artificial muscles, and controlling stepper motors, pre-bending structures and pneumatic artificial muscles, the steering and path adjustment are achieved using environmental interaction.
It improves the accuracy and flexibility of the movement of self-growth robots in a narrow space, reduces path following errors, and improves the ability to adapt to complex environments.
Smart Images

Figure CN120116209A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of flexible robots, and particularly to a self-growing robot and a motion planning method. Background Art
[0002] A self-growing robot is a new type of flexible robot that can grow by turning the material outwards and can infinitely extend its own length. Therefore, it is very suitable for application in narrow spaces. However, in narrow spaces, there are still great problems with the motion accuracy and flexibility of the self-growing robot. Therefore, a self-growing robot that can move flexibly and accurately in narrow spaces is needed. Summary of the Invention
[0003] The purpose of the present application is to provide a self-growing robot and a motion planning method, which can improve the accuracy and flexibility of motion in narrow spaces.
[0004] To achieve the above purpose, the present application provides the following solutions:
[0005] In a first aspect, the present application provides a self-growing robot, including: a stepper motor, a drive storage mechanism, a flexible body, a pre-bending structure, a pneumatic artificial muscle, a gyroscope chip, and a controller;
[0006] The head end of the flexible body is connected to the drive storage mechanism; the stepper motor is connected to the drive storage mechanism; the stepper motor is used to drive the flexible material in the drive storage mechanism to achieve the self-growth of the flexible body; the stepper motor is also used to obtain the length information of the self-growing robot; the gyroscope chip is used to obtain the heading angle information of the end of the flexible body;
[0007] The pre-bending structure is used to adjust the growth direction of the flexible body; the pneumatic artificial muscles are arranged on both sides of the flexible body; the pneumatic artificial muscles are used to control the attitude of the end of the flexible body; the controller is used to control the stepper motor, the pre-bending structure, and the pneumatic artificial muscles to achieve turning according to the length information, the heading angle information, and the attitude by means of environmental interaction.
[0008] In one embodiment, the drive storage mechanism includes a coupling and a material reel;
[0009] The stepper motor is connected to the material reel through the coupling; the flexible material is wound on the material reel.
[0010] In one embodiment, the drive storage mechanism further includes an outlet plug, an air inlet, and a drive storage housing;
[0011] The flexible body is connected to the outlet plug; the outlet plug and the air inlet are both arranged on the driving and storage housing; the coupling and the material reel are both arranged inside the driving and storage housing.
[0012] In one embodiment, the self-growing robot further includes a terminal bracket;
[0013] The terminal bracket is arranged at the end of the flexible body; the gyroscope chip is arranged on the terminal bracket.
[0014] In one embodiment, the flexible body is a polyethylene plastic film.
[0015] In a second aspect, the present application provides a motion planning method for a self-growing robot. The motion planning method for the self-growing robot is applied to the self-growing robot, and the method includes:
[0016] Modeling the self-growing robot in a virtual scenario; the virtual scenario includes a virtual real scenario and a virtual ideal scenario;
[0017] Obtaining the state information of the self-growing robot in the virtual scenario; the state information includes the length information of the self-growing robot and the heading angle information of the end of the flexible body;
[0018] Using a pre-bending action selection network according to the state information of the self-growing robot in the virtual scenario to obtain a pre-bending action;
[0019] Controlling to add a new capsule to the end of the flexible body of the self-growing robot in the virtual scenario according to the pre-bending action; the flexible body is modeled using capsules in the virtual scenario;
[0020] Obtaining the state information of the self-growing robot in the current virtual scenario and performing Kalman filtering to obtain the state information in the current virtual scenario;
[0021] Determining the motion data of the pneumatic artificial muscle according to the state information in the current virtual scenario using a perception and error correction network;
[0022] Controlling the pneumatic artificial muscle to perform motion according to the motion data.
[0023] In one embodiment, the pre-bending action selection network includes an input layer, a first fully connected layer, a first dropout layer, a second fully connected layer, a second dropout layer, a third fully connected layer, and an output layer connected in sequence.
[0024] In one embodiment, the training process of the pre-bending action selection network specifically includes:
[0025] Initialize the virtual ideal scenario and the flexible body of the self-growing robot; the flexible body is modeled using a capsule body;
[0026] Extract the state information of the self-growing robot in the initialized virtual ideal scenario as the first training sample;
[0027] According to the first training sample, use the first reinforcement learning network to select the pre-bending action information corresponding to the first training sample;
[0028] Control to add a new capsule body at the end of the flexible body in the virtual ideal scenario according to the pre-bending action information corresponding to the first training sample to obtain the turned growth direction and the turned included angle; the included angle is the included angle between the self-growing robot and the target point;
[0029] Calculate the turning behavior reward according to the first training sample, the turned growth direction and the turned included angle;
[0030] Determine the pre-bending action selection network according to the turning behavior reward, the capsule body at the end of the flexible body and the target point.
[0031] In one embodiment, the perception and error correction network and the pre-bending action selection network have the same structure.
[0032] In one embodiment, the training process of the perception and error correction network specifically includes:
[0033] Initialize the virtual scenario;
[0034] Obtain the state information of the self-growing robot in the initialized virtual scenario as the second training sample;
[0035] According to the second training sample, use the pre-bending action selection network to determine the pre-bending action of the self-growing robot in the initialized virtual scenario;
[0036] Add a new capsule body at the end of the flexible body in the virtual scenario according to the pre-bending action of the self-growing robot in the initialized virtual scenario;
[0037] Obtain the state information of the self-growing robot in the current training virtual scenario and perform Kalman filtering to obtain the state information in the current training virtual scenario;
[0038] According to the state information in the current training virtual scenario, use the second reinforcement learning network to determine the training motion data of the pneumatic artificial muscle;
[0039] Control the pneumatic artificial muscle to move in the virtual real scenario according to the training motion data to obtain the motion result;
[0040] Calculate the reward function of the second reinforcement learning network according to the motion result
[0041] Determine the perception and error correction network according to the reward function, the capsule at the end of the flexible body, and the target point.
[0042] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application:
[0043] This application provides a self-growing robot and a motion planning method, which uses a stepping motor to replace the flexible material to realize the self-growth of the flexible body, adjusts the growth direction of the flexible body through a pre-bending structure, realizes movement in a narrow space, and controls the posture of the end of the flexible body through the pneumatic artificial muscles arranged on both sides of the flexible body to reduce the path tracking error of the self-growing robot, thereby improving the accuracy and flexibility of the self-growing robot moving in a narrow space. Brief Description of the Drawings
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 Schematic diagram of the structure of a self-growing robot in an embodiment of the present application;
[0046] Figure 2 Schematic diagram of adding a capsule to the virtual scene simulation in an embodiment of the present application;
[0047] Figure 3 Schematic diagram of the overall virtual scene simulation in an embodiment of the present application;
[0048] Figure 4 Schematic diagram of simulating pneumatic artificial muscles in the virtual scene in an embodiment of the present application;
[0049] Figure 5 Comparison diagram of the virtual scene and the actual self-growing robot in an embodiment of the present application;
[0050] Figure 6 Schematic diagram of the pre-bending manufacturing error in an embodiment of the present application;
[0051] Figure 7 Schematic diagram of the structure of the reinforcement learning network in an embodiment of the present application;
[0052] Figure 8 Schematic diagram of the state vector of the first reinforcement learning network in an embodiment of the present application;
[0053] Figure 9Schematic diagram of the training process of the pre-bending motion selection network in an embodiment of the present application;
[0054] Figure 10 Schematic diagram of the state vector of the second reinforcement learning network in an embodiment of the present application;
[0055] Figure 11 Schematic diagram of the training process of the perception and error correction network in an embodiment of the present application;
[0056] Figure 12 Schematic diagram of the process of the motion planning method of the self-growing robot in an embodiment of the present application;
[0057] Figure 13 Top view of a self-growing robot in an embodiment of the present application.
[0058] Reference numerals: end bracket - 1, gyroscope chip - 2, pneumatic artificial muscle - 3, stepper motor - 4, drive storage mechanism - 5, flexible body - 6, pre-bending structure - 7, outlet plug - 8, coupling - 9, material reel - 10, air inlet - 11. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0060] The currently existing driving technologies limit the flexibility and motion accuracy of self-growing robots. The pre-bending method can easily access narrow spaces, but this method is irreversible and it is difficult to correct errors during motion. Using pneumatic artificial muscles to control the steering of the robot, it can control the flexible movement of the end, but it is difficult to control the overall shape of the self-growing robot to cross complex obstacles. At the same time, the existing control methods or means mostly use pre-bending environment interaction or active steering methods for navigation. The method of environment interaction navigation can give full play to the compliance of the self-growing robot itself, but it lacks the ability of active obstacle avoidance. The active steering method also has the problem of insufficient accuracy during motion. At the same time, the traditional reinforcement learning control method also has problems such as high training cost and great training difficulty. In view of the above problems, the self-growing robot provided in this application uses a combination of pre-bending and pneumatic artificial muscles to control the self-growing robot to navigate in a narrow space, which can maximize the flexibility of the self-growing robot. At the same time, this application proposes a method of combining virtual and real reinforcement learning to control the navigation of the self-growing robot. The method of combining virtual and real improves the training speed and reduces the training cost. This algorithm can effectively reduce the path following error of the self-growing robot and improve the adaptability of the self-growing robot to complex environments.
[0061] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in conjunction with the accompanying drawings and specific embodiments.
[0062] As Figure 1 and Figure 13 shown, this application provides a self-growing robot, including: a stepper motor 4, a drive storage mechanism 5, a flexible body 6, a pre-bending structure 7, pneumatic artificial muscles 3, a gyroscope chip 2, and a controller; the head end of the flexible body 6 is connected to the drive storage mechanism 5; the stepper motor 4 is connected to the drive storage mechanism 5; the stepper motor 4 is used to drive the flexible material in the drive storage mechanism 5 to achieve the self-growth of the flexible body 6; the stepper motor 4 is also used to obtain the length information of the self-growing robot; the gyroscope chip 2 is used to obtain the heading angle information of the end of the flexible body 6; the pre-bending structure 7 is used to adjust the growth direction of the flexible body 6; the pneumatic artificial muscles 3 are arranged on both sides of the flexible body 6; the pneumatic artificial muscles 3 are used to control the attitude of the end of the flexible body 6; the controller is used to control the stepper motor 4, the pre-bending structure 7, and the pneumatic artificial muscles 3 according to the length information, the heading angle information, and the attitude by using the environment interaction method to achieve steering.
[0063] The two - stage driven self - growing robot proposed in this application uses the method of pre - bending combined with pneumatic artificial muscles 3 to navigate in a narrow space. The robot uses the pre - bending method and can navigate in a narrow space through means of environmental interaction. At the same time, pneumatic artificial muscles 3 are installed on both sides of the self - growing robot, which can flexibly control the end pose of the self - growing robot and reduce the path - following error of the robot.
[0064] In an exemplary embodiment, the drive storage mechanism 5 includes a coupling 9 and a material reel 10; the stepper motor 4 is connected to the material reel 10 through the coupling 9; a flexible material is wound on the material reel 10.
[0065] In an exemplary embodiment, the drive storage mechanism 5 further includes an outlet plug 8, an air inlet 11 and a drive storage housing; the flexible body 6 is connected to the outlet plug 8; both the outlet plug 8 and the air inlet 11 are provided on the drive storage housing; both the coupling 9 and the material reel 10 are provided inside the drive storage housing.
[0066] In practical applications, an encoder is set in the stepper motor 4, and the stepper motor 4 is connected to the material reel 10 of the drive storage mechanism 5 through the coupling 9. Among them, the material required for the growth of the self - growing robot is wound and stored on the material reel 10. By filling compressed air from the air inlet 11 of the drive storage mechanism 5, and at the same time, the stepper motor 4 drives the material reel 10 to rotate to release the material required for the growth of the robot, so as to achieve the purpose of increasing the length of the self - growing robot. The encoder of the stepper motor 4 calculates the length of the released material by checking the rotation angle of the stepper motor 4 to obtain the length information of the self - growing robot.
[0067] In an exemplary embodiment, the self - growing robot further includes a terminal bracket 1; the terminal bracket 1 is arranged at the end of the flexible body 6; the gyroscope chip 2 is arranged on the terminal bracket 1.
[0068] In an exemplary embodiment, the flexible body 6 is a polyethylene plastic film. The flexible body 6 of the self - growing robot is composed of a polyethylene plastic film and is connected to the outlet plug 8 of the drive storage mechanism 5 through tape.
[0069] Before the self - growing robot grows, the growth direction of the self - growing robot is controlled by presetting the pre - bending structure 7. During the process of setting the pre - bending structure 7, first, the polyethylene plastic film is flattened, then folded along the length direction of the film to form a Z - shaped structure, and then the folded part is fixed with plastic, so that a length difference is formed on both sides of the robot, and the robot will grow towards the side with a shorter length.
[0070] The pneumatic artificial muscle 3 is adhered to the left and right sides of the self-growing robot to control the posture of the end of the self-growing robot and reduce the path following error of the self-growing robot.
[0071] Since the self-growing robot grows by turning the material outwards, the material at its end is constantly updated. Therefore, the gyroscope chip 2 is installed on the end bracket 1 so that the gyroscope chip 2 can grow together with the self-growing robot. The gyroscope chip 2 is used to sense the heading angle information of the end of the self-growing robot. Figure 1 One implementation of pre-bending plus the drive of the pneumatic artificial muscle 3 is only for demonstration, and it can also be other drive methods that can achieve pre-bending plus the pneumatic artificial muscle 3.
[0072] Compared with the current self-growing robots, the self-growing robot in this application adopts a steering method combining pre-bending and the pneumatic artificial muscle 3. By setting the pre-bending in advance and based on the method of environmental interaction, it can enter narrow spaces. The pneumatic artificial muscle 3 is installed at the end of the self-growing robot, which can control the movement of the end of the self-growing robot, increase the flexibility of the self-growing robot, and reduce errors.
[0073] Based on the same inventive concept, the embodiment of this application also provides a self-growing robot motion planning method applied to the self-growing robot involved above. The solution provided by this method to solve the problem is similar to the solution described in the above device. Therefore, the specific limitations in one or more embodiments of the self-growing robot motion planning method provided below can refer to the limitations on the self-growing robot in the above text, and will not be repeated here.
[0074] In an exemplary embodiment, a self-growing robot motion planning method is provided. The self-growing robot motion planning method is applied to the self-growing robot, and the method includes:
[0075] Model the self-growing robot in a virtual scene; the virtual scene includes a virtual real scene and a virtual ideal scene.
[0076] Obtain the state information of the self-growing robot in the virtual scene; the state information includes the length information of the self-growing robot and the heading angle information of the end of the flexible body.
[0077] Use the pre-bending action selection network according to the state information of the self-growing robot in the virtual scene to obtain a pre-bending action.
[0078] Control the addition of new capsule bodies to the end of the flexible body of the self-growing robot in the virtual scene according to the pre-bending action; the flexible body is modeled using capsule bodies in the virtual scene.
[0079] Obtain the state information of the self-growing robot in the current virtual scene and perform Kalman filtering to obtain the state information in the current virtual scene.
[0080] Determine the motion data of the pneumatic artificial muscle according to the state information in the current virtual scene by using the perception and error correction network.
[0081] Control the pneumatic artificial muscle to move according to the motion data.
[0082] Compared with the previous reinforcement learning algorithms, the method provided in this application, the method combining virtual and real can greatly accelerate the training process of the self-growing robot, while reducing the training cost and improving the accuracy. At the same time, compared with the traditional motion planning algorithms, this application can effectively improve the adaptability of the self-growing robot to narrow spaces, improve the motion accuracy of the self-growing robot, and effectively improve the autonomous perception and decision-making ability of the robot in narrow spaces.
[0083] In practical applications, first, virtual modeling needs to be carried out for the flexible body, pneumatic artificial muscle, and pre-bending structure in the above Figure 1 structure. The self-growing robot is simulated by connecting multiple capsule bodies in series. Connecting multiple capsule bodies in series represents the flexible body of the self-growing robot. Using the method of connecting multiple capsule bodies in series to simulate the flexible body of the self-growing robot can achieve the simulation of the flexible body of the self-growing robot with less computing power consumption.
[0084] At the same time, the self-growing robot grows by turning out the material, and it is difficult to achieve a complete simulation in the virtual scene. Therefore, the method of adding a new capsule body at the end is used to simulate the increase in the length of the self-growing robot.
[0085] As Figure 2 shown, two adjacent capsule bodies are connected by configurable joints. The upper center of the sphere of the previous capsule body coincides with the lower center of the sphere of the next capsule body, and the collision between adjacent capsule bodies is cancelled. Only consider the behavior of the self-growing robot in the plane. For this purpose, the rotational degrees of freedom in the XZ direction are locked, and only the rotational degree of freedom in the Y direction is retained. At the same time, to simulate the behavior of the inflatable flexible beam, appropriate springs and dampers are added at the joints. And the material of the capsule body is set to add a certain elasticity to it. The pre-bending structure of the self-growing robot can be simulated by changing the angle between two adjacent capsule bodies.
[0086] As Figure 3 shown, a coordinate system is established at the position of the bottom center of the sphere of the first capsule body, and the first capsule body is fixed as the base of the self-growing robot. The growth process of the self-growing robot is simulated by adding a new rigid capsule body at the end.
[0087] As Figure 4As shown in the figure, in order to simplify the whole process, an equal-curvature model is adopted to model the self-growing robot. The self-growing robot is regarded as an arc, where α is the central angle of the arc. At this time, drives are added at all joints. Assuming there are n joints in total, the rotation angle of a single joint is α / n. At this time, the simulation of the pneumatic artificial muscle can be completed.
[0088] In summary, the comparison between the specific self-growing robot and the virtual scene is as follows Figure 5 As shown in the figure, the comparison is carried out from two aspects: structural characteristics and driving strategy. The self-growing robot body is simulated by connecting capsules in series. The growth process of the robot is simulated by adding a new capsule at the end. The pre-bending is simulated by adjusting the initial angles of the joints between the capsules. The equal-curvature modeling method is used to control the uniform rotation of the joints to simulate the motion mode of the pneumatic artificial muscle.
[0089] As Figure 6 shown in the figure, there are manufacturing errors in the pre-bending during the actual application process. There may be an error of dα in the pre-bending angle; at the same time, there is also an error of dl in the set position of the pre-bending.
[0090] In an exemplary embodiment, the structures of the first reinforcement learning network and the second reinforcement learning network are both Figure 7 as shown in the figure. The pre-bending action selection network is the trained first reinforcement learning network. The pre-bending action selection network includes an input layer, a first fully connected layer, a first dropout layer, a second fully connected layer, a second dropout layer, a third fully connected layer, and an output layer connected in sequence. The input of the reinforcement learning network is: the state vector of the self-growing robot, where the state vector is the state information represented by a vector, and the state information of the self-growing robot is obtained from the environmental state information. The output is: the action of the self-growing robot. The neural network mainly includes three fully connected layers, each fully connected layer is composed of a ReLU activation function, and there are two dropout layers between the fully connected layers to discard some neurons of the neural network with a certain probability, so as to prevent the model from overfitting and enhance the generalization ability of the model.
[0091] In an exemplary embodiment, the training process of the pre-bending action selection network specifically includes:
[0092] Initialize the virtual ideal scene and the flexible body of the self-growing robot; the flexible body is modeled using capsules.
[0093] Extract the state information of the self-growing robot in the initialized virtual ideal scene as the first training sample.
[0094] Select the pre-bending action information corresponding to the first training sample using the first reinforcement learning network according to the first training sample.
[0095] Control to add a new capsule at the end of the flexible body according to the pre-bending action information corresponding to the first training sample in a virtual ideal scenario to obtain the turned growth direction and the turned included angle; the included angle is the included angle between the self-growing robot and the target point.
[0096] Calculate the turning behavior reward according to the first training sample, the turned growth direction and the turned included angle.
[0097] Determine the pre-bending action selection network according to the turning behavior reward, the capsule at the end of the flexible body and the target point.
[0098] Specifically, according to the virtual model established above, perform the control process of the pre-bending action so that the robot can reach the target position. The algorithm flow chart of the pre-bending action selection network is as Figure 9 shown.
[0099] S1-1, Initialize the virtual ideal environment, and initialize the environmental obstacles and the capsules at the base.
[0100] S1-2, Obtain the state information of the self-growing robot in the virtual ideal scenario.
[0101] S1-3, Select the pre-bending action information of the self-growing robot based on the reinforcement learning algorithm. Control the turning angle of the capsule through the reinforcement learning algorithm, where the reinforcement learning algorithm is a process of training the reinforcement learning network. The virtual scenario sends the state vector to the reinforcement learning network for storing experience and subsequent learning. For the reinforcement learning algorithm of the pre-bending action selection of the self-growing robot, as Figure 8 shown, the state vector x of the self-growing robot can be defined as:
[0102] x = [x robot x target x obs (1)
[0103] where
[0104]
[0105] p x ,p y are the coordinates at the end of the self-growing robot, θ end is the included angle between the growth direction of the self-growing robot and the line connecting to the target point, t x ,t y are the coordinates of the target point. x robot is the sub-state vector of the self-growing robot, x target is the sub-state vector of the target point, x obs is the sub-state vector of the obstacle. Assume that all obstacles in the environment are rectangles, O xi ,O yiis the coordinate of the centroid of the i-th obstacle, w i , h i are the length and width of the obstacle, and θ i is the angle of the obstacle.
[0106] S1-4, send pre-bending action information to the virtual scene.
[0107] S1-5, the virtual scene receives the pre-bending action information.
[0108] S1-6, the virtual scene adds a new capsule at the end of the flexible body based on the received pre-bending action information. S1-7, at this time, the synchronous reinforcement learning algorithm synchronously executes the calculation of the reward function r:
[0109] r = w action ·r action +w target ·r target (3)
[0110] R action , r target are the reward functions for the executed action and the distance to the target point respectively, and w action and w target are the weights of the action reward function and the target reward function respectively.
[0111]
[0112] Among them,
[0113]
[0114] The reward function is used to measure the quality of the robot's turning strategy, can guide the agent to learn the optimal strategy, and directly affects step S1-8. d is the Euclidean distance between the target point and the end of the robot, and θ t is the angle between the growth direction of the robot and the target direction.
[0115] The reward for approaching the target adopts an exponential function form to make the reward smoother. According to the angle between the growth direction of the self-growing robot before and after turning and the target point, calculate the turning behavior reward, encourage the robot to turn reasonably, and punish large-scale ineffective turning behaviors.
[0116] S1-8, after the reward function calculation is completed, store the experience and update the network. Store the current state of the self-growing robot, the actions taken, and the current reward function as experience, and update the weight information of each connection layer of the network using the gradient policy algorithm.
[0117] S1-9. Determine whether the end capsule body reaches the target point. If so, execute step S1-10 to delete all newly added capsule bodies in the environment and reset the current environment. Otherwise, return to execute step S1-2.
[0118] If the robot reaches the target point or the distance from the target point is less than a certain threshold, it is determined that the experiment is successful this time; otherwise, the experiment fails this time.
[0119] S1-11. Complete a single experiment, and further execute step S1-12 to determine whether the success rate of reaching the target point currently reaches the threshold. If so, execute step S1-13 to store the current model and complete the training. Otherwise, return to execute step S1-2 to start the next round of loop.
[0120] In an exemplary embodiment, the structures of the perception and error correction network and the pre-bending action selection network are the same.
[0121] In an exemplary embodiment, the training process of the perception and error correction network specifically includes:
[0122] Initialize the virtual scene.
[0123] Obtain the state information of the self-growing robot in the initialized virtual scene as the second training sample.
[0124] According to the second training sample, use the pre-bending action selection network to determine the pre-bending action of the self-growing robot in the initialized virtual scene.
[0125] Add new capsule bodies to the end of the flexible body in the virtual scene according to the pre-bending action of the self-growing robot in the initialized virtual scene.
[0126] Obtain the state information of the self-growing robot in the current training virtual scene and perform Kalman filtering to obtain the state information in the current training virtual scene.
[0127] According to the state information in the current training virtual scene, use the second reinforcement learning network to determine the training motion data of the pneumatic artificial muscle.
[0128] Control the pneumatic artificial muscle to move in the virtual real scene according to the training motion data to obtain the motion result.
[0129] Calculate the reward function of the second reinforcement learning network according to the motion result.
[0130] Determine the perception and error correction network according to the reward function, the capsule bodies at the end of the flexible body, and the target point.
[0131] Specifically, the state vector of the self-growing robot error correction algorithm is as Figure 10As shown, it can be defined as:
[0132] x = [x ideal x real x target x obs (4)
[0133] Where
[0134] x ideal = [p x p y θ end (5)
[0135] x real = [p x0 p y0 θ end0
[0136] x target = [t x t y
[0137] x obs = [O xi , O yi , w i , h i , θ i
[0138] Where p x , p y , θ end are the coordinates and orientation of the end of the self-growing robot in the ideal case, and p x0 , p y0 , θ end0 are the coordinates and orientation of the end of the self-growing robot in the actual case; x ideal is the ideal environmental sub-state vector, and x real is the actual environmental sub-state vector. Where O xi O yi is the coordinate of the i-th obstacle, and w i h i are the length and width of the obstacle respectively, and θ oi is the angle of the obstacle.
[0139] During the above growth process, that is, during the process of growing the target position, on the basis of pre-bending the selection action, considering the manufacturing error and sensor information, the pneumatic artificial muscle error correction of the self-growing robot is carried out. The flowchart for eliminating the algorithm is as Figure 11 shown.
[0140] The algorithm mainly consists of two layers of networks. The first layer of network is the pre-bending action selection network that has been trained in the previous section. The pre-bending action output by it serves as the input of the second layer of network. The second layer of network is the perception and error correction network, which is used to control the movement of the pneumatic artificial muscle to achieve error compensation.
[0141] There are two environments in the virtual scene, the virtual ideal scene and the virtual real scene. Among them, the virtual real scene is the ideal scene without any errors, and the virtual ideal scene simulates the real scene with certain errors.
[0142] S2-1, obtain the state vector of the self-growing robot as shown in Figure 8 from the environmental information.
[0143] S2-2, input the environmental information into the trained first reinforcement learning network to obtain the pre-bending action of the self-growing robot.
[0144] S2-3, add new capsules in the virtual real scene based on this action.
[0145] S2-4, add new capsules in the virtual ideal scene based on this action.
[0146] After completing step S2-3, execute step S2-5 to read the encoder and gyroscope data in the current virtual ideal scene, and perform Kalman filtering for multi-sensor fusion to obtain the state information in the current scene.
[0147] S2-6, based on the state vector of the robot in Figure 10 , receive the state information of the robot in the two virtual scenes.
[0148] Input the state information into the second reinforcement learning network, and execute step S2-7 to obtain the motion data of the pneumatic artificial muscle.
[0149] And based on this data, execute step S2-8 to control the movement of the pneumatic artificial muscle in the virtual ideal scene to achieve the purpose of eliminating errors.
[0150] Subsequently, based on the motion result of the pneumatic artificial muscle, execute step S2-9 to calculate the reward function of the second reinforcement learning network.
[0151] As shown in the following formula, obtain the reward information of reinforcement learning based on the difference between the actual end position and the theoretical end position. If the error is eliminated close to the target point after movement, a reward is given; if it is far from the target point, a penalty is given.
[0152]
[0153] where d t0 is the actual distance from the target point, and d tis the theoretical distance to the target point.
[0154] Based on the obtained reward function, step S2-10 is executed to store the experience and update the current network.
[0155] The current state of the self-growing robot, the actions taken, and the current reward function are stored as experience, and the weight information of each connection layer of the network is updated using the gradient policy algorithm.
[0156] S2-11, determine whether the capsule at the end in the virtual real scenario has reached the target point. If so, execute step S2-12 to reset the virtual ideal scenario and the virtual real scenario. Otherwise, execute step S2-1.
[0157] Before and after the movement of the pneumatic artificial muscle, calculate the deviation between the current end position of the self-growing robot and the end position of the self-growing robot in the ideal scenario respectively. If the deviation after the movement is less than the deviation before the movement, the pneumatic artificial muscle can eliminate the movement error, and the experiment is considered successful. Otherwise, the experiment fails.
[0158] Subsequently, execute step S2-13 to determine whether the success rate of the current second reinforcement learning network has reached a certain threshold. If so, execute step S2-14 to store the current model and complete the training. Otherwise, execute step S2-1 to start a new round of loop.
[0159] It should be noted that the data exchange between the virtual scenario and the reinforcement learning network uses the file system method. Based on the current action or environment information, it is overwritten and written into the txt file. There is an independent thread in the receiving segment that reads the content of the txt file and its timestamp at regular intervals. If the timestamp of the file has changed, it means that new information has been received, and the next instruction is executed based on this information. This method has lower efficiency compared to using socket or tcp communication methods, but it has higher stability, which can ensure the smooth exchange of information during the long sequential (usually more than 10 hours) simulation process and avoid algorithm crashes caused by information loss.
[0160] The following takes Figure 12 as an example to specifically elaborate on the execution steps of the algorithm, where the movement processes of the self-growing robot in the theoretical scenario and the actual scenario are respectively shown.
[0161] Figure 12 In (a) is the initial position of the self-growing robot. In Figure 12 In (b) and Figure 12 In (c), actions are generated based on the pre-bending action selection network, the pre-bending is set, and the turning of the self-growing robot is controlled. In Figure 12There is no pneumatic artificial muscle at (b) in [description], so the perception and error correction network has no output.
[0162] At Figure 12 In (c), the deviation between the theoretical value and the actual value is small, the output value of the perception and error correction network is small, and the movement of the pneumatic artificial muscle can be ignored. Subsequently, the self-growing robot grows to Figure 12 The position shown in (d) in [description]. At this time, it can be clearly observed that there is a certain deviation between the theoretical scenario and the actual scenario. Based on this deviation, the perception and error correction network outputs the movement data of the pneumatic artificial muscle, as Figure 12 shown in (e) in [description]. Through the movement of the pneumatic artificial muscle, the movement error of the self-growing robot can be effectively reduced.
[0163] In this application, the reinforcement learning model is trained in a virtual scenario and finally migrated to the actual scenario. The motion planning of the self-growing robot adopts the DDPG reinforcement learning algorithm. This algorithm includes a first reinforcement learning network and a second reinforcement learning network. The first reinforcement learning network is the pre-bending action selection network for the self-growing robot. It learns how the self-growing robot sets the pre-bending so as to navigate in a narrow space by interacting with the environment. The second reinforcement learning network is the perception and error correction network. It takes into account the pre-bending manufacturing error and sensor error of the self-growing robot and adopts the Kalman filtering algorithm of multi-sensor fusion to obtain the robot state information. The second reinforcement learning network controls the movement of the pneumatic artificial muscle to eliminate the path following error of the robot. The method provided in this application can improve the flexibility and motion accuracy of the robot in a narrow space, and improve the perception and autonomous decision-making ability of the self-growing robot. It provides a new method for the navigation of the self-growing robot.
[0164] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0165] In this article, specific examples are used to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A self-growing robot, characterized in that: The self-growing robot comprises: a stepping motor, a drive storage mechanism, a flexible body, a pre-bent structure, a pneumatic artificial muscle, a gyroscope chip and a controller; The head end of the flexible body is connected to the drive storage mechanism; the stepper motor is connected to the drive storage mechanism; the stepper motor is used to drive the flexible material in the drive storage mechanism to achieve the self-growth of the flexible body; the stepper motor is also used to obtain the length information of the self-growing robot; the gyroscope chip is used to obtain the heading angle information of the end of the flexible body; The pre-bending structure is used to adjust the growth direction of the flexible body; the pneumatic artificial muscles are arranged on both sides of the flexible body; the pneumatic artificial muscles are used to control the posture of the end of the flexible body; the controller is used to control the stepper motor, the pre-bending structure and the pneumatic artificial muscles in an environmental interactive manner according to the length information, the heading angle information and the posture to achieve steering.
2. The self-growing robot according to claim 1, characterized in that: The drive storage mechanism includes a coupling and a material reel; The stepper motor is connected to the material reel through the coupling; the flexible material is wound on the material reel.
3. The self-growing robot according to claim 2, characterized in that: The drive storage mechanism also includes an outlet plug, an air inlet, and a drive storage housing; The flexible body is connected to the outlet plug; the outlet plug and the air inlet are both arranged on the drive storage housing; the coupling and the material reel are both arranged in the drive storage housing.
4. The self-growing robot according to claim 1, characterized in that: The self-growing robot also includes an end bracket; The end bracket is arranged at the end of the flexible body; the gyroscope chip is arranged on the end bracket.
5. The self-growing robot according to claim 1, characterized in that: The flexible body is a polyethylene plastic film.
6. A self-growing robot motion planning method, characterized in that: The self-growing robot motion planning method is applied to the self-growing robot according to any one of claims 1 to 5, and the method comprises: Modeling the self-growing robot in a virtual scene; the virtual scene includes a virtual real scene and a virtual ideal scene; Acquire state information of the self-growing robot in the virtual scene; the state information includes length information of the self-growing robot and heading angle information of the end of the flexible body; According to the state information of the self-growing robot in the virtual scene, a pre-bending action is obtained by using a pre-bending action selection network; According to the pre-bending action, a new capsule is added to the end of the flexible body of the self-growing robot in the virtual scene; the flexible body is modeled by the capsule in the virtual scene; Obtain the state information of the self-growing robot in the current virtual scene and perform Kalman filtering to obtain the state information in the current virtual scene; Determine the motion data of the pneumatic artificial muscle using a perception and error correction network according to the state information in the current virtual scene; The pneumatic artificial muscle is controlled to move according to the movement data.
7. The self-growing robot motion planning method according to claim 6, characterized in that: The pre-bending action selection network includes an input layer, a first fully connected layer, a first random dropout layer, a second fully connected layer, a second random dropout layer, a third fully connected layer and an output layer which are connected in sequence.
8. The self-growing robot motion planning method according to claim 6, characterized in that: The training process of the pre-bending action selection network specifically includes: Initializing a virtual ideal scene and a flexible body of a self-growing robot; the flexible body is modeled using a capsule body; Extracting state information of the self-growing robot in the initialized virtual ideal scene as the first training sample; Selecting pre-bending action information corresponding to the first training sample using a first reinforcement learning network according to the first training sample; Controlling, in a virtual ideal scene, adding a new capsule body at the end of the flexible body according to the pre-bending action information corresponding to the first training sample to obtain a growth direction after turning and an angle after turning; the angle is the angle between the self-growing robot and the target point; Calculate the turning behavior reward according to the first training sample, the growth direction after turning, and the angle after turning; The pre-bending action selection network is determined based on the steering behavior reward, the capsule body at the end of the flexible body, and the target point.
9. The self-growing robot motion planning method according to claim 7, characterized in that: The perception and error correction network and the pre-bending action selection network have the same structure.
10. The self-growing robot motion planning method according to claim 6, characterized in that: The training process of the perception and error correction network specifically includes: Initialize the virtual scene; Obtaining state information of the self-growing robot in the virtual scene after initialization as a second training sample; Determine the pre-bending action of the self-growing robot in the initialized virtual scene using a pre-bending action selection network according to the second training sample; Adding a new capsule body at the end of the flexible body in the virtual scene according to the pre-bending action of the self-growing robot in the initialized virtual scene; Obtain the state information of the self-growing robot in the current training virtual scene and perform Kalman filtering to obtain the state information in the current training virtual scene; Determine the training motion data of the pneumatic artificial muscle by using a second reinforcement learning network according to the state information in the current training virtual scene; Controlling the pneumatic artificial muscle to move in the virtual real scene according to the training motion data to obtain a motion result; Calculate a reward function of a second reinforcement learning network according to the movement result; A perception and error correction network is determined according to the reward function, a capsule body at the end of the flexible body, and a target point.
Citation Information
Patent Citations
Underwater narrow space detection orientated flexible robot system
CN108818521A
Self-growing flexible arm gripper device
CN113103212A
Control driving system and method of self-growing soft robot for target tracking
CN118386263A
Self-growing robot and environment exploration method thereof
CN119589640A
Robotic Mobility and Construction by Growth
US20190217908A1