Robot sensing and control method and system based on dynamic nerve symbol distance field
By introducing dynamic neural symbol distance field technology into robot control, combined with reinforcement learning and imitation learning, the problem of robots relying on the simulator punishment coefficient to avoid collision in the existing technology is solved, and the robots are able to judge and actively perceive potential risks in advance, and control stability and robustness are improved.
Patent Information
- Application Number
- CN202510435517.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing robot control methods rely on the huge penalty coefficient in the emulator to avoid collisions, lack of early judgment of potential risks, and the inconsistency between the simulation and the real machine leads to inaccurate collision detection.
A robot perception and control method based on dynamic neural symbol distance field is adopted. By acquiring the robot's grid files, constructing static and dynamic neural symbol distance fields, and combining reinforcement learning and imitation learning techniques, student strategies that can perform autonomous perception and collision avoidance are trained.
The robot is able to determine and actively perceive potential risks in advance, improve path planning and collision avoidance capabilities in compact space, reduce the problem of simulation and real machine inconsistency, and improve control stability and robustness.
Smart Images

Figure CN119927933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a robot perception and control method and system based on a dynamic neural signed distance field. Background Art
[0002] With the advancement of artificial intelligence, especially deep reinforcement learning technology, robotics has ushered in a new revolutionary change. More and more companies are beginning to use reinforcement learning for robot control. This method has shown strong control stability and robustness, pushing the human control level to a new stage. However, it also brings greater challenges to the robot's proprioception solution.
[0003] For example, the invention with the publication number CN116203945A discloses a quadruped robot motion planning method based on privileged knowledge distillation. The current mainstream legged robot motion control scheme has the following principles: 1) By taking privileged information as input, a teacher strategy is trained. Since the privileged information is taken as input, the strategy cannot be deployed. 2) Using the imitation learning method, the teacher strategy is further distilled, and the implicit latent encoding of the privileged information is predicted using the existing proprioceptive information, thereby training a student strategy for deployment.
[0004] The above solution can complete the task correctly, but it does not reflect the effect of human intuition. Unlike robots, humans can judge the approximate distance from joints to obstacles based on visual information in advance, and make a pre-judgment of the possible collision risk. The above method and the existing raw sensor data cannot complete the modeling process of the robot's super-proprioceptive perception system, which will cause the robot to passively rely on the huge penalty coefficient in the simulator to avoid collisions, rather than making a pre-judgment of potential risks based on existing information; Furthermore, since collision detection in simulators usually simplifies robots into a series of simple regular polyhedrons rather than detecting mesh surfaces, this can cause serious discrepancies between the simulation and the actual machine. Summary of the invention
[0005] The purpose of the present invention is to provide a robot perception and control method and system based on dynamic neural symbolic distance field in order to overcome the defect of the above-mentioned prior art that the robot only passively relies on the huge penalty coefficient in the simulator to avoid collision, rather than making early judgments on potential risks through existing information.
[0006] The purpose of the present invention can be achieved by the following technical solutions: A robot perception and control method based on dynamic neural signed distance field comprises the following steps: Obtain each original mesh file of the robot and pre-process the original mesh files; The preprocessed mesh file is used to construct the whole robot, and random sampling is performed to calculate the true SDF value from the sampling point to each robot joint. The true SDF value of each robot joint is fitted to obtain the static neural signed distance field. The joint positions of the robot are sampled in the simulator, and the forward kinematics results are obtained as input. The static neural signed distance field is distilled to obtain the dynamic neural signed distance field. The robot is trained in a simulator using reinforcement learning technology. During the training process, the dynamic neural signed distance field is used as the robot's external perception input, and the joint angle is used as the robot's output to obtain the teacher strategy. Distilling the teacher strategy using an imitation learning method to obtain a student strategy; Adopt student strategy for robot motion control.
[0007] Furthermore, the static neural signed distance field is: a neural network having all static grids, wherein the static grids are all 0s for the joint positions of the robot; the input of the static neural signed distance field is the three-dimensional query point coordinates, and the shortest distance from the query point coordinates to the robot surface at the 0 position is returned; The input of the dynamic neural signed distance field includes the three-dimensional query point coordinates, and the joint angles and positions of the robot, and the output is the distance value corresponding to any given robot joint posture; The dynamic neural signed distance field is used to calculate the shortest distance from a sampling point to the surface of the robot according to a set of sampling points within a certain range around the robot.
[0008] Furthermore, the acquisition process of the static neural signed distance field is specifically as follows: Read all preprocessed mesh files, get the overall size of the robot, and normalize it to the unit sphere; Randomly sample three-dimensional point coordinates in the unit sphere, and calculate the true value of the SDF from each sampling point to each robot joint; For each joint of the robot, a multi-layer perceptron is used to fit the corresponding SDF true value to obtain the neural network of the entire static grid.
[0009] Furthermore, in the fitting process of the true value of the SDF, a supervised learning method is used for fitting, and a mean square error function is used as a loss function. The expression of the mean square error function is: In the formula, is the mean square error function, N is the batch data size, is the true value of SDF, Evaluate the function for a multilayer perceptron.
[0010] Furthermore, the acquisition process of the dynamic neural signed distance field is specifically as follows: By performing the forward motion of the robot in the simulator, the joint pose of the robot is calculated. Given the homogeneous transformation matrix, the joint coordinate system and the base coordinate system, the global sampling points in the simulator are mapped back to the joint coordinate system through the coordinate system transformation formula. In the simulator, multiple sets of joint angles of the robot are sampled, and the joint positions corresponding to each set of joint angles are calculated using forward kinematics. Then, the true SDF value corresponding to each set of joint angles is calculated using the coordinate system transformation formula. After distilling the neural network corresponding to the static neural signed distance field, a new neural network is obtained, and training is performed according to the true value of the SDF corresponding to each group of joint angles to obtain a distilled neural network with joint postures as a dynamic neural signed distance field.
[0011] Furthermore, the process of acquiring the teacher strategy specifically includes the following steps: A teacher network is constructed using deep reinforcement learning technology. The input of the teacher network includes privileged input and the robot's own sensor data, and the output is the expected joint angle. The privileged input includes a first part of privileged information and a second part of privileged information. The first part of privileged information includes an elevation map of a specific area, which is compressed by a double-layer neural network and transmitted to the teacher network as input; the second part of privileged information includes a group of elevation map sampling points in the local coordinate area of the robot, and the SDF true value between each elevation map sampling point and the robot body is obtained through the dynamic neural signed distance field. The calculated SDF true value is compressed by a double-layer neural network and transmitted to the teacher network as input; The teacher network is trained using an implicit model estimation module to predict the robot state at the next moment and obtain the teacher strategy.
[0012] Furthermore, the process of acquiring the student strategy is specifically as follows: The depth map of the robot's environment is used as the policy input, and convolutional neural networks and gated recurrent networks are used to extract features from the input depth map. At the same time, the DAGGER algorithm of imitation learning is used to distill the teacher strategy to obtain the student strategy and train it based on the extracted features.
[0013] Furthermore, the preprocessing process of the original grid file includes: The simplified surface method is used to compress all triangle meshes of the original mesh file, and the redundant joints for calculating mass and collision during simulation are deleted. Only the mesh file matching the actual robot is retained to obtain the preprocessed mesh file.
[0014] Furthermore, the robot includes a quadruped robot.
[0015] The present invention also provides a robot perception and control system based on dynamic neural signed distance field, comprising a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program to execute the steps of the method described above.
[0016] Compared with the prior art, the present invention has the following advantages: (1) The present invention takes into account that the privileged information considered in the existing legged robot motion control scheme generally only includes terrain information and contact status, and it can only rely on the huge penalty coefficient in the simulator to avoid collision; This application uses dynamic neural distance field technology to characterize the robot's own proprioception, that is, the SDF true value corresponding to each elevation map sampling point in the robot's local coordinate area calculated by the dynamic neural distance field technology is added to the privileged information of the teacher strategy, which is equivalent to providing a local collision probability prior. The robot has a preliminary modeling of its own surface properties and joint shape information, which significantly improves the path planning and collision avoidance of legged robots in compact spaces; and the teacher strategy predicts the state at the next moment through the implicit model estimation module. The introduction of this scheme enables the robot to actively perceive dangerous scenes and predict situations that may lead to collisions in advance, while existing algorithms are all passive perception.
[0017] (2) The present invention is different from the existing SDF true value calculation process. The existing solution needs to use the joint angle to calculate the forward kinematics, and then use the coordinate transformation relationship to remap the sampling points to the joint coordinate system; then use the pre-trained neural network to calculate the SDF. The forward kinematics solution process of this process will be very slow when the number of parallel environments is relatively high, which will seriously slow down the simulation speed; The present invention samples multiple groups of joint angles, calculates joint positions using forward kinematics, and then uses coordinate transformation to calculate the SDF value corresponding to each group of joint angles. Finally, supervised learning technology is used to distill the truth value of this step into another new group of networks. In this way, the input becomes the global coordinates of the sampling points and the posture information of the corresponding joints to be queried, and the output is the expected joint rotation angle. The dynamic neural signed distance field obtained in this way can obtain results for any posture, and can efficiently build an accurate collision estimator in end-to-end learning tasks, with the following advantages: 1. Fast calculation speed, a large number of coordinate points can be sampled for query in a short time, which makes it possible to perform SDF calculation and accurate joint collision distance solution in real-time tasks.
[0018] 2. Better adapted to end-to-end training of reinforcement learning and parallel simulation. Due to faster calculation speed and consideration of dynamic joint changes, the Dynamic Neural Signed Distance Field (DNSDF) technology can be better integrated with parallel training of reinforcement learning.
[0019] 3. This method has broad application prospects in extreme scenarios such as narrow spaces. On legged robots or other robots with complex surface structures, DNSDF will bring precise posture planning capabilities that existing methods cannot achieve.
[0020] (3) The present invention uses an implicit model prediction module to enhance the robustness of the strategy. This method enables the robot to have corresponding stability and robustness on various unstructured surfaces and surfaces with different friction.
[0021] (4) The present invention has been extensively tested and verified on real machines, which shows that the method has strong adaptability in different scenarios.
[0022] (5) This invention realizes the active collision perception capability of a legged robot for the first time. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A schematic flow chart of a robot perception and control method based on a dynamic neural signed distance field provided in an embodiment of the present invention; Figure 2 A flowchart of a robot perception and control method based on a dynamic neural signed distance field provided in an embodiment of the present invention; Figure 3 A schematic diagram of a method flow of a dynamic neural signed distance field provided in an embodiment of the present invention; Figure 4 A schematic diagram of a deployment framework used in an experiment provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0027] Example 1 like Figure 1 As shown, this embodiment provides a robot perception and control method based on dynamic neural signed distance field, comprising the following steps: S1: Obtain each original mesh file of the robot and preprocess the original mesh files; S2: Use the preprocessed mesh file to build the whole robot, perform random sampling, calculate the true SDF value from the sampling point to each robot joint, fit the true SDF value of each robot joint, and obtain the static neural signed distance field; S3: Sample the joint positions of the robot in the simulator, obtain the forward kinematics results as input, distill the static neural signed distance field, and obtain the dynamic neural signed distance field; S4: Use reinforcement learning technology to train the robot in the simulator. During the training process, the dynamic neural signed distance field is used as the external perception input of the robot, and the joint angle is used as the output of the robot to train the teacher strategy; S5: Use imitation learning method to distill the teacher strategy to obtain the student strategy; S6: Using student strategy for robot motion control.
[0028] Specifically, the static neural signed distance field is: a neural network with all static grids, where the joint positions of the robot are all zero; the input of the static neural signed distance field is the three-dimensional query point coordinates, and the shortest distance from the query point coordinates to the robot surface at the zero position is returned; The input of the dynamic neural signed distance field includes the three-dimensional query point coordinates, as well as the robot's joint angles and positions, and the output is the distance value corresponding to any given robot joint posture; The dynamic neural signed distance field is used to calculate the shortest distance from a sampling point to the robot surface based on a set of sampling points within a certain range around the robot.
[0029] Preferably, the preprocessing process of the original grid file in step S1 includes: The simplified surface method is used to compress all triangle meshes of the original mesh file, and the redundant joints for calculating mass and collision during simulation are deleted. Only the mesh file matching the actual robot is retained to obtain the preprocessed mesh file.
[0030] In this embodiment, the preprocessing process is specifically as follows: simplify the triangular mesh, delete redundant meshes, and use them to make input files for training neural signed distance fields. Since the original mesh file exported from CAD has a huge number of triangular facets, training the signed distance field network in this way will cause storage overflow, so the original mesh file needs to be preprocessed to obtain a simplified three-dimensional mesh file. Here, the built-in simplification algorithm of Blender (three-dimensional graphics and image software Blender) is used for processing, and all triangular meshes of the original mesh file are compressed using the streamlined face method. This step can compress the original mesh file size by more than 90%, greatly reducing storage consumption. At the same time, the input URDF (a format based on XML specifications for describing robot structures) file is simplified, and redundant joints for calculating mass and collision during simulation are deleted, leaving only mesh files that match the actual machine.
[0031] Preferably, in step S2, the process of acquiring the static neural signed distance field is specifically as follows: Read all preprocessed mesh files, get the overall size of the robot, and normalize it to the unit sphere; Randomly sample the coordinates of three-dimensional points in the unit sphere and calculate the true value of the SDF from each sampling point to each robot joint; For each joint of the robot, a multi-layer perceptron is used to fit the corresponding SDF true value to obtain the neural network of the entire static grid.
[0032] In this embodiment, the above process of step S2 is specifically as follows: Using the preprocessed grid file, artificial intelligence technology is used for supervised fitting to train a static neural signed distance field. This step can be completed in several sub-steps: 1. Read all mesh files, get the overall size of the robot, and normalize it to a unit sphere for subsequent sampling.
[0033] 2. Randomly sample 3D point coordinates in the interval [-1,1], and use the mesh-to-sdf library to calculate the true value of the signed distance field (SDF) from the sampling point to each robot joint. The number of sampling points is set to 500,000. After sampling, save all the sampling points and their corresponding SDF true values.
[0034] 3. For each joint of the robot, a three-layer multi-layer perceptron (MLP) is used to fit its true SDF value. The sampling results obtained in the previous step are used to make a data set. The supervised learning method is used to fit the true value of SDF. The mean square error function is used as the loss function. Given the batch data size (N), the true value of SDF ( ), MLP calculation function , the loss function is calculated as follows: In the formula, is the mean square error function.
[0035] In step S3, the process of acquiring the dynamic neural signed distance field is specifically as follows: By performing the forward motion of the robot in the simulator, the joint pose of the robot is calculated. Given the homogeneous transformation matrix, the joint coordinate system and the base coordinate system, the global sampling points in the simulator are mapped back to the joint coordinate system through the coordinate system transformation formula. In the simulator, multiple sets of joint angles of the robot are sampled, and the joint positions corresponding to each set of joint angles are calculated using forward kinematics. Then, the true SDF value corresponding to each set of joint angles is calculated using the coordinate system transformation formula. After distilling the neural network corresponding to the static neural signed distance field, a new neural network is obtained, and trained according to the true value of the SDF corresponding to each group of joint angles to obtain the distilled neural network with joint posture as the dynamic neural signed distance field.
[0036] In this embodiment, the above process of step S3 is specifically as follows: Sampling joint positions, taking the forward kinematics results as input, distilling the static neural signed distance field to obtain the dynamic neural signed distance field. The flowchart of step S2 and step S3 is as follows: Figure 3 .
[0037] Since the SDF calculation model obtained in step S2 is a static distance field, that is, it is assumed that all joints of the robot are located at position 0, all query results are only the closest distance to position 0. If the robot's joints are not at position 0, it is necessary to use the joint angles to calculate forward kinematics, and then use the coordinate transformation relationship to remap the sampling points to the joint coordinate system. Then use the pre-trained neural network to calculate the SDF. The problem with this step is that the forward kinematics solution process will be very slow when the number of parallel environments is relatively high, which will cause it to seriously slow down the simulation speed.
[0038] Therefore, this scheme uses the forward kinematics in the simulator to directly calculate the joint pose to save time. Given the homogeneous transformation matrix T, the joint coordinate system (link), and the base coordinate system (base), the global sampling points can be mapped back to the joint coordinate system using the following formula: In the formula, is the homogeneous transformation matrix from the joint coordinate system to the base, is the homogeneous transformation matrix from the base to the world coordinate system, is the homogeneous transformation matrix from the joint coordinate system to the world coordinate system, is the coordinate of the sampling point in the joint coordinate system, is the transformation matrix from the base to the joint coordinate system, is the coordinate of the sampling point in the base coordinate system.
[0039] Since the SDF field obtained in step S2 is a static field, after the sampling point is moved to the joint coordinate system, the rotation of the joint still needs to be considered. In order to simplify this step, this scheme performs a secondary distillation on the SDF field. The specific method is to sample multiple sets of joint angles, use forward kinematics to calculate the joint position, and then use coordinate transformation to calculate the SDF value corresponding to each set of joint angles. Finally, supervised learning technology is used to distill the true value of this step into another set of new networks. After completing the secondary distillation, the input of this scheme becomes the global coordinates of the sampling point and the posture information corresponding to the joint to be queried. In this way, the obtained neural signed distance field is no longer relative to the 0 position, but the result can be obtained for any posture.
[0040] Therefore, this proposal calls this technology Dynamic Neural Signed Distance Field (DNSDF). This technology enables efficient construction of accurate collision estimators in end-to-end learning tasks. Compared with existing technologies, this method has significant advantages: 1. Fast calculation speed, a large number of coordinate points can be sampled for query in a short time, which makes it possible to perform SDF calculation and accurate joint collision distance solution in real-time tasks.
[0041] 2. Better adapted to end-to-end training of reinforcement learning and parallel simulation. Due to faster calculation speed and consideration of dynamic joint changes, DNSDF technology can be better integrated with parallel training of reinforcement learning.
[0042] 3. This method has broad application prospects in extreme scenarios such as narrow spaces. On legged robots or other robots with complex surface structures, DNSDF will bring precise posture planning capabilities that existing methods cannot achieve.
[0043] In step S4, the process of acquiring the teacher strategy specifically includes the following steps: The teacher network is constructed by using deep reinforcement learning technology. The input of the teacher network includes privileged input and the robot's own sensor data, and the output is the expected joint angle. The privileged input includes the first part of privileged information and the second part of privileged information. The first part of privileged information includes an elevation map of a specific area, which is compressed by a double-layer neural network and transmitted to the teacher network as input; the second part of privileged information includes a set of elevation map sampling points in the local coordinate area of the robot. The SDF true value between each elevation map sampling point and the robot body is obtained through the dynamic neural signed distance field. The calculated SDF true value is compressed by a double-layer neural network and transmitted to the teacher network as input. The teacher network is trained using an implicit model estimation module to predict the robot state at the next moment and obtain the teacher strategy.
[0044] In this embodiment, the above process of step S4 is specifically as follows: The quadruped robot is trained in a simulator using reinforcement learning technology. During the training process, the dynamic neural signed distance field is used as the robot's external perception input to improve its risk resistance. The implicit model estimation module is used to train a robust teacher strategy.
[0045] Specifically, we use deep reinforcement learning technology to obtain a set of vision-based end-to-end motion control strategies. The input is the depth image and the robot's own sensor (such as IMU, encoder) data, and the output is the expected joint angle. We use Isaac gym as a simulator, and run 5,500 environments in parallel. We use the proximal policy optimization (PPO) algorithm for reinforcement learning training.
[0046] In the first stage of training, the privileged input of this embodiment is divided into two parts. The first part includes an elevation map with a range of 1.6mx 1m. This information is compressed through a two-layer neural network and given to the control network as input; The second part includes a set of elevation map sampling points within the local coordinate area of the robot. The SDF values between all sampling points in this local area and the robot body will be calculated by the DNSDF module mentioned above. The calculated SDF true value will be compressed by a two-layer network in a similar way and then given to the underlying controller network as input.
[0047] The role of the DNSDF module here is that it provides a local collision probability prior, which is the first time that collision prior information is considered in a legged robot. The existing technology relies heavily on the passive collision detection performed by the simulator itself, and then imposes corresponding penalties to avoid collisions. After adding the SDF prior information, the robot has a pre-modeling of its own surface properties and joint shape information, which significantly improves the path planning and collision avoidance of legged robots in compact spaces.
[0048] The training framework of this part is as follows Figure 2 As shown in the figure, the underlying actor (the policy function responsible for generating actions) uses a framework called implicit model prediction to complete training. The significance of this method is to predict the state at the next moment in order to cope with stronger external interference. The training process is completed using an independent encoder. After the PPO update, the following formula is used for additional updates. Experiments have shown that this method can ensure the stable convergence of the strategy under a large range of random disturbances. This is due to the prediction of external interference to the strategy and the clustering of body perception, which ensures the motion robustness of the strategy for different unstructured surfaces.
[0049] The loss function of the proximal policy optimization (PPO) algorithm is calculated as: In the formula, is a hidden variable predicted based on historical information, is the hidden variable predicted for the next moment, In order to cancel the gradient, the gradient information here is only returned by the gradient obtained by the hidden variable predicted based on the historical information.
[0050] Preferably, the process of acquiring the student strategy in step S5 is specifically as follows: The specific process of acquiring student strategies is as follows: The depth map of the robot's environment is used as the policy input, and convolutional neural networks and gated recurrent networks are used to extract features from the input depth map. At the same time, the DAGGER algorithm of imitation learning is used to distill the teacher strategy to obtain the student strategy and train it based on the extracted features.
[0051] In this embodiment, the above process of step S5 is specifically as follows: The teacher strategy is distilled using the imitation learning method to obtain the student strategy, with the input being a depth image.
[0052] The strategy obtained in step S4 cannot be sim2real (Simulation to Reality, which means migrating the algorithms, strategies or models trained or developed in the simulation environment to the real physical world). This is because the input of the strategy contains privileged information, which is not available in the real machine. In order to be able to perform sim2real, this embodiment performs another step of strategy distillation after completing step S4, as shown in the process Figure 2 As shown in the figure, when training the distillation strategy, the front-end depth map is used as the strategy input, a lightweight convolutional neural network (CNN) and a gated recurrent network (GRU) are used to extract features, and the DAGGER algorithm of imitation learning is used to distill the teacher strategy, and the output results of the teacher strategy are distilled to the student strategy.
[0053] This approach is reasonable for the following reasons: 1. For elevation map prediction, with the help of the 3D scene information provided by the depth camera, it is possible to reconstruct the shape and structure of the environment around the robot.
[0054] 2. For the SDF module, the depth map can provide scene information of the environment. At the same time, with the help of the state estimation function of the underlying GRU module, it can theoretically predict the SDF value of the local scene around the current robot. Therefore, this solution is reasonable.
[0055] In this embodiment, the processing process of step S6 is specifically as follows: Export the student strategy and perform sim2real testing. The deployment flowchart is as follows: Figure 4 As shown in the figure, during the deployment phase, ROS2 is used to allow an independent desktop computer to communicate with the robot's own control motherboard. A 3D printed bracket is designed on the front of the robot to fix the realsense (depth camera RealSense). The visual algorithm runs at a frequency of 10Hz and the underlying control strategy runs at a frequency of 50Hz.
[0056] Example 2 This embodiment provides a robot perception and control system based on a dynamic neural signed distance field, including a memory and a processor, the memory stores a computer program, and the processor calls the computer program to execute the steps of the robot perception and control method based on a dynamic neural signed distance field as described in Example 1.
[0057] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A robot perception and control method based on dynamic neural signed distance field, characterized in that: The following steps are involved: Obtain each original mesh file of the robot and pre-process the original mesh files; The preprocessed mesh file is used to construct the whole robot, and random sampling is performed to calculate the true SDF value from the sampling point to each robot joint. The true SDF value of each robot joint is fitted to obtain the static neural signed distance field. The joint positions of the robot are sampled in the simulator, and the forward kinematics results are obtained as input. The static neural signed distance field is distilled to obtain the dynamic neural signed distance field. The robot is trained in a simulator using reinforcement learning technology. During the training process, the dynamic neural signed distance field is used as the robot's external perception input, and the joint angle is used as the robot's output to obtain the teacher strategy. Distilling the teacher strategy using an imitation learning method to obtain a student strategy; Adopt student strategy for robot motion control.
2. A robot perception and control method based on dynamic neural signed distance field according to claim 1, characterized in that: The static neural signed distance field is: a neural network with all static grids, wherein the static grids are all 0 for the joint positions of the robot; the input of the static neural signed distance field is the three-dimensional query point coordinates, and the shortest distance from the query point coordinates to the robot surface at the 0 position is returned; The input of the dynamic neural signed distance field includes the three-dimensional query point coordinates, and the joint angles and positions of the robot, and the output is the distance value corresponding to any given robot joint posture; The dynamic neural signed distance field is used to calculate the shortest distance from a sampling point to the surface of the robot according to a set of sampling points within a certain range around the robot.
3. The robot perception and control method based on dynamic neural signed distance field according to claim 1 is characterized in that: The acquisition process of the static neural signed distance field is specifically as follows: Read all preprocessed mesh files, get the overall size of the robot, and normalize it to the unit sphere; Randomly sample three-dimensional point coordinates in the unit sphere, and calculate the true value of the SDF from each sampling point to each robot joint; For each joint of the robot, a multi-layer perceptron is used to fit the corresponding SDF true value to obtain the neural network of the entire static grid.
4. A robot perception and control method based on dynamic neural signed distance field according to claim 3, characterized in that: In the fitting process of the true value of the SDF, a supervised learning method is used for fitting, and a mean square error function is used as a loss function. The expression of the mean square error function is: In the formula, is the mean square error function, N is the batch data size, is the true value of SDF, Evaluate the function for a multilayer perceptron.
5. The robot perception and control method based on dynamic neural signed distance field according to claim 1 is characterized in that: The acquisition process of the dynamic neural signed distance field is specifically as follows: By performing the forward motion of the robot in the simulator, the joint pose of the robot is calculated. Given the homogeneous transformation matrix, the joint coordinate system and the base coordinate system, the global sampling points in the simulator are mapped back to the joint coordinate system through the coordinate system transformation formula. In the simulator, multiple sets of joint angles of the robot are sampled, and the joint positions corresponding to each set of joint angles are calculated using forward kinematics. Then, the true SDF value corresponding to each set of joint angles is calculated using the coordinate system transformation formula. After distilling the neural network corresponding to the static neural signed distance field, a new neural network is obtained, and training is performed according to the true value of the SDF corresponding to each group of joint angles to obtain a distilled neural network with joint postures as a dynamic neural signed distance field.
6. The robot perception and control method based on dynamic neural signed distance field according to claim 1 is characterized in that: The acquisition process of the teacher strategy specifically includes the following steps: A teacher network is constructed using deep reinforcement learning technology. The input of the teacher network includes privileged input and the robot's own sensor data, and the output is the expected joint angle. The privileged input includes a first part of privileged information and a second part of privileged information. The first part of privileged information includes an elevation map of a specific area, which is compressed by a double-layer neural network and transmitted to the teacher network as input; the second part of privileged information includes a group of elevation map sampling points in the local coordinate area of the robot, and the SDF true value between each elevation map sampling point and the robot body is obtained through the dynamic neural signed distance field. The calculated SDF true value is compressed by a double-layer neural network and transmitted to the teacher network as input; The teacher network is trained using an implicit model estimation module to predict the robot state at the next moment and obtain the teacher strategy.
7. The robot perception and control method based on dynamic neural signed distance field according to claim 1 is characterized in that: The process of obtaining the student strategy is specifically as follows: The depth map of the robot's environment is used as the policy input, and convolutional neural networks and gated recurrent networks are used to extract features from the input depth map. At the same time, the DAGGER algorithm of imitation learning is used to distill the teacher strategy to obtain the student strategy and train it based on the extracted features.
8. The robot perception and control method based on dynamic neural signed distance field according to claim 1 is characterized in that: The preprocessing process of the original grid file includes: The simplified surface method is used to compress all triangle meshes of the original mesh file, and the redundant joints for calculating mass and collision during simulation are deleted. Only the mesh file matching the actual robot is retained to obtain the preprocessed mesh file.
9. The robot perception and control method based on dynamic neural signed distance field according to claim 1 is characterized in that: The robot includes a quadruped robot.
10. A robot perception and control system based on dynamic neural signed distance field, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor calls the computer program to execute the steps of any one of the methods according to claims 1 to 9.
Citation Information
Patent Citations
Quadruped robot motion planning method based on privileged knowledge distillation
CN116203945A
Footprint prediction method, footprint prediction device and humanoid robot
CN112084853A
Robot three-dimensional measurement path planning method based on deep reinforcement learning
CN116604571A
Mechanical arm collision detection method based on SDF function
CN117601124A
Robot reinforcement learning control method and system based on gating circulation unit
CN119536333A
Cited By
Hexapod robot leg and arm multiplexing control method and device based on reinforcement learning
CN120993711A
Progressive motion interpolation training method for whole-body motion learning of humanoid robot
CN121267943A
Progressive motion interpolation training method for whole body motion learning of humanoid robot
CN121267943B