Robot local navigation method based on RGB pictures
Through the robot local navigation method based on RGB pictures, the Transformer model is used for imitation learning and a trained navigation model is generated, which solves the problem of insufficient dynamic adaptability of robot navigation in frequent changing environments in the prior art, and improves the generalization ability and robustness of navigation.
Patent Information
- Application Number
- CN202510223478.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
When existing robot navigation technology faces frequently changing indoor public spaces, warehouses or construction site environments, it is difficult for reinforcement learning methods to effectively deal with new obstacle layouts, terrain or dynamic factors in new environments with large differences, affecting the accuracy and safety of navigation.
The robot local navigation method based on RGB images is adopted to generate twin room data by collecting and processing room data, analyze the picture using a visual big model, select candidate centroids and map them to a three-dimensional data set, use the path planning algorithm to generate feasible paths, and use the Transformer model to imitate and learn to generate a trained navigation model.
The generalization ability of local navigation models is improved, the migration ability of navigation experience is enhanced, and the robot's robustness and adaptability to environmental changes is enhanced, and the robot's training cost in new environments is reduced.
Smart Images

Figure CN120141479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent recognition and local navigation, and specifically provides a local navigation method for a robot based on RGB images. Background Art
[0002] After obtaining the destination information, the robot first plans a generally feasible route through global path planning, and then calls the local path planner to plan the specific action strategy of the robot locally based on this route and obstacle information.
[0003] Local navigation technologies can generally be divided into two categories: traditional methods and reinforcement learning.
[0004] Traditional navigation methods (such as dynamic window, time elastic band, model predictive control, etc.) all rely on prior exploration of the environment. However, in indoor public spaces, warehouses or construction sites, the environment changes frequently, and the dynamic adaptation ability of current robot navigation technologies is difficult to meet the requirements;
[0005] At the same time, for existing reinforcement learning methods, the effect of local navigation experience transfer depends on the similarity between the old and new environments. In a new environment with large differences, the transferred experience may not be able to effectively handle completely new obstacle layouts, terrains or dynamic factors, and may even cause the robot to travel on an inappropriate path, affecting the accuracy and safety of navigation.
[0006] At the same time, after the robot enters an unfamiliar environment, it needs to rebuild the map or retrain, increasing the training and deployment costs of the robot.
[0007] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention
[0008] The purpose of the present invention is to provide a local navigation method for a robot based on RGB images to solve the technical problems proposed in the background art.
[0009] To achieve the above purpose, the present invention provides the following technical solution: A local navigation method for a robot based on RGB images, at least including the following steps:
[0010] S1: Collect data of the room and process it to generate twin room data;
[0011] S2: After selecting a room, randomly generate a collision-free robot model on the floor, control the robot camera to randomly take pictures in the room; use a vision large model to analyze and segment the pictures, select the object closest to the center of the picture as the candidate picture, and map the centroid of the mask to the three-dimensional dataset;
[0012] S3: After screening and merging the centroids, starting from the robot and using the projection points of the candidate centroids on the horizontal plane where the camera is located as the end points, use a path planning algorithm to obtain a feasible path; use the robot's optional action set to fit the path and obtain the sampling points of the path, and retain the camera pictures at the sampling points; merge the action set and the camera pictures at the sampling points as the navigation data set. In particular, mark the centroid coordinates in the start and end pictures.
[0013] S4: Based on the Transformer model, use imitation learning to learn on the robot navigation data set to obtain a trained robot navigation model.
[0014] S5: Transfer the trained model to the real robot, and after manually marking the pictures, the robot navigates autonomously.
[0015] Furthermore, the S1 at least includes the following steps:
[0016] Collect the 3D data of a real or virtual single room, segment and label each object in the room, and search for candidate labels in the digital assets that are similar to each object label.
[0017] According to the similarity between the snapshot of the candidate label and the snapshot of the object, retain the most relevant and closest label.
[0018] According to the size of each type of object, scale the digital asset corresponding to the closest label of the object, rotate and align it, and replace the original object to form a twin room data.
[0019] Furthermore, for the set of objects O=(o 1 ,…,o i ,…,o N ) in the room, where N is the number of objects in the room;
[0020] After feature embedding the object labels, retrieve the obtained candidate label set. At the same time, use the DINOv2 model to extract the features of the candidate label set snapshot and the object label snapshot, and calculate using the cosine similarity of formula (1) below. Take the label with the largest similarity as the closest label;
[0021]
[0022] where A and B are feature vectors, representing the feature vector of the object label and the feature vector of the candidate label snapshot respectively;
[0023] Find the axial bounding box of the object, scale the 3D model corresponding to the closest label according to the bounding width size, and rotate and translate the 3D model to the room where the object has been removed to form a twin room data.
[0024] Further, in S2, the Grounded-SAM model is used to perform semantic segmentation on the camera image, the object with the closest center of the object mask to the center of the image is selected, and the centroid of the object mask is calculated;
[0025] When the homogeneous coordinates of the centroid of the object mask are with a depth of and the transformation matrix between the robot and the camera is then the coordinates of the image in the robot coordinate system are K -1 z c p obj .
[0026] Further, in S3, the set of discrete actions available to the robot is set as P = (p 1 , …, p i , …, p M ), where M is the number of actions;
[0027] Then a piece of data in the navigation dataset is represented as:
[0028] {a 1 , …, a t , … a T ; I 1 , …, I t , … I T} (2)
[0029] where a t and I t are the action taken by the robot and the image captured at the t-th discrete point respectively, and T is taken as a fixed constant;
[0030] Examples of the discrete action set: move forward, turn left 45°, turn right 45°, rotate the camera upward, rotate the camera downward.
[0031] Further, S4 includes at least the following steps:
[0032] Given the training samples on the path: {a 1 , …, a t , … a T ; I 1 , …, I t , … I T}, after grouping the data, it is sent into the Transformer model, and the output action set satisfies formula (3):
[0033] a′ t-J:t = Trans(MLP(a t-J:t ), Conv(I t-J:t )) (3)
[0034] a′t-J:t After passing through different MLPs (Multi-Layer Perceptrons), relevant distances for action policy classification and prediction can be obtained;
[0035] where MLP(a t-J:t ) processes the action sequence through the MLP;
[0036] where Conv(I t-J:t ) processes the image sequence using Conv (convolutional layer);
[0037] Through the Transformer model, the robot can generate an action a′ t , that is, a′ t is the action generated by the model;
[0038] The loss function during training is formula (4):
[0039] L = L act + L track (4)
[0040] where L act = MSE(π θ (a′ t ), a t ), and π θ is the policy function;
[0041] L track = MSE(s t , d θ (s t ))), s t is the centroid of the object mask, and d θ is the relevant distance of s t between different pictures.
[0042] Compared with the prior art, the beneficial effects of the present invention are:
[0043] The present invention proposes an RGB-based local navigation method, which can improve the generalization ability of the local navigation model, is conducive to the transfer of navigation experience and the deployment of robots. At the same time, through the training of the robot in different scenarios, the robustness and adaptability of the robot to environmental changes are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 Schematic diagram of the original data and twin data of the rooms of the present invention;
[0046] Figure 2 Schematic diagram of the network framework of the present invention;
[0047] Figure 3 Schematic diagram of the local navigation of the first robot of the present invention;
[0048] Figure 4 Schematic diagram of the local navigation of the second robot of the present invention;
[0049] Figure 5 Schematic diagram of the local navigation of the third robot of the present invention. Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0051] Please refer to Figures 1 - 5 , a method for local navigation of a robot based on RGB pictures, at least including the following steps:
[0052] S1: Collect data of the room, and process it to generate twin room data;
[0053] S2: After selecting a room, randomly generate a collision-free robot model on the floor, and control the robot camera to randomly take pictures in the room; use a vision large model to analyze and segment the pictures, select the object closest to the center of the picture as the candidate picture, and map the centroid of the mask to the three-dimensional data set;
[0054] S3: After screening and merging the centroids, starting from the robot and using the projection point of the candidate centroid on the horizontal plane where the camera is located as the end point, use a path planning algorithm to obtain a feasible path; use the optional action set of the robot to fit the path and obtain the sampling points of the path, and retain the camera pictures at the sampling points; the action set and the sampling point camera pictures are merged as the navigation data set. In particular, mark the particle coordinates in the start and end pictures;
[0055] S4: Based on the Transformer model, use imitation learning to learn on the robot navigation data set to obtain a trained robot navigation model;
[0056] S5: Transfer the trained model to the real robot, and the robot autonomously navigates after the pictures are manually marked.
[0057] S1 at least includes the following steps:
[0058] Collect three-dimensional data of a real or virtual single room, segment and label each object in the room, and search for candidate labels in the digital assets that are similar to each object label;
[0059] Retain the most relevant and closest label according to the similarity between the snapshot of the candidate label and the snapshot of the object;
[0060] According to the size of each type of object, scale the digital asset corresponding to the closest label of the object, rotate and align it, and replace the original object to form twin room data.
[0061] For the set of objects O=(o 1 ,…,o i ,…,o N ) in the room, where N is the number of objects in the room;
[0062] After feature embedding of the object labels, retrieve the obtained candidate label set. At the same time, use the DINOv2 model to extract features from the candidate label set snapshot and the object label snapshot, and calculate using the cosine similarity of formula (1) below. Take the label with the maximum similarity as the closest label;
[0063]
[0064] where A and B are feature vectors, representing the feature vector of the object label and the feature vector of the candidate label snapshot respectively;
[0065] Find the axial bounding box of the object, scale the 3D model corresponding to the closest label according to the bounding width dimension, and rotate and translate the 3D model to the room where the object has been removed to form twin room data.
[0066] In S2, use the Grounded-SAM model to perform semantic segmentation on the camera image, select the object with the closest object mask center to the image center, and calculate the centroid of the object mask;
[0067] When the homogeneous coordinates of the centroid of the object mask are the depth is the transformation matrix between the robot and the camera is then the coordinates of the image in the robot coordinate system are K -1 z c p obj .
[0068] In S3, set the set of discrete actions available to the robot as P=(p 1 ,…,p i ,…,p M ), where M is the number of actions;
[0069] Then a piece of data in the navigation dataset is represented as:
[0070] {a 1 ,…,a t ,…a T ; I 1 ,…,I t ,…I T} (2)
[0071] where a t and I t are the action taken and the image captured by the robot at the t-th discrete point respectively, and T is taken as a fixed constant;
[0072] Example of the discrete action set: move forward, turn left by 45 degrees, turn right by 45 degrees, rotate the camera upward, rotate the camera downward.
[0073] S4 includes at least the following steps:
[0074] Given the training samples on the path: {a 1 ,…,a t ,…a T ; I 1 ,…,I t ,…I T}, after grouping its data, send it into the Transformer model, and the output action set satisfies formula (3):
[0075] a′ t-J:t = Trans(MLP(a t-J:t ), Conv(I t-J:t )) (3)
[0076] a′ t-J:t is the relevant distance for action policy classification and prediction after passing through different MLPs (Multi-Layer Perceptrons);
[0077] where MLP(a t-J:t ) is to process the action sequence through MLP;
[0078] where Conv(I t-J:t ) is to process the image sequence using Conv (Convolutional layer);
[0079] Through the Transformer model, the robot can generate the action a′ t at a future time, that is, a′ t is the action generated by the model;
[0080] The loss function during training is formula (4):
[0081] L = L act + Ltrack (4)
[0082] where L act = MSE(π θ (a′ t ), a t ), and π θ is a policy function;
[0083] L track = MSE(s t , d θ (s t ))), where s t is the centroid of the object mask, and d θ is the correlation distance of s t between different images.
[0084] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claimed rights.
Claims
1. A robot local navigation method based on RGB images, characterized in that: At least the following steps are included: S1: Collect room data, process it, and generate twin room data; S2: After selecting a room, randomly generate a collision-free robot model on the floor and control the robot camera to randomly take photos of the room; use the visual big model to parse and segment the picture, select the object away from the center of the picture as the candidate picture, and map the centroid of the mask to the three-dimensional data set; S3: After screening and merging the centroids, take the robot as the starting point and the projection point of the candidate centroid on the horizontal plane where the camera is located as the end point, use the path planning algorithm to obtain a feasible path; use the robot's optional action set to fit the path and obtain the sampling points of the path, and retain the camera images at the sampling points; merge the action set and the camera images of the sampling points as the navigation data set, especially mark the coordinates of the mass points in the start and end points images; S4: Based on the Transformer model, use imitation learning to learn on the robot navigation dataset to obtain a trained robot navigation model; S5: Transfer the trained model to the real robot, and let the robot navigate autonomously after manually labeling the pictures.
2. The robot local navigation method based on RGB images according to claim 1, characterized in that: The S1 at least comprises the following steps: Collect 3D data of a real or virtual single room, segment and label each object in the room, and search for candidate labels similar to each object label in the digital assets; According to the similarity between the snapshot of the candidate label and the snapshot of the object, the most relevant and closest labels are retained; According to the size of each type of object, the digital asset corresponding to the object's closest label is scaled, rotated, aligned and replaced to form twin room data.
3. The robot local navigation method based on RGB images according to claim 2, characterized in that: By calculating the object set O = (o1,…,o i ,…,o N ), N is the number of objects in the room; After embedding the object label, the candidate label set is retrieved, and the DINOv2 model is used to extract features from the candidate label set snapshot and the object label snapshot. The cosine similarity is calculated using the following formula (1), and the label with the largest similarity is taken as the most similar label. Where A and B are feature vectors, representing the feature vector of the object label and the feature vector of the candidate label snapshot respectively; Find the axial bounding box of the object, scale the 3D model corresponding to the closest label according to the bounding width, and rotate and translate the 3D model to the room where the object is removed to form twin room data.
4. The robot local navigation method based on RGB images according to claim 3 is characterized in that: In S2, the Grounded-SAM model is used to perform semantic segmentation on the camera image, the object whose object mask center is closest to the image center is selected, and the centroid of the object mask is calculated; When the homogeneous coordinates of the object mask centroid are Depth The transformation matrix between the robot and the camera is Then the coordinates of the image in the robot coordinate system are K -1 z c p obj .
5. The robot local navigation method based on RGB images according to claim 4, characterized in that: In S3, the set of optional discrete actions of the robot is set to P = (p1, ..., p i ,…,p M ), M is the number of actions; Then a piece of data in the navigation data set is represented as: {a1,…,a t ,…a T ;I1,…,I t ,…I T } (2) where a t and I t are the action taken by the robot at the tth discrete point and the image taken, respectively, and T is taken as a fixed constant.
6. The robot local navigation method based on RGB images according to claim 5, characterized in that: The S4 at least comprises the following steps: Given a training sample on a path: {a1,…,a t ,…a T ; I1,…,I t ,…I T }, after grouping the data, it is sent to the Transformer model, and the output action set satisfies formula (3): a′ t-J:t =Trans(MLP(a t-J:t ),Conv(I t-J:t )) (3) a′ t-J:t After passing through different MLPs (Multi-Layer Perceptrons), the relevant distances between action strategy classification and prediction can be obtained; Among them, MLP (a t-J:t ) is the action sequence processed by MLP; Where Conv(I t-J:t ) is to use Conv (convolution layer) to process the image sequence; Through the Transformer model, the robot can generate action a′ at the future moment t , that is, a′ t The actions generated for the model; The loss function during training is formula (4): L=L act +L track (4) Where L act =MSE(π θ (a′ t ),a t ), π θ is the strategy function; L track =MSE(s t ,d θ (s t )),s t is the object mask centroid, d θ For t The relative distances between different images.
Citation Information
Patent Citations
Autonomous navigation and obstacle avoidance system and method of indoor mobile robot
CN102359784A
Foot type robot navigation and positioning method based on variable configuration sensing device
CN113390411A
Indoor mobile robot navigation method
CN114018268A
Robot intelligent guide machining method and system
CN115147437A
Robot visual language navigation method suitable for real indoor environment
CN116518973A