Robot navigation method based on three-dimensional semantic map and large language model
Through the collaborative architecture, hierarchical planning and dynamic feedback mechanism of three-dimensional semantic maps and large language models, the problems of motion jitter and trajectory deviation of robot navigation systems in complex dynamic scenarios are solved, and the navigation success rate and robustness are improved.
Patent Information
- Application Number
- CN202510765310.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing robot navigation systems have problems of motion jitter and trajectory deviation in complex dynamic scenarios, and lack kinematic constraint optimization and inter-hierarchical dynamic feedback mechanisms, resulting in inefficient navigation.
The collaborative architecture of three-dimensional semantic maps and large language models is adopted to decompose the navigation tasks into high-level planning, middle-level planning and underlying execution goals. Through three-layer collaborative optimization and inter-level dynamic feedback mechanism, combined with three-dimensional semantic maps, SPFA algorithms and model prediction control, smooth trajectories are generated and environmental mutations are responded in real time.
It significantly improves the navigation success rate and robustness in narrow spaces and high dynamic scenarios, and improves the robot's autonomous navigation ability and task execution accuracy in complex environments.
Smart Images

Figure CN120293150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot navigation, and particularly to a robot navigation method based on a three-dimensional semantic map and a large language model. Background Art
[0002] With the rapid development of intelligent robot technology, the navigation system with a hierarchical architecture has become one of the core technologies for robots to efficiently execute complex tasks. The existing mainstream methods usually adopt a classic three-layer hierarchical structure. The high-level planning conducts path planning based on static environment information, the middle-level planning adjusts the path in real time, and the local execution layer processes dynamic obstacle avoidance according to real-time sensor data.
[0003] Although such an architecture has improved the navigation efficiency to a certain extent, it still faces significant technical bottlenecks in complex dynamic scenarios, which are specifically manifested in the following two aspects: First, there is a lack of optimization of kinematic constraints. The local planning layer usually only focuses on dynamic obstacle avoidance and real-time path adjustment, but does not deeply optimize the smoothness of the path and kinematic constraints. The path is directly transmitted to the underlying control module, resulting in trajectory jitter or unstable turning when the robot moves. Especially in narrow spaces or high-dynamic scenarios, the robot is prone to motion jitter or even deviate from the expected path due to sudden changes in trajectory curvature or lag in obstacle avoidance response. Second, there is a lack of a dynamic feedback mechanism between levels. Once the path generated by the global planning layer fails due to environmental mutations (such as non-bypassable obstacles), the local execution layer cannot trigger the replanning of the global layer, resulting in the robot falling into a local optimum or repeating ineffective actions, ultimately leading to task failure. Summary of the Invention
[0004] In view of the above-mentioned defects of the prior art, the present invention proposes a robot navigation method based on a three-dimensional semantic map and a large language model. Through a three-layer collaborative architecture of high-level planning goals, middle-level planning goals, and bottom-level execution goals, it solves the problems of motion jitter and trajectory deviation caused by the lack of optimization of kinematic constraints in traditional methods, and significantly improves the success rate of passing through narrow channels; by establishing a dynamic feedback mechanism between levels, it improves the second-level response ability to sudden obstacles; further, it improves the autonomous navigation and task execution ability of the robot in complex dynamic scenarios.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A robot navigation method based on a three-dimensional semantic map and a large language model, comprising the following steps:
[0007] S1. Task planning:
[0008] Encode the navigation task input by the user to generate a semantic vector; decompose the navigation task into three levels: high-level planning goals, middle-level planning goals, and low-level execution goals, and encode them into One-hot vectors respectively. Then, map them to the same dimension as the semantic vector through a fully connected layer to obtain the corresponding task-level encodings.
[0009] Define the high-level planning goal as determining the target area; select the target area through collaborative analysis of the 3D semantic map and the navigation task. The 3D semantic map is an advanced map representation that not only contains the 3D information of the space but also integrates the semantic understanding of various objects in the environment.
[0010] Define the middle-level planning goal as planning the travel path; obtain the travel path through the SPFA algorithm based on the 2D map of the target area.
[0011] Define the low-level execution goal as controlling the robot to complete the travel action according to the travel path.
[0012] S2. Obtain environmental features:
[0013] Obtain the point cloud data of the image of the robot's surrounding environment, extract the scene-level global features from the point cloud data through a 3D convolutional neural network, and extract the object-level fine-grained features through a feature pyramid network.
[0014] S3. Multi-scale feature association:
[0015] Fuse the task-level encodings, the semantic vector, the scene-level global features, and the object-level fine-grained features to form a unified multi-scale feature.
[0016] S4. Semantic understanding and navigation instruction generation:
[0017] Perform semantic understanding and parsing on the multi-scale feature through a large language model, and dynamically generate specific navigation instructions for the three levels of the navigation task.
[0018] Preferably, in step S1, defining the high-level planning goal as determining the target area includes:
[0019] Divide the environment into multiple candidate areas according to the semantic label set of the 3D semantic map; perform similarity matching between the semantic vector and the semantic label set of each candidate area, and calculate and generate the regional semantic score.
[0020] Based on the 2D map, use the SPFA algorithm to calculate the shortest path from the current position of the robot to the center of the candidate area; calculate and generate the path reachability score through an exponential decay function according to the length of the shortest path and the obstacle density of the shortest path.
[0021] The target area is selected by linearly weighted fusion of the regional semantic score and the path reachability score.
[0022] Preferably, in the step S1, the defined underlying execution objectives include:
[0023] Establish a robot dynamics model: The robot dynamics model establishes safety constraints according to the perceived obstacle position information; the robot dynamics model dynamically generates motion instructions according to the travel path and the safety constraints; if it is detected that the travel path is blocked during the process of controlling the robot to complete the travel action, the middle-level planning objective is re-planned.
[0024] Preferably, in the step S3, after forming the multi-scale features, an attention mechanism is used to dynamically allocate weights to the scene-level global features and the object-level fine-grained features.
[0025] Compared with the prior art, the beneficial effects of the present invention are reflected in:
[0026] 1. By adopting the hierarchical planning of the navigation tasks at three levels of the above-mentioned high-level planning objective, middle-level planning objective and underlying execution objective, the tasks at each level are decoupled and collaboratively optimized. Among them, the high-level planning objective dynamically selects the global target area based on the three-dimensional semantic map; the middle-level planning objective generates a smooth trajectory by combining path reachability and kinematic constraints; the underlying execution objective uses model predictive control to achieve dynamic obstacle avoidance. This hierarchical division effectively solves the problems of motion jitter and trajectory deviation caused by the lack of kinematic constraints in the traditional end-to-end model, and significantly improves the navigation success rate in narrow spaces and high-dynamic scenarios.
[0027] 2. By establishing a dynamic feedback mechanism between levels, the system can respond to environmental mutations in real time and trigger replanning. This cross-level dynamic feedback mechanism significantly improves the second-level response ability to sudden obstacles and ensures the robustness of the robot in complex dynamic environments.
[0028] 3. In the multi-scale feature fusion stage, by fusing task-level encoding, semantic vectors, scene-level global features and object-level fine-grained features, and dynamically allocating feature weights based on the attention mechanism, the system can autonomously adjust the focus of environmental perception according to different task requirements, making the weight allocation more in line with actual needs, thereby improving the scene adaptability of the navigation system.
[0029] 4. In addition, through in-depth analysis and reasoning of the semantic and context semantic information of the navigation task, the large language model can accurately understand the instruction intention of the user, thereby improving the accuracy of task completion. Description of the Drawings
[0030] Figure 1 This is the flowchart of the method according to Embodiment 1 of the present invention. Detailed implementation manners
[0031] In order to make the technical means, creative features, achieved purposes and functions of the invention easy to understand, the present invention will be further described below in conjunction with specific illustrations. However, the present invention is not limited to the following embodiments.
[0032] It should be noted that the structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in the art to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have any technical significance. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the functions that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed by the present invention.
[0033] Embodiment 1:
[0034] As Figure 1 shown, a robot navigation method based on a three-dimensional semantic map and a large language model includes the following steps:
[0035] S1: Encode the navigation task input by the user into a semantic vector; decompose the navigation task into three levels: high-level planning goals, middle-level planning goals, and low-level execution goals, and encode them into One-hot vectors respectively, and then map them to the same dimension as the semantic vector through a fully connected layer to obtain task level encoding. The main approach is as follows:
[0036] Use the BERT text encoder to encode the navigation task into a semantic vector , and the calculation formula is as follows:
[0037] (1)
[0038] The purpose of doing this is to make the task encoding and the semantic vector in the same dimensional space, which is convenient for subsequent feature fusion and decision-making.
[0039] In addition, a hierarchical dynamic feedback mechanism is established, and the system can respond to environmental mutations in real time and trigger replanning. This cross-hierarchical dynamic feedback mechanism significantly improves the second-level response ability to sudden obstacles and ensures the robustness of the robot in a complex dynamic environment.
[0040] S1-1: Define the high-level planning goal as determining the target area; select the target area through the collaborative analysis of the three-dimensional semantic map and the navigation task. The main approach is as follows:
[0041] First, divide the environment into several candidate areas according to the semantic label set of the three-dimensional semantic map The three - dimensional semantic map is a digital model that includes three - dimensional spatial geometric structures (such as position, shape, size), object semantic attributes (such as "kitchen", "corridor") and topological relationships. Using the semantic vector of the navigation task and the semantic label set of each candidate area to perform similarity matching and calculate the regional semantic score ( ), and the calculation formula is as follows:
[0042] (2)
[0043] where is the cosine similarity function, is the pre - trained word vector of the label . The higher the semantic score, the stronger the relevance between the area and the navigation task.
[0044] Then, based on the two - dimensional grid map , use the SPFA algorithm to calculate the shortest path length from the current position of the robot to the center of the candidate area , and combine the obstacle density on the path. Calculate the path reachability score through the exponential decay function: The calculation formula is as follows:
[0045] (3)
[0046] where is the maximum length of all possible paths in the map, , are used to balance the influence of path length and obstacle density (default , ).
[0047] Finally, linearly weight and fuse the regional semantic score and the path reachability score to select the target area D. The objective function of the target area D is as follows:
[0048] (4)
[0049] where , are weight coefficients used to balance semantic relevance and path reachability. (Default value = 0.7, )
[0050] Encapsulate the coordinate range and semantic label of the target area D into high - level planning output parameters and input them into the middle - level planning target.
[0051] S1-2: Define the middle-level planning goal as planning the travel path; according to the two-dimensional map of the target area, obtain the travel path through the SPFA algorithm. The main approach is as follows:
[0052] According to the output of the high-level planning goal Plan a feasible travel path. According to the two-dimensional map , output the set of feasible paths from the starting point to the target area D based on the SPFA algorithm , The SPFA algorithm is a shortest path algorithm based on dynamic relaxation. It optimizes the relaxation process through queue management, improves the path search efficiency, and combines curvature optimization to ensure that the path meets the kinematic constraints of the robot. Its cost function is:
[0053] (5)
[0054] where is the actual path length from the starting point to the center of the target area D, is the path curvature, is the curvature weight coefficient (default = 0.2). Subsequently, perform B-spline interpolation on the path set P to generate a smooth curve , reduce the steering jitter of the robot, and improve the motion stability.
[0055] Pack the coordinate sequence of into the middle-level planning output parameter and input it to the system's underlying execution module.
[0056] S1-3: Define the underlying execution goal as controlling the robot to complete the travel action according to the travel path. The main approach is as follows:
[0057] First, establish the dynamic model of the robot:
[0058] (6)
[0059] where represents the state of the robot at time k (including position and orientation), is the linear velocity, is the angular velocity, represents the control period.
[0060] Then, according to the real-time perceived obstacle position information, establish safety constraints to ensure that the robot does not collide with obstacles:
[0061] (7)
[0062] where is the predicted position at the i-th step within the prediction horizon, is the position of the j-th obstacle, is the safety margin (default = 0.2 m).
[0063] Generate a reference trajectory based on the output of the middle-level planning goal , the reference point at the k-th step, and optimize the control sequence within the prediction horizon of N steps :
[0064] (8)
[0065] where is the cost term imposed on the control input, which is used to limit the drastic change or excessive value of the control quantity, so as to ensure the smoothness and stability of the system.
[0066] The above model predictive control module receives the output of the middle-level planning goal , and dynamically generates motion instructions that take into account both obstacle avoidance safety and global goal reachability , ensuring that the task can still be safely executed under sudden obstacles.
[0067] If a path is detected to be permanently blocked during the bottom layer execution, trigger middle-level replanning and update .
[0068] S2: Obtain the point cloud data of the robot's surrounding environment, and extract the scene-level global features from the point cloud data through a 3D convolutional neural network Extract the object-level fine-grained features through a feature pyramid network. The main approach is as follows:
[0069] The navigation system extracts the scene-level global features from the point cloud data through a 3D convolutional neural network , and constructs a three-dimensional semantic map containing spatial topological relationships. Among them, contains 5 layers of 3D convolution, with the number of channels in each layer being [64, 128, 256, 512, 1024], the stride being [2, 2, 1, 1, 1], and the output feature dimension .
[0070] The navigation system extracts the geometric features of the object from the low-level feature map and extracts the semantic information of the object from the high-level feature map , and outputs the object-level feature .
[0071] S3: Integrate the task hierarchy encoding, navigation task encoding, scene-level global features, and object-level fine-grained features to form a unified multi-scale feature. The main approach is as follows:
[0072] First, integrate the task hierarchy encoding obtained in S1 and the navigation task semantic encoding to generate a joint context and enhance the context expression ability oriented to tasks. The calculation formula is:
[0073] (9)
[0074] where , are learnable parameters.
[0075] Project the scene-level features and the object-level features into the query space respectively to generate the corresponding query vectors and :
[0076] (10)
[0077] Use the joint task context as the key (Key) to calculate the matching degree between the two types of features and the task context through dot product. The calculation formula is as follows:
[0078] (11)
[0079] (12)
[0080] Among them, the ReLU activation function is used to filter negative correlations and highlight positive associations.
[0081] Then, generate the scene-level feature weight and the object-level feature weight through Softmax normalization. The calculation formula is as follows:
[0082] (13)
[0083] Perform weighted summation on the normalized features to generate a unified multi-scale feature representation.
[0084] (14)
[0085] Through the dual perception of language instructions and task hierarchy, the weights and can be automatically adjusted according to different tasks, making the weight distribution more in line with actual needs.
[0086] S4: Semantically understand and analyze the multi-scale features through a large language model, and dynamically generate specific navigation instructions for different hierarchical tasks. The main approach is as follows:
[0087] Semantically understand and analyze the multi-scale features obtained in S3 through a large language model (such as DeepSeek-R1) and utilize the reasoning and decision-making capabilities of the large language model to dynamically generate specific navigation instructions for different hierarchical tasks.
[0088] S4-1: High-level planning:
[0089] Task encoding T = [1, 0, 0], scene-level feature weight = 0.85 and object-level feature weight = 0.15.
[0090] Input: 3D semantic map : Point cloud data containing office environment; language instruction : "Please go to the kitchen to get a glass of water."
[0091] Decision: Calculate the semantic similarity of candidate areas (meeting room, corridor, kitchen) , and the kitchen has the highest score (0.92). The shortest path length from the current robot position to the kitchen = 4.5m, obstacle density = 0.1, and calculate according to the formula .78, total score = 0.88, select the kitchen as the target area.
[0092] Output: Target area D: Coordinate range of the kitchen area ( [5.0, 8.0], y [3.0, 6.0], z [0, 2.5]).
[0093] Navigation instruction: Global target: Move southeast to the kitchen area, coordinate range: , ,
[0094] z , pay attention to avoiding table and chair obstacles in the path.
[0095] S4-2: Middle-level path planning:
[0096] Task encoding T = [0, 1, 0], scene-level feature weight = 0.55 and object-level feature weight = 0.45.
[0097] Input: 2D map : Grid map of target area D (including obstacle distribution); Current position of the robot: x = 1.0, y = 1.0).
[0098] Decision: First, generate an initial path through the SPFA algorithm. The path has a length of 5.8m and a curvature = 0.12; The path has a length of 6.1m and a curvature = 0.07. The cost function = 5.94, = 6.24. Therefore, select , and perform B-spline interpolation on to eliminate the sharp angles at the path turning points and generate a smooth curve .
[0099] Output: : [(1.0, 1.0) (3.2, 2.6) (5.1, 4.2) (6.5, 5.0)]
[0100] Navigation instruction: Starting from the current position, move northeast to (3.2, 2.6), then move north-northeast to (5.1, 4.2), and finally reach the kitchen center (6.5, 5.0).
[0101] S4-3: Low-level execution actions:
[0102] Task encoding T = [0, 0, 1], scene-level feature weight = 0.25 and object-level feature weight = 0.75.
[0103] Input: Path output by mid-level path planning ; Obstacle information: A pedestrian is detected at (4.0, 3.5).
[0104] Decision: Through model predictive control optimization, within the prediction horizon N = 5 steps, it is detected that the distance between the obstacle and the predicted position of the robot is 0.08m, which is less than the safety margin . Add an obstacle penalty term to the optimization objective function to generate instructions to decelerate and turn left to ensure compliance with safety constraint conditions.
[0105] Output: Control instruction :
[0106] = 0.5m / s, = 0.1 rad / s] (for 2 seconds)
[0107] = 0.3 m / s, = -0.2 rad / s] (obstacle avoidance adjustment)
[0108] Navigation instruction: Execution action: Go straight at a speed of 0.5 m / s for 2 seconds, then decelerate to 0.3 m / s and turn left (angular velocity -0.2 rad / s) to avoid pedestrians ahead.
Claims
1. A robot navigation method based on a three-dimensional semantic map and a large language model, characterized in that, It includes the following steps: S1. Task Planning: Encode the navigation task input by the user to generate a semantic vector; decompose the navigation task into three levels of high-level planning objectives, middle-level planning objectives, and low-level execution objectives and encode them as One-hot vectors respectively, and then map them to the same dimension as the semantic vector through fully connected layers to obtain the corresponding task-level encodings; Define the high-level planning objective as determining the target area; Select the target area through the collaborative analysis of the 3D semantic map and the navigation task; Define the middle-level planning objective as planning the travel path; obtain the travel path through the SPFA algorithm according to the 2D map of the target area; Define the low-level execution objective as controlling the robot to complete the travel action according to the travel path; S2. Obtain Environmental Features: Obtain the point cloud data of the environmental image around the robot, extract the scene-level global features from the point cloud data through a 3D convolutional neural network, and extract the object-level fine-grained features through a feature pyramid network; S3. Multi-scale Feature Association: Fuse the task-level encoding, the semantic vector, the scene-level global features, and the object-level fine-grained features to form a unified multi-scale feature; S4. Semantic Understanding and Navigation Instruction Generation: Perform semantic understanding and parsing on the multi-scale feature through a large language model, and dynamically generate specific navigation instructions for the navigation tasks at the three levels.
2. The robot navigation method based on a three-dimensional semantic map and a large language model according to claim 1, wherein In the step S1, defining the high-level planning objective as determining the target area includes: Divide the environment into multiple candidate areas according to the semantic label set of the 3D semantic map; perform similarity matching between the semantic vector and the semantic label set of each candidate area, and calculate and generate the regional semantic score; Based on the 2D map, use the SPFA algorithm to calculate the shortest path from the current position of the robot to the center of the candidate area; calculate and generate the path reachability score through an exponential decay function according to the length of the shortest path and the obstacle density of the shortest path; Select the target area by linearly weighted fusion of the regional semantic score and the path reachability score.
3. The robot navigation method based on the three-dimensional semantic map and the large language model according to claim 2, wherein, In the step S1, defining the low-level execution objective includes: Establish a robot dynamics model: The robot dynamics model establishes safety constraints according to the perceived obstacle position information; the robot dynamics model dynamically generates motion instructions according to the travel path and the safety constraints; if it is detected that the travel path is blocked during controlling the robot to complete the travel action, re-plan the middle-level planning objective.
4. The robot navigation method based on a three-dimensional semantic map and a large language model according to claim 1, characterized in that, In the step S3, after forming the multi-scale feature, use the attention mechanism to perform dynamic weight allocation on the scene-level global features and the object-level fine-grained features.
Citation Information
Patent Citations
Robot navigation method based on semantic map and dynamic search
CN118603096A
Mobile robot indoor semantic map construction and path planning method and system
CN118896617A
Semantic navigation method based on large model, article taking and delivering method and robot
CN119188791A
Exhibition hall robot visual language navigation method based on large model
CN119309580A
System and method for navigating a vehicle using language instructions
WO2021058090A1
Cited By
Robot path planning method based on large language model guidance
CN121163528A
Quadruped robot efficient semantic perception method and device based on instance point cloud
CN121236720A
Semantic-based unmanned aerial vehicle autonomous navigation method and device, equipment and medium
CN121383998A
Humanoid robot autonomous decision-making method and system facing complex tasks
CN122363235A