Robot semantic navigation method and system, terminal and storage medium
By constructing semantic maps and using large language models to process natural language commands, the problems of low navigation efficiency and success rate in the existing technology are solved, and efficient and accurate navigation and precise positioning are achieved in complex environments.
Patent Information
- Application Number
- CN202411881045.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art cannot meet practical application requirements in terms of navigation efficiency and success rate, and it is difficult to perform accurate positioning and navigation tasks with high spatial accuracy in complex environments.
By building a semantic map containing semantic information, and using a large language model to convert natural language commands into instructions that the robot can recognize, combining the semantic map to determine the position point of the target navigation object, and performing path planning to ensure that the robot accurately reaches the target point.
It realizes efficient and accurate navigation and precise positioning in complex environments, meets the needs of high spatial accuracy, and improves navigation efficiency and success rate.
Smart Images

Figure CN120063242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics engineering, and particularly to a robot semantic navigation method, system, terminal, and storage medium. Background Art
[0002] Currently, existing navigation methods prioritize open exploration and highly value extracting semantic information, but they tend to overlook the high spatial accuracy required for inspection tasks, and cannot ensure that the robot can accurately reach the target point, nor can it perform precise positioning and navigation tasks in complex environments. Moreover, traditional navigation methods relying on deep neural networks or reinforcement learning require a large amount of training resources and large datasets, making the existing navigation technologies unable to meet the actual application requirements in terms of navigation efficiency and success rate.
[0003] Therefore, there are deficiencies in the existing technology and it needs to be improved and developed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a robot semantic navigation method, system, terminal, and storage medium, which can meet the requirements of high spatial accuracy, ensure that the robot can accurately reach the target point, and thus can perform precise positioning and navigation tasks in complex environments, aiming at the above-mentioned deficiencies of the existing technology.
[0005] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0006] A robot semantic navigation method, wherein the method includes:
[0007] Input the received natural language command and the pre-constructed semantic map into a pre-trained large language model, so that the large language model extracts the target navigation object in the natural language command, and matches the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map;
[0008] Perform path planning based on the current position of the robot and the target position point to obtain a corresponding target path, so that the robot navigates to the target position point based on the target path;
[0009] Wherein, the semantic map is a global map corresponding to the current spatial environment and containing semantic information.
[0010] In one implementation, before inputting the natural language command and the pre-constructed semantic map into the pre-trained large language model, it further includes:
[0011] Perform text processing on the pre-constructed semantic map to obtain a texturized semantic map;
[0012] Among them, inputting the natural language command and the pre-constructed semantic map into the pre-trained large language model includes:
[0013] Inputting the natural language command and the text-based semantic map into the pre-trained large language model.
[0014] In one implementation, during the process of constructing a semantic map containing semantic information corresponding to the current spatial environment, it further includes:
[0015] Based on a preset laser mapping algorithm, and using a preset lidar device to scan the current spatial environment to construct a corresponding three-dimensional point cloud;
[0016] At the same scanning time as the preset lidar device, using a preset imaging device to take pictures of the current spatial environment to obtain a corresponding RGB image;
[0017] Using a pre-trained semantic segmentation model to perform semantic segmentation on the RGB image to obtain a corresponding segmentation mask;
[0018] Matching the three-dimensional point cloud with the segmentation mask to determine the target point cloud, and assigning corresponding labels to the target point cloud in the three-dimensional point cloud to obtain a semantic map containing semantic information.
[0019] In one implementation, matching the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map includes:
[0020] Matching the target navigation object with the semantic map to extract the target semantic information of the target navigation object from the semantic map; wherein, the target semantic information includes the position information and size information of the target navigation object;
[0021] Based on the semantic information of the target navigation object and the current position of the robot, determine the target position point of the target navigation object on the semantic map.
[0022] In one implementation, based on the semantic information of the target navigation object and the current position of the robot, determining the target position point of the target navigation object on the semantic map includes:
[0023] Select corresponding position points at the edge of the target navigation object to obtain candidate position points of the target navigation object on the semantic map;
[0024] Based on the semantic information of the target navigation object, calculate the cost value gradient from the center of the target navigation object to each candidate position point to obtain the target cost value gradient corresponding to each candidate position point;
[0025] Calculate the Euclidean distance between the current position of the robot and each of the candidate position points to obtain the target distance corresponding to each candidate position point;
[0026] Determine the score of each candidate position point based on the target cost value gradient and the target distance corresponding to each candidate position point;
[0027] Compare the scores of each candidate position point to determine the candidate position point with the highest score, and determine the candidate position point with the highest score as the target position point of the target navigation object on the semantic map.
[0028] In one implementation, the determining the score of each candidate position point based on the target cost value gradient and the target distance corresponding to each candidate position point includes:
[0029] Use a preset scoring formula and determine the score of each candidate position point based on the target cost value gradient and the target distance corresponding to each candidate position point;
[0030] Wherein, the preset scoring formula is:
[0031]
[0032] Wherein, S(θ) represents the score, w g represents the gradient weight parameter, w d represents the distance weight parameter, G(θ) represents the target cost value gradient, and D(θ) represents the target distance.
[0033] In one implementation, after performing path planning based on the current position of the robot and the target position point to obtain a corresponding target path so that the robot navigates to the target position point based on the target path, it further includes:
[0034] When the received natural language command is a command carrying a corresponding detection task, input the target image captured at the target position point into a pre-trained vision-language model, so that the vision-language model performs image understanding based on the target image to obtain a picture description output by the vision-language model, and send the picture description to the large language model so that the large language model outputs a corresponding detection result based on the picture description.
[0035] The present invention also discloses a robot semantic navigation system, wherein the system includes:
[0036] A command processing module, configured to input the received natural language command and the pre-constructed semantic map into a pre-trained large language model, so that the large language model extracts the target navigation object in the natural language command, and matches the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map;
[0037] A path planning module, configured to perform path planning based on the current position of the robot and the target position point to obtain a corresponding target path, so that the robot navigates to the target position point based on the target path;
[0038] Wherein, the semantic map is a global map corresponding to the current spatial environment and containing semantic information.
[0039] The present invention also discloses a terminal, which includes: a memory, a processor, and a robot semantic navigation program stored on the memory and executable on the processor. When the robot semantic navigation program is executed by the processor, the steps of the robot semantic navigation method described above are implemented.
[0040] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the robot semantic navigation method described above.
[0041] A robot semantic navigation method, system, terminal and storage medium provided by the present invention. The robot semantic navigation method includes: inputting the received natural language command and the pre-constructed semantic map into a pre-trained large language model, so that the large language model extracts the target navigation object in the natural language command, and matches the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map; performing path planning based on the current position of the robot and the target position point to obtain a corresponding target path, so that the robot navigates to the target position point based on the target path; wherein, the semantic map is a global map corresponding to the current spatial environment and containing semantic information. It can be seen that by constructing a semantic map containing semantic information, the present invention can meet the requirements of high spatial accuracy, and use the large language model to convert natural language commands into instructions that can be recognized and executed by the robot, so that the robot can understand human natural language, extract the target navigation object in the natural language command through the large language model, then combine the semantic map to determine the target position point of the target navigation object on the semantic map, and finally perform path planning based on the current position and the target position point of the robot to ensure that the robot can reach the target position point efficiently and accurately, and thus can perform precise positioning and navigation tasks in a complex environment. Description of the Drawings
[0042] Figure 1 is a flowchart of a preferred embodiment of the robot semantic navigation method in the present invention;
[0043] Figure 2 is a specific semantic mapping schematic diagram disclosed by the present invention;
[0044] Figure 3 is a schematic diagram of semantic maps from different perspectives disclosed by the present invention;
[0045] Figure 4 is a schematic diagram of the internal working process of the robot completing the detection task under the guidance of the large language model disclosed by the present invention;
[0046] Figure 5 is a functional principle block diagram of a preferred embodiment of the robot semantic navigation system in the present invention;
[0047] Figure 6 is a specific schematic diagram of the robot semantic navigation system disclosed by the present invention;
[0048] Figure 7 is a specific initialization prompt analysis schematic diagram disclosed by the present invention;
[0049] Figure 8 is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. Detailed implementation manners
[0050] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the present invention will be further described in detail below with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0051] Please refer to Figure 1 , Figure 1 is a flowchart of the robot semantic navigation method in the present invention. As Figure 1 shown, the robot semantic navigation method described in the embodiments of the present invention includes:
[0052] Step S11: Input the received natural language command and the pre-constructed semantic map into a pre-trained large language model, so that the large language model extracts the target navigation object in the natural language command, and matches the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map; wherein, the semantic map is a global map corresponding to the current spatial environment and containing semantic information.
[0053] In this embodiment, the user issues a natural language command to the robot. Then, the robot receives the natural language command issued by the user, inputs the natural language command and the pre-constructed semantic map into a pre-trained large language model, so as to use the large language model to convert the natural language command into an instruction recognizable and executable by the robot, extract the target navigation object in the natural language command, and then match the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map.
[0054] In this embodiment, before inputting the natural language command and the pre-constructed semantic map into the pre-trained large language model, it may specifically further include: performing text processing on the pre-constructed semantic map to obtain a text-based semantic map, and then inputting the natural language command and the text-based semantic map into the pre-trained large language model. It can be understood that encoding the semantic map into text and providing it to the LLM (Large Language Model) in the form of a parameter configuration file can facilitate the LLM's understanding of the environment, and by adding concise and clear task details and context descriptions to the initial prompt input to the large language model, the LLM can generate more accurate responses.
[0055] It should be noted that the semantic map is a pre-constructed global map corresponding to the current spatial environment and containing semantic information. And in the process of constructing the semantic map containing semantic information corresponding to the current spatial environment, it may specifically include: based on a preset lidar mapping algorithm, using a preset lidar device to scan the current spatial environment to construct a corresponding three-dimensional point cloud; at the same scanning time as the preset lidar device, using a preset imaging device to take pictures of the current spatial environment to obtain a corresponding RGB image; using a pre-trained semantic segmentation model to perform semantic segmentation on the RGB image to obtain a corresponding segmentation mask; matching the three-dimensional point cloud with the segmentation mask to determine the target point cloud, and assigning corresponding labels to the target point cloud in the three-dimensional point cloud to obtain a semantic map containing semantic information. It can be understood that a semantic map containing semantic information is constructed by combining lidar and vision, that is, a semantic map containing semantic information is constructed by combining the point cloud data generated by lidar scanning and the RGB image data. Among them, rich semantic information can be obtained from the RGB image data, and lidar can provide stable depth estimation, so as to ensure that the map has certain spatial details and can also associate the objects in the map with meaningful labels, which is crucial for downstream navigation tasks.
[0056] For example, as shown in Figure 2 for a frame of point cloud L P, first through the extrinsic parameter matrix of the lidar and the cameraC T L Convert the point cloud L P from the lidar coordinate system to the camera coordinate system to obtain C P, and then use the camera intrinsic matrix K to project the point cloud C P from the camera coordinate system to the image coordinate system to obtain the point cloud in the image coordinate system I P. At the same time, use a pre-trained semantic segmentation model to perform semantic segmentation on the RGB image captured at the same time as the point cloud L P to generate a corresponding segmentation mask, and then match the point cloud I P with the segmentation mask. If any point falls within the segmentation mask, the points in the point cloud L P will be assigned corresponding labels, and then a global map containing semantic information will be generated in combination with the laser mapping algorithm. The semantic maps containing semantic information from different perspectives are shown in Figure 3 as follows. Among them, the purple, bright green, dark green, orange, and brown point clouds represent the TV, whiteboard, chair, keyboard, and monitor respectively.
[0057] In this embodiment, matching the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map may specifically include: matching the target navigation object with the semantic map to extract the target semantic information of the target navigation object; wherein, the target semantic information includes the position information and size information of the target navigation object; determining the target position point of the target navigation object on the semantic map based on the semantic information of the target navigation object and the current position of the robot. It can be understood that the target semantic information of the object (target navigation object) and the current position of the robot are used to calculate the optimal target position point for the robot. That is to say, the target position point can be determined by analyzing the cost value gradient from the center to the edge of the object and considering the relative position of the robot, so as to ensure that the robot approaches the object from the direction of maximizing safety and efficiency. Considering the cost value gradient when determining the optimal target position point can solve the problem of difficult to accurately locate the target position point in the case of object occlusion or space limitation.
[0058] It should be noted that in the top-down semantic map, the robot often fails to accurately locate the target point (target position point). For example, when the target is against the wall, the robot sometimes navigates to the other side. Although from the top-down map, the robot does reach near the target object, this is not the correct target point. To solve this problem, the method of using semantic information and the gradient of the target cost value can be utilized to ensure the accuracy of the inspection and positioning. That is, corresponding position points are selected at the edge of the target navigation object to obtain the candidate position points of the target navigation object on the semantic map; based on the semantic information of the target navigation object, the cost value gradient from the center of the target navigation object to each candidate position point is calculated to obtain the target cost value gradient corresponding to each candidate position point; the Euclidean distance between the current position of the robot and each candidate position point is calculated to obtain the target distance corresponding to each candidate position point; based on the target cost value gradient and the target distance corresponding to each candidate position point, the score of each candidate position point is determined; the scores of each candidate position point are compared to determine the candidate position point with the highest score, and the candidate position point with the highest score is determined as the target position point of the target navigation object on the semantic map.
[0059] For example, an object instance in the current spatial environment can be represented by its spatial coordinates and size. For example, each object is modeled as a rectangular area with its center coordinates being (x obj , y θobj ), and its size can be represented by x size and y size . Since the goal of the robot is to approach the target from the direction with the steepest cost gradient while considering the relative position between the current position of the robot and the target object, the best target position point can be selected by performing a 360-degree scan and evaluation on the candidate points at the edge of the object. That is, for each candidate position point, two main parameters, namely the cost value gradient from the center of the object to the candidate point and the Euclidean distance between the robot and the candidate point, can be calculated, and then the best candidate position point is selected as the target position point based on the cost value gradient and the Euclidean distance corresponding to each candidate position point. Among them, the magnitude of the cost value gradient can represent safety and navigability, and the Euclidean distance can represent proximity and efficiency.
[0060] It should be noted that the calculation of the cost value gradient in the cost map is performed along each direction starting from the center of the object, that is:
[0061]
[0062] where G(θ) represents the cost value gradient calculated along the direction θ starting from the center of the object, and Cost(x θ , y θ ) represents the cost map value of the object edge along the direction θ, and Cost(x obj , yθobj ) represents the cost map value at the center of the object, d θ represents the Euclidean distance between the center of the object and the edge point, d safe represents the preset safety distance, r represents the resolution of the cost map, with the unit of meters per grid cell, and d safe is sized according to the actual size of the robot in advance.
[0063] And, the distance along the θ direction between the current position of the robot and the candidate position point can be calculated through the Euclidean distance formula, that is:
[0064] D(θ) = [(x θ - x robot ) 2 + (y θ - y robot ) 2 · r 2 ;
[0065] Among them, D(θ) represents the physical distance along the θ direction between the current position of the robot and the candidate position point. The smaller this value is, the shorter the moving distance. (x θ , y θ ) represents the coordinates of the candidate position point along the θ direction, and (x robot , y robot ) represents the coordinates of the current position of the robot.
[0066] Among them, the scores of each candidate position point are determined based on the target cost value gradient and the target distance corresponding to each candidate position point. Specifically, it can include: using a preset scoring formula and determining the scores of each candidate position point based on the target cost value gradient and the target distance corresponding to each candidate position point, so as to be able to combine the influences of the gradient and the distance, comprehensively consider the cost value gradient and the distance of the candidate position point, and select the best target position point.
[0067] Among them, the preset scoring formula is:
[0068]
[0069] Among them, S(θ) represents the score, w g represents the gradient weight parameter, w d represents the distance weight parameter, G(θ) represents the target cost value gradient, and D(θ) represents the target distance. w g and w d are respectively used to balance the importance of the gradient and the distance. The importance of the gradient and the distance can be adjusted by adjusting w g and w dWhether to prioritize safety or efficiency. If the gradient value is high, it indicates that safety is prioritized. If the distance is short, it indicates that efficiency is prioritized. Then, the candidate position point with the highest score is determined as the best target position point for the robot to approach the object. At the same time, the normal direction of the candidate position point towards the center point of the object is used as the orientation of the robot's target position point.
[0070] Step S12: Based on the current position of the robot and the target position point, perform path planning to obtain a corresponding target path so that the robot can navigate to the target position point based on the target path.
[0071] In this embodiment, after extracting the target navigation object in the natural language command through the large language model and determining the target position point of the target navigation object on the semantic map in combination with the semantic map containing semantic information, path planning can be performed based on the current position of the robot and the target position point to obtain a corresponding target path so that the robot can navigate to the target position point based on the target path.
[0072] In this embodiment, after performing path planning based on the current position of the robot and the target position point to obtain a corresponding target path so that the robot can navigate to the target position point based on the target path, it may further specifically include: when the received natural language command is a command carrying a corresponding detection task, input the target image captured at the target position point into a pre-trained vision-language model, so that the vision-language model can perform image understanding based on the target image to obtain the picture description output by the vision-language model, and send the picture description to the large language model so that the large language model can output a corresponding detection result based on the picture description.
[0073] For example, see Figure 4As shown in the figure, when the robot is an inspection robot for performing inspection tasks, the user issues an inspection command (user instruction) in the form of natural language to the inspection robot, such as "Go to the table to check if the computer is on the table". The robot obtains the user instruction and sends it to the large language model. Through the large language model, the inspection robot can convert the inspection command into a detection primitive recognizable by itself, extract the target navigation object in the inspection command, and output it in the JSON structure format, and then output a corresponding response, that is, "Currently, I will go to the table to check if the computer is on it". Furthermore, the large language model combines the semantic map containing semantic information to determine the target position point of the target navigation object on the semantic map, and then performs path planning based on the current position of the robot and the target position point to obtain the corresponding target path. The robot navigates to the target position point of the object based on the target path and takes a photo at the target position point, and sends the target image taken at the target position point to the pre-trained vision language model, so that the vision language model can perform image understanding based on the target image to obtain the picture description output by the vision language model, and send the picture description to the large language model so that the large language model can output the corresponding detection result based on the picture description, that is, "Yes, the computer, keyboard and mouse are all on the white table". That is to say, the LLM is used to interpret and extract the target navigation object in the natural language command, and at the same time determine whether a detection task needs to be performed, that is, through the large language model, a structured detection task understandable by the robot can be generated, so as to realize the autonomous navigation of the robot and call the VLM (Vision Language Model) for image understanding, so as to complete the corresponding inspection. It can be seen that by using the LLM as a zero-shot advanced task planner and combining it with the traditional low-level task planning framework, a more general navigation scheme can be formed, solving the problems of requiring a large amount of training resources and lacking generality in actual situations, realizing the generality, stability and ease of use of the inspection robot, significantly improving the task success rate, navigation efficiency and stability, and can be used in the actual application scenarios of various detection tasks with only minimal initialization.
[0074] It can be seen that in the embodiment of the present invention, by constructing a semantic map containing semantic information, the requirement of high spatial accuracy can be met, and the large language model can be used to convert the natural language command into an instruction recognizable and executable by the robot, so that the robot can understand human natural language. The large language model extracts the target navigation object in the natural language command, and then combines the semantic map to determine the target position point of the target navigation object on the semantic map. Finally, path planning is performed based on the current position of the robot and the target position point to ensure that the robot can reach the target position point efficiently and accurately, and thus can perform precise positioning and navigation tasks in a complex environment.
[0075] In one embodiment, as Figure 5 shown, based on the above-mentioned robot semantic navigation method, the present invention also correspondingly provides a robot semantic navigation system, including:
[0076] A command processing module 11, configured to input the received natural language command and the pre-constructed semantic map into a pre-trained large language model, so that the large language model extracts the target navigation object in the natural language command, and matches the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map; wherein, the semantic map is a global map corresponding to the current spatial environment and containing semantic information;
[0077] A path planning module 12, configured to perform path planning based on the current position of the robot and the target position point to obtain a corresponding target path, so that the robot navigates to the target position point based on the target path.
[0078] For example, as shown in Figure 6 shown, the robot semantic navigation system can be composed of two main parts, namely a perception part and a task planning part. In the perception part, the construction of a semantic map containing semantic information is mainly realized, that is, semantic segmentation is performed on the RGB image to obtain a semantic segmentation mask, and the semantic segmentation mask is used to perform semantic association on the 3D lidar scan-generated point cloud, so as to construct a semantic map containing semantic information. After text processing of the semantic map, it is input into the large language model, and an initialization prompt is input into the large language model, that is, "You are a detection robot, and your function is to read the semantic layer of the map and output navigation instructions in JSON structure format according to the user's instructions, and generate simple answers according to the received image description", so that the large language model can generate more accurate responses. For a specific analysis of the initialization prompt, see Figure 7 shown, the initialization prompt describes the tasks and capabilities of the robot, that is, the tasks of the robot: "You are a detection robot responsible for approaching and inspecting an object"; the capabilities of the robot: according to the user's language instructions (natural language commands) and the provided environmental information (semantic map), output structured detection raw data, that is, generate an answer in JSON format according to the user's request, and only provide JSON data.
[0079] Among them, the function str(get_semantic_info(map.yaml)) provides information about the objects in the working environment, that is, represents the semantic map containing semantic information in a text-based manner. The semantic information after the map is converted into text is as follows:
[0080] -id: The specific number of the object;
[0081] - Name: Name of the object:
[0082] - Location: Location of the object on the map;
[0083] - Size: Two-dimensional size of the object.
[0084] A task example is provided for the robot. That is, if the natural language command issued by the user is "Go to the largest table and check what's on the table", it is necessary to find the largest table from the available data and output it in the JSON structure format.
[0085] Moreover, in the task planning section, accurate positioning and navigation are mainly achieved through the command acquisition module, the command processing module, and the path planning module, and the corresponding inspection tasks are completed. For example, when a language task issued by the user is obtained, that is, "Check if the door is closed?" and this language task is input into the large language model. The large language model (LLM) is used to interpret and extract the target navigation object in the natural language command, and at the same time, it is determined whether the inspection task needs to be executed, and the output of the LLM can be constrained by structured parameters to ensure that the LLM can correctly convert the natural language command into a detection primitive that the robot can execute in real time. Then, the large language model combines the semantic map containing semantic information to determine the target position point of the target navigation object on the semantic map. This target position point is the target point with the highest score obtained after optimization. Then, based on the current position of the robot and the target position point, path planning is performed to obtain the corresponding target path so that the robot can navigate to the target position point based on the target path and take a photo at the target position point. The target image taken at the target position point is input into the pre-trained vision language model so that the vision language model can perform image understanding based on the target image to obtain the image description output by the vision language model, and the image description is input into the large language model so that the large language model can output the corresponding detection result based on the picture description, that is, "Yes, this door is closed". It can be seen that the large language model converts human language instructions into detection primitives that the robot can execute, including navigation instructions and calling the vision language model for image understanding. Finally, the large language model will generate an anthropomorphic response for the user.
[0086] Figure 8 It is a schematic structural diagram of the terminal provided by the embodiment of the present application. The terminal may include:
[0087] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0088] When the processor 502 executes the program, it implements the robot semantic navigation method provided in the above embodiment.
[0089] Further, the terminal further includes:
[0090] A communication interface 503 for communication between the memory 501 and the processor 502.
[0091] A memory 501 for storing a computer program that can run on the processor 502.
[0092] The memory 501 may include a high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0093] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected via a bus to complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0094] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a chip, the memory 501, the processor 502, and the communication interface 503 can complete communication with each other through an internal interface.
[0095] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0096] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned robot semantic navigation method is implemented.
[0097] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in this application. The specification and examples are only illustrative, and the true scope and spirit of the present invention are pointed out by the claims.
[0098] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0099] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read instructions from the instruction execution system, apparatus, or device and execute the instructions), or in connection with these instruction execution systems, apparatus, or devices.
[0100] It should be understood that the various parts of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination of them can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA, Programmable Gate Array), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), etc.
[0101] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A robot semantic navigation method, characterized in that: The method comprises: Inputting the received natural language command and the pre-built semantic map into a pre-trained large language model so that the large language model can extract the target navigation object in the natural language command, and match the target navigation object with the semantic map to determine the target location point of the target navigation object on the semantic map; Performing path planning based on the current position of the robot and the target position point to obtain a corresponding target path so that the robot can navigate to the target position point based on the target path; The semantic map is a global map corresponding to the current spatial environment and containing semantic information.
2. The robot semantic navigation method according to claim 1, characterized in that: Before inputting the natural language command and the pre-built semantic map into the pre-trained large language model, the method further includes: Performing text processing on the pre-built semantic map to obtain a textual semantic map; The step of inputting the natural language command and the pre-built semantic map into a pre-trained large language model includes: The natural language command and the textual semantic map are input into a pre-trained large language model.
3. The robot semantic navigation method according to claim 1, characterized in that: The process of constructing a semantic map containing semantic information corresponding to the current spatial environment also includes: Based on the preset laser mapping algorithm, the preset laser radar device is used to scan the current space environment to construct the corresponding three-dimensional point cloud; At the same scanning time as the preset laser radar device, the current space environment is photographed using a preset photographing device to obtain a corresponding RGB image; Perform semantic segmentation on the RGB image using a pre-trained semantic segmentation model to obtain a corresponding segmentation mask; The three-dimensional point cloud is matched with the segmentation mask to determine a target point cloud, and a corresponding label is assigned to the target point cloud in the three-dimensional point cloud to obtain a semantic map containing semantic information.
4. The robot semantic navigation method according to claim 1, characterized in that: The matching the target navigation object with the semantic map to determine a target location point of the target navigation object on the semantic map includes: Matching the target navigation object with the semantic map to extract target semantic information of the target navigation object from the semantic map; wherein the target semantic information includes position information and size information of the target navigation object; The target position point of the target navigation object on the semantic map is determined based on the semantic information of the target navigation object and the current position of the robot.
5. The robot semantic navigation method according to claim 4, characterized in that: The step of determining a target location point of the target navigation object on the semantic map based on the semantic information of the target navigation object and the current position of the robot includes: Selecting corresponding position points on the edge of the target navigation object to obtain candidate position points of the target navigation object on the semantic map; Calculate the cost value gradient from the center of the target navigation object to each of the candidate position points based on the semantic information of the target navigation object, and obtain the target cost value gradient corresponding to each of the candidate position points; Calculate the Euclidean distance between the current position of the robot and each of the candidate position points to obtain the target distance corresponding to each of the candidate position points; Determine a score for each of the candidate location points based on the target cost value gradient and the target distance corresponding to each of the candidate location points; The scores of the candidate location points are compared to determine the candidate location point with the highest score, and the candidate location point with the highest score is determined as the target location point of the target navigation object on the semantic map.
6. The robot semantic navigation method according to claim 5, characterized in that: The determining the score of each candidate location point based on the target cost value gradient and the target distance corresponding to each candidate location point includes: Determine the score of each candidate location point by using a preset scoring formula and based on the target cost value gradient and the target distance corresponding to each candidate location point; Wherein, the preset scoring formula is: Among them, S(θ) represents the score, w g represents the gradient weight parameter, w d represents the distance weight parameter, G(θ) represents the target cost value gradient, and D(θ) represents the target distance.
7. The robot semantic navigation method according to any one of claims 1 to 6, characterized in that: After performing path planning based on the current position of the robot and the target position point to obtain a corresponding target path so that the robot can navigate to the target position point based on the target path, the method further includes: When the received natural language command is a command carrying a corresponding detection task, the target image captured at the target location is input into a pre-trained visual language model so that the visual language model performs image understanding based on the target image to obtain a picture description output by the visual language model, and the picture description is sent to the large language model so that the large language model outputs a corresponding detection result based on the picture description.
8. A robot semantic navigation system, characterized in that: The system comprises: A command processing module, used for inputting the received natural language command and the pre-built semantic map into a pre-trained large language model, so that the large language model can extract the target navigation object in the natural language command, and match the target navigation object with the semantic map to determine the target position point of the target navigation object on the semantic map; A path planning module, used to perform path planning based on the current position of the robot and the target position point to obtain a corresponding target path so that the robot can navigate to the target position point based on the target path; The semantic map is a global map corresponding to the current spatial environment and containing semantic information.
9. A terminal, characterized in that: include: A memory, a processor, and a robot semantic navigation program stored in the memory and executable on the processor, wherein the robot semantic navigation program, when executed by the processor, implements the steps of the robot semantic navigation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed to implement the steps of the robot semantic navigation method according to any one of claims 1 to 7.
Citation Information
Cited By
Robot moving operation method and device, electronic equipment, storage medium and computer program product
CN120697017A
Wheeled inspection robot and camera motion understanding method and system thereof
CN120766376A
Construction site quality inspection navigation method, quality inspection method and system
CN121026151A
A construction site quality inspection navigation method, quality inspection method and system
CN121026151B
Semantic-based unmanned aerial vehicle autonomous navigation method and device, equipment and medium
CN121383998A