A semantic navigation method, a semantic navigation device, and a robot
By constructing a semantic two-dimensional raster map and analyzing user statements, the robot can intelligently navigate to a specified semantic object, solving the problem of insufficient flexibility in robot navigation in the prior art and achieving higher mobility and flexibility.
Patent Information
- Application Number
- CN202210279083.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-03-21
AI Technical Summary
In the prior art, robot navigation requires users to manually specify target points on the activity map, resulting in low mobility and flexibility.
Build a semantic two-dimensional raster map of the environment, mark the semantic object categories, determine the target semantic object through user statements entered by speech or text, and navigate based on the semantic two-dimensional raster map.
It realizes intelligent interactive navigation of robots, improves mobility and flexibility, and allows users to specify the task execution location through language or text interaction at any time.
Smart Images

Figure CN114739408B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of navigation technology, and particularly relates to a semantic navigation method, a semantic navigation device, a robot, and a computer-readable storage medium. Background Art
[0002] With the intelligent development of robots, various types of robots can already replace humans to perform various tasks. Generally, a robot usually needs to first navigate from its current position to a certain target point before it can start performing the corresponding task.
[0003] Currently, when allowing a robot to perform a new task, the user needs to manually specify the coordinates of the target point for performing the new task in the robot's activity map. This specification method is relatively single and requires the user to operate on the map, resulting in low maneuverability and flexibility of the robot. Summary of the Invention
[0004] This application provides a semantic navigation method, a semantic navigation device, a robot, and a computer-readable storage medium, enabling the robot to achieve intelligent navigation starting from interactive behaviors, effectively improving the maneuverability and flexibility of the robot.
[0005] In a first aspect, this application provides a semantic navigation method, including:
[0006] Construct a semantic two-dimensional grid map of the environment, where the semantic two-dimensional grid map contains at least one semantic object, and each semantic object is marked with the semantic object category to which it belongs:
[0007] When the robot receives a user statement, if the intention of the user statement is navigation, determine the target semantic object indicated by the user statement in the semantic two-dimensional grid map, where the user statement is input by voice or text;
[0008] Perform navigation according to the target semantic object and the semantic two-dimensional grid map to control the robot to move towards the target semantic object.
[0009] In a second aspect, this application provides a semantic navigation device, including:
[0010] A construction module for constructing a semantic two-dimensional grid map of the environment, where the semantic two-dimensional grid map contains at least one semantic object, and each semantic object is marked with the semantic object category to which it belongs:
[0011] A determination module for, when the robot receives a user statement, if the intention of the user statement is navigation, determining the target semantic object indicated by the user statement in the semantic two-dimensional grid map, where the user statement is input by voice or text;
[0012] A navigation module, which is used to perform navigation according to a target semantic object and a semantic two-dimensional grid map to control the robot to move towards the target semantic object.
[0013] In a third aspect, the present application provides a robot. The robot includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method in the first aspect are implemented.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.
[0015] In a fifth aspect, the present application provides a computer program product. The computer program product includes a computer program. When the computer program is executed by one or more processors, the steps of the method in the first aspect are implemented.
[0016] The beneficial effects of the present application compared with the prior art are as follows: First, the robot constructs a semantic two-dimensional grid map of the environment, where the semantic two-dimensional grid map contains at least one semantic object, and each semantic object is marked with the category of the semantic object to which it belongs; then, when the robot receives a user statement, if the intention of the user statement is navigation, the target semantic object indicated by the user statement is determined in the semantic two-dimensional grid map, where the user statement is input by voice or text; finally, the robot can perform navigation according to the target semantic object and the semantic two-dimensional grid map to control the robot to move towards the target semantic object. The above process enables the robot to achieve intelligent interactive navigation, allowing the user to specify a new task execution location at any time through language or text interaction with the robot, effectively improving the mobility and flexibility of the robot.
[0017] It can be understood that the beneficial effects of the second to fifth aspects can refer to the relevant descriptions in the first aspect and will not be elaborated here. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of the implementation of the semantic navigation method provided by the embodiment of the present application;
[0020] Figure 2is an example diagram of the range to be traversed provided in an embodiment of the present application;
[0021] Figure 3 is another example diagram of the range to be traversed provided in an embodiment of the present application;
[0022] Figure 4 is another example diagram of the range to be traversed provided in the embodiment of the present application;
[0023] Figure 5 is a structural block diagram of a semantic navigation device provided in an embodiment of the present application;
[0024] Figure 6 It is a schematic diagram of the structure of the robot provided in the embodiment of the present application. DETAILED DESCRIPTION
[0025] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0026] In order to illustrate the technical solution proposed in this application, a specific embodiment is provided below for illustration.
[0027] The following is an explanation of the semantic navigation method proposed in the embodiment of the present application. Figure 1 ,The implementation process of the semantic navigation method is described in detail as follows:
[0028] Step 101, constructing a semantic two-dimensional grid map of the environment.
[0029] In the embodiment of the present application, the robot can build a semantic two-dimensional grid map of the environment in a semi-automatic or fully automatic manner by walking in the environment in which it is located (i.e., the environment in which the task is to be performed). It can be considered that the semantic two-dimensional grid map is the global map that the robot relies on in the subsequent navigation and movement process, and is the basis for navigation and movement.
[0030] In the semantic two-dimensional grid map, some or all obstacles are marked with the semantic object category to indicate what type of obstacle it is. For the convenience of distinction, these obstacles marked with the semantic object category are recorded as semantic objects. It can be understood that the difference between the semantic two-dimensional grid map and the two-dimensional grid map used in traditional navigation is:
[0031] In the two-dimensional grid map used in traditional navigation, only the positions occupied by obstacles are marked. Through this two-dimensional grid map used in traditional navigation, the robot can know which places in the environment have obstacles and which places are passable, but it cannot know what type of object a certain obstacle is.
[0032] The semantic two-dimensional grid map not only marks the positions occupied by obstacles, but also marks the semantic object categories to which the obstacles belong. Through this semantic two-dimensional grid map, the robot can not only know which places in the environment have obstacles, but also know what type of object a certain obstacle is.
[0033] In some embodiments, the process of constructing a semantic two-dimensional grid map in a semi-automatic manner can be briefly described as follows: The robot normally constructs a two-dimensional grid map without semantic objects using existing technologies and outputs it to the user for editing after completion. The user can compare the constructed two-dimensional grid map with the actual environment, manually confirm what type of object an obstacle in the two-dimensional grid map is, and mark the obstacles in the two-dimensional grid map based on the results of the manual confirmation, then a semantic two-dimensional grid map can be constructed.
[0034] In some embodiments, the process of constructing a semantic two-dimensional grid map in an automatic manner can be briefly described as follows: The robot obtains the depth information and color information of each object in the environment through the onboard environmental sensor. Here, the environmental sensor can be an RGB-D camera, or it can also be an RGB camera plus a radar (or other detection sensors), or it can also be a binocular RGB camera, which is not limited here. Among them, the color information of the object can be used as an analysis basis for the semantic object category to which the object belongs, and the depth information of the object can be used as an analysis basis for the position of the object in the environment. Based on this, the robot can construct a semantic two-dimensional grid map by analyzing the depth information and color information of each object in the environment.
[0035] Step 102, when the robot receives a user statement, if the intention of the user statement is navigation, determine the target semantic object indicated by the user statement in the semantic two-dimensional grid map.
[0036] In the embodiments of the present application, during the operation of the robot, the user can interact with the robot at any time according to their own needs. It can be understood that the interaction referred to here does not mean the interaction between the robot and the user in terms of physical actions, but rather the user conveys their intention to the robot through a user statement, and after the robot understands the intention, it performs corresponding operations. Among them, the user statement is input through voice or text.
[0037] When the robot receives a user statement, it can first analyze the user statement to determine the intention conveyed by the user through the user statement. Considering that this application focuses on the navigation and movement of the robot in the environment, therefore, only when the intention is navigation, does the robot have the need to perform navigation and movement. At this time, the target semantic object indicated by the user statement can be determined in the semantic two-dimensional grid map.
[0038] In some embodiments, the analysis process can be: set multiple keywords for the intention of navigation, such as "navigation", "go", "to", "move", and "travel" etc. If the user statement contains a keyword under the intention of navigation, it can be determined that the intention of the user statement is navigation.
[0039] In some embodiments, the analysis process can also be: a pre-trained statement intention classification model is built in the robot in advance. The robot inputs the user statement into the statement intention classification model to obtain the probability that the user statement hits the intention of navigation output by the statement intention classification model. If the probability is greater than a preset probability threshold, it can be determined that the intention of the user statement is navigation.
[0040] Merely as an example, the user can input the voice "Go to the table and wait" to the robot. Through the voice recognition technology of the robot, the user statement in voice form can be converted into a user statement in text form. By analyzing the user statement in text form in the way of the keywords proposed above or the statement intention classification model, it can be determined that the intention of the user statement is navigation, and subsequent operations can be executed.
[0041] Step 103, perform navigation according to the target semantic object and the semantic two-dimensional grid map to control the robot to move towards the target semantic object.
[0042] In the embodiments of this application, the essence of the target semantic object is an obstacle in the environment, and the robot actually cannot directly move to the position that coincides with the target semantic object (because that position is already occupied by the target semantic object). It can be understood that when the user drives the robot to move with the target semantic object as the reference for navigation, what is expected is that the robot can move as close as possible to the target semantic object. Based on this, the robot can take moving to the vicinity of the target semantic object as the goal, perform navigation according to the semantic two-dimensional grid map to control the robot to move towards the target semantic object, and realize the user's intention.
[0043] In some embodiments, when the robot performs navigation, it needs to first determine the end point of navigation, otherwise the robot cannot know where it should end the navigation. The robot can determine the target navigation point based on the goal of moving to the vicinity of the target semantic object. This target navigation point is the actual end point of this navigation. Step 103 can be specifically expressed as:
[0044] A1. Generate a cost map for the semantic 2D grid map, where the cost map includes an inflation layer.
[0045] As described above, the semantic 2D grid map is actually a more optimized 2D grid map. That is, based on the operations that can be performed on the 2D grid map, they can also be performed based on the semantic 2D grid map. Based on this, the robot can use the move_base navigation package based on the Robot Operating System (ROS) to generate the cost map of the semantic 2D grid map, and the specific generation process of this cost map will not be elaborated here.
[0046] It can be understood that the cost map contains several layers, such as the Static Map Layer, the Obstacle Map Layer, and the Inflation Layer, etc. The cost map will not be explained here.
[0047] A2. In the inflation layer, find the target navigation point according to the target semantic object.
[0048] Since the inflation layer is obtained by inflating (expanding outwards) on the basis of the static map layer and the obstacle map layer, when the robot searches for the target navigation point based on this inflation layer, it can ensure that there is a certain safety distance between the found target navigation point and the target semantic object.
[0049] A3. Navigate based on the target navigation point.
[0050] After finding the target navigation point, the robot can use this target navigation point as the end point of navigation and the current position point of the robot as the starting point of navigation to perform global path planning and local path tracking to achieve navigation movement. Specifically, the robot can use the method of Dijkstra global path planning and the method of TEB local path tracking, which is not limited here.
[0051] In some embodiments, to ensure that the target navigation point found by the robot is the nearest reachable point to the target semantic object, step A2 may specifically include:
[0052] A21. In the inflation layer, determine the range to be traversed.
[0053] Only as an example, the range to be traversed can be the entire inflation layer; or, the range to be traversed can also be a range that only includes some areas divided by the robot in the inflation layer according to certain limiting conditions, which is not limited here. However, it should be noted that the range to be traversed needs to include the centroid point of the target semantic object.
[0054] A22. Traverse outward starting from the centroid of the target semantic object within the range to be traversed.
[0055] The robot can adopt a breadth-first traversal method and traverse outward starting from the centroid of the target semantic object within the range to be traversed. It can be understood that since the breadth-first traversal method traverses the points closer to the centroid of the target semantic object first and then the points farther away from the centroid of the target semantic object, theoretically, the finally found target navigation point should be the point closest to the centroid among the points that meet the specified conditions.
[0056] A23. Determine whether the currently traversed point meets the specified conditions.
[0057] To ensure a certain safety distance between the target navigation point and the target semantic object, the robot can set the specified condition as: the cost value of the point is the preset value. Since when the cost value of a certain point is 0, it indicates that the environmental space corresponding to this point is the space where the robot can move freely, the preset value can be set to 0. That is, the robot can determine whether the cost value of the currently traversed point is 0. If it is not 0, it is determined that the currently traversed point does not meet the specified conditions. If it is 0, it is determined that the currently traversed point meets the specified conditions.
[0058] A24. If the currently traversed point does not meet the specified conditions, continue traversing and return to execute step A23 and subsequent steps until there are no un-traversed points within the range to be traversed.
[0059] In the case where the currently traversed point does not meet the specified conditions, traversing can continue, that is, select a new point as the currently traversed point and return to execute step A23 and subsequent steps. It can be understood that if the robot has not found a point that meets the specified conditions, it will keep traversing until all the points within the range to be traversed have been traversed before stopping.
[0060] A25. If the currently traversed point meets the specified conditions, stop traversing and determine the currently traversed point as the target navigation point.
[0061] In the case where the currently traversed point meets the specified conditions, the robot can immediately determine the currently traversed point as the target navigation point, stop traversing, and no longer continue to search.
[0062] In some extreme cases, all the points within the range to be traversed may not meet the specified conditions; that is, after traversing all the points within the range to be traversed, the target navigation point still cannot be determined. At this time, the robot can output a reminder message to remind the user that the current robot cannot move according to their requirements. Only as an example, the reminder message can be: "Unable to reach the destination you indicated. Please remove the obstacles near the destination or specify a new destination."
[0063] Among them, the output method of this reminder message can be unified with the input method of the user statement. For example, if the user inputs the user statement to the robot by voice, the robot can output the reminder message to the user by voice.
[0064] In some embodiments, to save the moving steps of the robot and avoid the robot going around behind the target semantic object, the robot can determine the range to be traversed in the following way:
[0065] In the dilation layer, the range sandwiched between the first straight line and the second straight line is determined as the range to be traversed, where the first straight line is the perpendicular line drawn from the centroid point of the target semantic object to the specified connection line, and the second straight line is the perpendicular line drawn from the position point where the robot is currently located to the specified connection line. The specified connection line is the connection line between the centroid point of the target semantic object and the position point where the robot is currently located.
[0066] For the robot, the orientation of the robot relative to the target semantic object is the front of the target semantic object. Please refer to Figure 2 , Figure 2 which gives an example of the range to be traversed determined in the dilation layer. In Figure 2 , the square is the target semantic object; point A is the centroid point of the target semantic object; point B is the position point where the robot is currently located; line l1 is the perpendicular line drawn from the centroid point of the target semantic object to the specified connection line AB; line l2 is the perpendicular line drawn from the position point where the robot is currently located to the specified connection line AB; the gray area is the range to be traversed. It should be noted that the range to be traversed includes the boundary, that is, the first straight line and the second straight line are also within the range to be traversed.
[0067] In some embodiments, to further save the moving steps of the robot, optimization can be performed on the basis of the range to be traversed proposed above. Then the robot can determine the range to be traversed in the following way:
[0068] In the dilation layer, the range sandwiched between the first straight line and the second straight line is determined as the first candidate range; the range to be traversed is determined within the first candidate range, where the distances from the points within the range to be traversed to the specified connection line are all less than or equal to the preset distance. The first straight line and the second straight line have been defined above and will not be elaborated here.
[0069] Please refer to Figure 3 , Figure 3 which gives another example of the range to be traversed determined in the dilation layer. In Figure 3 , the square is the target semantic object; point A is the centroid point of the target semantic object; point B is the position point where the robot is currently located; line l1 is the perpendicular line drawn from the centroid point of the target semantic object to the specified connection line AB; line l2 is the perpendicular line drawn from the position point where the robot is currently located to the specified connection line AB; the line segment CD is parallel to the specified connection line AB, and there is a preset distance d between them; the line segment EF is parallel to the specified connection line AB, and there is also a preset distance d between them; the gray area (i.e., the rectangle CDFE) is the range to be traversed. It should be noted that the range to be traversed includes the boundary, that is, the four sides of the rectangle CDFE are also within the range to be traversed.
[0070] In some embodiments, in order to prevent the target navigation point determined by the robot from being too far away from the target semantic object, the robot can also determine the range to be traversed in the following manner:
[0071] In the dilation layer, with the target semantic object as the center and the length of the specified connection line as the radius, a circular second candidate range is determined; the second candidate range is divided by the first line to obtain two semi - circles, and the semi - circle containing the specified connection line among the two semi - circles is determined as the range to be traversed. The first line has been defined above and will not be elaborated here.
[0072] Obviously, the points in the range to be traversed determined in this way are not farther from the centroid point of the target semantic object than the position point where the robot is currently located. Therefore, when navigating with the target navigation point determined within this range to be traversed, the situation where the robot walks farther and farther will not occur.
[0073] Please refer to Figure 4 , Figure 4 which gives yet another example of the range to be traversed determined in the dilation layer. In Figure 4 , the square is the target semantic object; point A is the centroid point of the target semantic object; point B is the position point where the robot is currently located; line l1 is the perpendicular line drawn from the centroid point of the target semantic object to the specified connection line AB; the gray area is the range to be traversed. It should be noted that the range to be traversed includes the boundary of this gray area.
[0074] In some embodiments, when the robot constructs a semantic two - dimensional grid map in an automatic manner, step 201 may include:
[0075] B1. Construct a three - dimensional semantic point cloud map of the environment through semantic segmentation of the RGBD image of the environment.
[0076] The robot can carry an RGB-D camera on its head or any part. It can be understood that when the RGB-D camera is installed on the head of the robot, the viewing angle of the RGB-D camera is close to that of a human. When the robot moves, it can collect the color image and depth image of the environment through the RGB-D camera at each moment. Usually, the sizes of the color image and depth image at the same moment are the same, and the corresponding pixel points (i.e., pixel points with the same coordinates) in the color image and depth image at the same moment actually point to the same location in the environment. For the color image and depth image collected at the same moment, the robot can perform the following operations:
[0077] Input the color image and depth image into the semantic segmentation model simultaneously to obtain the labeled image output by the semantic segmentation model. The size of the labeled image is the same as that of the color image and depth image, and each pixel point in the labeled image has been labeled with the semantic object category to which it belongs, and the semantic object category to which each pixel point belongs can be intuitively represented by the color of the pixel point. For example, if the point P1 in the labeled image belongs to the table, then the point P1 is red; if the point P2 in the labeled image belongs to the chair, then the point P2 is green.
[0078] Through the above operations, the depth image and labeled image at each moment can be obtained, such as the depth image and labeled image at time T1 (obtained by inputting the color image and depth image at time T1 into the semantic segmentation model), the depth image and labeled image at time T2 (obtained by inputting the color image and depth image at time T2 into the semantic segmentation model), and so on until the depth image and labeled image at time Tn (obtained by inputting the color image and depth image at time Tn into the semantic segmentation model). Based on the labeled images and depth images at each moment, the three-dimensional point cloud of the environment can be restored, and the semantic point cloud three-dimensional map of the environment can be reconstructed.
[0079] B2. Determine the semantic objects in the semantic point cloud three-dimensional map.
[0080] The robot can cluster each point in the semantic point cloud three-dimensional map. It should be noted that since there may be multiple objects in the environment belonging to the same semantic object category, when clustering, not only the semantic object category but also the Euclidean distance between points needs to be considered. It can be simply understood that points belonging to the same semantic object category and with an Euclidean distance less than the preset Euclidean distance threshold are clustered together, and these clustered points represent a semantic object.
[0081] Considering that there may be multiple semantic objects belonging to the same semantic object category in the environment, for the convenience of distinction, when the robot performs clustering, it can mark the serial numbers of each semantic object obtained by clustering. The process can be as follows: under each semantic object category, start from 1 and mark the serial numbers in an increasing manner. In this way, the uniqueness of the labels of each semantic object under the same semantic object category can be ensured. For example, there can be a table No. 1, a chair No. 1, a flower pot No. 1, etc., but there cannot be two tables No. 1.
[0082] It can be understood that the earlier the robot "sees" an object, the smaller its serial number (that is, the earlier it is marked with a serial number). If at the same moment, there are multiple semantic objects of the same semantic object category in the "field of view" of the robot, then the serial numbers of these semantic objects can be marked in the order from left to right as the first order and from bottom to top as the second order. It should be noted that the first order and the second order are only examples.
[0083] B3. Project the semantic objects in the semantic point cloud three-dimensional map onto a two-dimensional plane to obtain a semantic two-dimensional grid map.
[0084] The semantic objects in the semantic point cloud three-dimensional map are not only marked with the semantic object categories to which they belong, but also marked with serial numbers. These pieces of information are still recorded in the projected semantic two-dimensional grid map. In addition, for the convenience of users to consult, in the semantic two-dimensional grid map, the semantic object categories and serial numbers to which each semantic object belongs can also be displayed at each semantic object. When this semantic two-dimensional grid map is pushed by the robot to the user terminal (such as a smart phone) for the user to view; or when the semantic two-dimensional grid map is displayed on the screen of the robot for the user to view, the user can correspond the serial numbers of each semantic object with the actual objects in the environment, which is convenient for the user to more accurately control the robot to move to the desired location.
[0085] In some embodiments, the robot can first find the minimum bounding box of each semantic object in the semantic point cloud three-dimensional map, and then project the minimum bounding box of each semantic object onto a two-dimensional plane to obtain a semantic two-dimensional grid map. It can be understood that for any semantic object in the semantic two-dimensional grid map, its centroid point refers to: the point obtained by projecting the centroid point of the minimum bounding box of this semantic object in the semantic point cloud three-dimensional map onto the two-dimensional plane; that is, in the semantic two-dimensional grid map, the point obtained by projecting the centroid point of the minimum bounding box of this semantic object in the semantic point cloud three-dimensional map.
[0086] In some embodiments, when each semantic object in the semantic two-dimensional grid map is also marked with a serial number, the robot can determine the target semantic object indicated by the user's statement in the semantic two-dimensional grid map in the following manner:
[0087] C1. Analyze the user statement to determine the information about the semantic object category and the serial number carried by the user statement.
[0088] The robot can determine the information about the semantic object category carried by the user statement through the following process: For the robot, the semantic object category to which each semantic object in the semantic 2D grid map belongs is known. From this, the robot can expand and form a semantic object category keyword table by means of association. There are multiple semantic object category keywords in this semantic object category keyword table, and each semantic object category keyword in it corresponds to a semantic object category. The robot can match the user statement with each semantic object category keyword in the semantic object category keyword table and determine the information about the semantic object category carried by the user statement according to the matching result.
[0089] The robot can determine the information about the semantic object category carried by the user statement through the following process: After parsing the user statement, extract the number carried by the user statement, and this number is the information about the serial number.
[0090] Only as an example, when the semantic objects in the semantic 2D grid map are a table, a chair, and a flower pot, the semantic object category keyword table can include the following keywords: "table", "desk", "tea table", "chair", "stool", "flower pot", and "vase", etc. Among them, "table", "desk", and "tea table" correspond to the semantic object category of the table, "chair" and "stool" correspond to the semantic object category of the chair, and "flower pot" and "vase" correspond to the semantic object category of the flower pot. If the user statement is "Go and take a look at the 3rd table", then after parsing this user statement, the keyword "table" can be matched, and thus it can be known that the information about the semantic object category carried by the user statement is "table"; similarly, after parsing this user statement, the number "3" is extracted, and then the information about the serial number is known as "No. 3".
[0091] C2. Determine the target semantic object in the semantic 2D grid map according to the information about the semantic object category and the serial number carried by the user statement.
[0092] As can be seen from the previous description, a semantic object category and a serial number uniquely indicate a semantic object in the semantic 2D grid map. Thus, the robot can determine the target semantic object in the semantic 2D grid map according to the information about the semantic object category and the serial number carried by the user statement.
[0093] It can be understood that this method can help the robot quickly determine the target semantic object and improve the movement efficiency of the robot to a certain extent.
[0094] In some embodiments, if after parsing the user statement, the information about the semantic object category carried by the user statement can be determined, but the information about the serial number cannot be determined (for example, the user statement is "Go and take a look at the table"), then the following several possible operations can be considered:
[0095] For the first possible operation, the robot can, under this semantic object category, determine the semantic object with the serial number marked as 1 as the target semantic object and perform subsequent operations.
[0096] For the second possible operation, the robot can determine the semantic object closest to the position point where the robot itself is located (i.e., the semantic object closest to the robot itself) under this semantic object category as the target semantic object and perform subsequent operations.
[0097] For the third possible operation, the robot can output a reminder message to remind the user to specify the serial number. For example, the reminder message can be "Which number [xxx] do you mean? The serial numbers of each [xxx] can be viewed through the map!". Here, "[xxx]" can be filled according to the information about the semantic object category carried by the user statement. Then, based on the user's valid reply, the target semantic object is determined and subsequent operations are performed.
[0098] In some embodiments, if after parsing the user statement, the information about the serial number carried by the user statement can be determined, but the information about the semantic object category cannot be determined (for example, the user statement is "Go and take a look at number 3"), then the following several possible operations can be considered:
[0099] For the first possible operation, the robot can determine the semantic object category for this navigation based on the semantic object category with the most navigation times based on the user statement within a specified past time period; that is, the navigation times here do not consider the fixed navigation tasks completed by the robot at regular times / locations, but only consider the navigation tasks completed through the user statements input by the user. For example, within the past week, the user has controlled the robot to navigate to the table (regardless of which table number) 3 times through the user statement, controlled the robot to navigate to the chair (regardless of which chair number) 5 times through the user statement, and controlled the robot to navigate to the flower pot 0 times through the user statement. Then, based on the semantic object category with the most navigation times, "chair", and the determined information about the serial number, "number 3", the target semantic object is determined as the 3rd chair.
[0100] For the second possible operation, the robot can output a reminder message to remind the user to specify the semantic object category. For example, the reminder message can be "Where do you want [robot name] to go?". Then, based on the user's valid reply, the target semantic object is determined and subsequent operations are performed.
[0101] As can be seen from the above, through the embodiments of the present application, the robot can achieve intelligent interactive navigation, enabling users to specify new task execution locations at any time through language or text interaction with the robot, effectively improving the mobility and flexibility of the robot.
[0102] Corresponding to the semantic navigation method provided above, the embodiments of the present application also provide a semantic navigation device. As Figure 5 shown, the semantic navigation device 500 includes:
[0103] A construction module 501, configured to construct a semantic two-dimensional grid map of the environment, where at least one semantic object is included in the semantic two-dimensional grid map, and each of the above semantic objects is marked with the semantic object category to which it belongs:
[0104] A determination module 502, configured to, when the robot receives a user statement, if the intention of the user statement is navigation, determine the target semantic object indicated by the user statement in the semantic two-dimensional grid map, where the user statement is input by voice or text;
[0105] A navigation module 503, configured to perform navigation according to the target semantic object and the semantic two-dimensional grid map to control the robot to move towards the target semantic object.
[0106] Optionally, the above navigation module 503 includes:
[0107] A generation unit, configured to generate a cost map of the semantic two-dimensional grid map, where the cost map includes an inflation layer;
[0108] A search unit, configured to search for a target navigation point in the inflation layer according to the target semantic object;
[0109] A navigation unit, configured to perform navigation based on the target navigation point.
[0110] Optionally, the above search unit includes:
[0111] A range determination subunit, configured to determine a range to be traversed in the inflation layer;
[0112] A traversal subunit, configured to traverse outward from the centroid point of the target semantic object within the range to be traversed;
[0113] A judgment subunit, configured to judge whether the currently traversed point meets the specified conditions;
[0114] A first processing subunit, configured to, if the currently traversed point does not meet the specified conditions, continue traversing and trigger the execution of the judgment subunit again until there are no untraversed points in the range to be traversed;
[0115] A second processing sub-unit, configured to stop traversing if the currently traversed point meets the specified condition, and determine the currently traversed point as the target navigation point.
[0116] Optionally, the determination sub-unit is specifically configured to determine whether the cost value of the currently traversed point is a preset value. If the cost value of the currently traversed point is not the preset value, it is determined that the currently traversed point does not meet the specified condition. If the cost value of the currently traversed point is the preset value, it is determined that the currently traversed point meets the specified condition.
[0117] Optionally, the range determination sub-unit is specifically configured to, in the dilation layer, determine the range sandwiched by the first straight line and the second straight line as the range to be traversed, where the first straight line is: a perpendicular line drawn from the centroid point of the target semantic object to the specified connection line, and the second straight line is: a perpendicular line drawn from the position point where the robot is currently located to the specified connection line, and the specified connection line is: a connection line between the centroid point of the target semantic object and the position point where the robot is currently located.
[0118] Optionally, the building module 501 includes:
[0119] Construct a three-dimensional semantic point cloud map of the environment through semantic segmentation of the RGBD image of the environment;
[0120] Determine the semantic objects in the three-dimensional semantic point cloud map;
[0121] Project the semantic objects in the three-dimensional semantic point cloud map onto a two-dimensional plane to obtain the semantic two-dimensional grid map.
[0122] Optionally, each of the semantic objects in the semantic two-dimensional grid map is also marked with a serial number; the determination module 502 includes:
[0123] A statement parsing unit, configured to parse the user statement to determine the information about the category of the semantic object and the information about the serial number carried by the user statement;
[0124] An object determination unit, configured to determine the target semantic object in the semantic two-dimensional grid map according to the information about the category of the semantic object and the information about the serial number carried by the user statement.
[0125] As can be seen from the above, through the embodiments of the present application, the robot can achieve intelligent interactive navigation, enabling the user to specify a new task execution location at any time through language or text interaction with the robot, effectively improving the mobility and flexibility of the robot.
[0126] Corresponding to the semantic navigation method provided above, an embodiment of the present application further provides a robot. Please refer to Figure 6 , the robot 6 in the embodiment of the present application includes: a memory 601, one or more processors 602 ( Figure 6 only one is shown in the figure) and a computer program stored on the memory 601 and executable on the processor. Among them: the memory 601 is used to store software programs and units, and the processor 602 executes various functional applications and data processing by running the software programs and units stored in the memory 601 to obtain resources corresponding to the above preset events. Specifically, when the processor 602 runs the above computer program stored in the memory 601, the following steps are implemented:
[0127] Construct a semantic two-dimensional grid map of the environment, where at least one semantic object is included in the above semantic two-dimensional grid map, and each of the above semantic objects is marked with the semantic object category to which it belongs:
[0128] When the robot receives a user statement, if the intention of the above user statement is navigation, determine the target semantic object indicated by the above user statement in the above semantic two-dimensional grid map, where the above user statement is input by voice or text;
[0129] Navigate according to the above target semantic object and the above semantic two-dimensional grid map to control the above robot to move towards the above target semantic object.
[0130] Assuming the above is the first possible implementation manner, then in the second possible implementation manner provided on the basis of the first possible implementation manner, the above navigation according to the above target semantic object and the above semantic two-dimensional grid map includes:
[0131] Generate a cost map of the above semantic two-dimensional grid map, where the above cost map includes a dilation layer;
[0132] In the above dilation layer, find a target navigation point according to the above target semantic object;
[0133] Navigate based on the above target navigation point.
[0134] In the third possible implementation manner provided on the basis of the above first possible implementation manner, the above finding a target navigation point according to the above target semantic object in the above dilation layer includes:
[0135] In the above dilation layer, determine the range to be traversed;
[0136] In the above range to be traversed, traverse outward starting from the centroid point of the above target semantic object;
[0137] Determine whether the currently traversed point meets the specified conditions;
[0138] If the currently traversed point does not meet the specified conditions, continue traversing, and return to execute the step of determining whether the currently traversed point meets the specified conditions and subsequent steps until there are no untraversed points within the range to be traversed;
[0139] If the currently traversed point meets the specified conditions, stop traversing and determine the currently traversed point as the target navigation point.
[0140] In the fourth possible implementation provided based on the above third possible implementation, the determination of whether the currently traversed point meets the specified conditions includes:
[0141] Determine whether the cost value of the currently traversed point is a preset value;
[0142] If the cost value of the currently traversed point is not the preset value, determine that the currently traversed point does not meet the specified conditions;
[0143] If the cost value of the currently traversed point is the preset value, determine that the currently traversed point meets the specified conditions.
[0144] In the fifth possible implementation provided based on the above third possible implementation, in the above inflation layer, determining the range to be traversed includes:
[0145] In the above inflation layer, determine the range sandwiched between the first straight line and the second straight line as the range to be traversed, where the first straight line is: the perpendicular line drawn from the centroid point of the above target semantic object to the specified connection line, and the second straight line is: the perpendicular line drawn from the position point where the above robot is currently located to the above specified connection line, and the above specified connection line is: the connection line between the centroid point of the above target semantic object and the position point where the above robot is currently located.
[0146] In the sixth possible implementation provided based on the above first possible implementation, the construction of the semantic 2D grid map of the environment includes:
[0147] Construct the semantic point cloud 3D map of the above environment through semantic segmentation of the RGBD image of the environment;
[0148] Determine the semantic objects in the above semantic point cloud 3D map;
[0149] Project the semantic objects in the above semantic point cloud 3D map onto a 2D plane to obtain the above semantic 2D grid map.
[0150] In a seventh possible implementation provided based on the above first possible implementation, each of the semantic objects in the semantic two-dimensional grid map is also marked with a serial number; determining the target semantic object indicated by the user statement in the semantic two-dimensional grid map includes:
[0151] Parse the user statement to determine the information on the category of the semantic object and the information on the serial number carried by the user statement;
[0152] Determine the target semantic object in the semantic two-dimensional grid map according to the information on the category of the semantic object and the information on the serial number carried by the user statement.
[0153] It should be understood that in the embodiments of the present application, the so-called processor 602 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0154] The memory 601 may include a read-only memory and a random access memory, and provide instructions and data to the processor 602. A part or all of the memory 601 may also include a non-volatile random access memory. For example, the memory 601 may also store information on the device category.
[0155] As can be seen from the above, through the embodiments of the present application, the robot can achieve intelligent interactive navigation, enabling the user to specify a new task execution location at any time through language or text interaction with the robot, effectively improving the mobility and flexibility of the robot.
[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0157] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0158] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of external device software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0159] In the embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0160] The units described as separate components above may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of this application, it can also be completed by a computer program instructing related hardware. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The above computer-readable storage medium can include: any entity or device that can carry the above computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer-readable memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the above computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0162] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A semantic navigation method, characterized in that, Including: Construct a semantic two-dimensional grid map of the environment, where the semantic two-dimensional grid map contains at least one semantic object, and each semantic object is labeled with the semantic object category it belongs to: When the robot receives a user statement, if the intention of the user statement is navigation, determine the target semantic object indicated by the user statement in the semantic two-dimensional grid map, where the user statement is input by voice or text; Generate a cost map of the semantic two-dimensional grid map, and the cost map includes a dilation layer; In the dilation layer, determine the range sandwiched between the first straight line and the second straight line as the first candidate range, and determine the range to be traversed within the first candidate range; the distance from each point within the range to be traversed to the specified connection line is less than or equal to a preset distance. The first straight line is a perpendicular line drawn from the centroid point of the target semantic object to the specified connection line, and the second straight line is a perpendicular line drawn from the position point where the robot is currently located to the specified connection line. The specified connection line is the connection line between the centroid point of the target semantic object and the position point where the robot is currently located; Within the range to be traversed, start traversing outward from the centroid point of the target semantic object; Judge whether the currently traversed point meets the specified conditions, including: judging whether the cost value of the currently traversed point is a preset value; if the cost value of the currently traversed point is not the preset value, determine that the currently traversed point does not meet the specified conditions; if the cost value of the currently traversed point is the preset value, determine that the currently traversed point meets the specified conditions, and the preset value is used to judge whether the current point is passable; If the currently traversed point does not meet the specified conditions, continue traversing, and return to execute the steps of judging whether the currently traversed point meets the specified conditions and subsequent steps until there are no untraversed points within the range to be traversed; if the currently traversed point meets the specified conditions, stop traversing, and determine the currently traversed point as the target navigation point; Navigate based on the target navigation point to control the robot to move towards the target semantic object.
2. The semantic navigation method according to claim 1, characterized in that, The construction of the semantic two-dimensional grid map of the environment includes: Construct a semantic point cloud three-dimensional map of the environment through semantic segmentation of the RGBD image of the environment; Determine the semantic objects in the semantic point cloud three-dimensional map; Project the semantic objects in the semantic point cloud three-dimensional map onto a two-dimensional plane to obtain the semantic two-dimensional grid map.
3. The semantic navigation method according to claim 1, wherein Each semantic object in the semantic two-dimensional grid map is also labeled with a serial number; the determination of the target semantic object indicated by the user statement in the semantic two-dimensional grid map includes: Parse the user statement to determine the information of the semantic object category and the information of the serial number carried by the user statement; Determine the target semantic object in the semantic two-dimensional grid map according to the information of the semantic object category and the information of the serial number carried by the user statement.
4. A semantic navigation device, characterized in that, The semantic navigation device includes: A building module for building a semantic two-dimensional grid map of the environment, where the semantic two-dimensional grid map contains at least one semantic object, and each semantic object is labeled with the semantic object category to which it belongs: A determination module for, when the robot receives a user statement, if the intention of the user statement is navigation, determining the target semantic object indicated by the user statement in the semantic two-dimensional grid map, where the user statement is input by voice or text; A generation unit for generating a cost map of the semantic two-dimensional grid map, the cost map including a dilation layer; A range determination subunit for, in the dilation layer, determining the range sandwiched between a first straight line and a second straight line as a first candidate range, and determining a range to be traversed within the first candidate range; the distances from all points within the range to be traversed to a specified connection line are less than or equal to a preset distance, the first straight line is a perpendicular line drawn from the centroid point of the target semantic object to the specified connection line, the second straight line is a perpendicular line drawn from the position point where the robot is currently located to the specified connection line, and the specified connection line is the connection line between the centroid point of the target semantic object and the position point where the robot is currently located; A traversal subunit for traversing outward from the centroid point of the target semantic object within the range to be traversed; A judgment subunit for judging whether the currently traversed point meets a specified condition, including: judging whether the cost value of the currently traversed point is a preset value; if the cost value of the currently traversed point is not the preset value, determining that the currently traversed point does not meet the specified condition; if the cost value of the currently traversed point is the preset value, determining that the currently traversed point meets the specified condition, and the preset value is used to judge whether the current point is passable; A first processing subunit for, if the currently traversed point does not meet the specified condition, continuing to traverse and returning to execute the step of judging whether the currently traversed point meets the specified condition and subsequent steps until there are no untraversed points within the range to be traversed; A second processing subunit for, if the currently traversed point meets the specified condition, stopping the traversal and determining the currently traversed point as the target navigation point; A navigation unit for navigating based on the target navigation point to control the robot to move towards the target semantic object.
5. A robot, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Mobile robot indoor navigation method based on semantic information
CN107063258A
Method for constructing indoor two-dimensional semantic grid map with object navigation point
CN111486855A
Navigation target position determination method and device, readable storage medium and robot
CN111854751A