A navigation method and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-28
- Publication Date
- 2026-08-14
AI Technical Summary
[0114]从以上技术方案可以看出,本申请实施例具有以下优点:云端设备接收终端设备的导航请求后,使用第一多层拓扑结构确定坐标,并向终端设备发送坐标,使得终端设备可以根据第一多层拓扑结构确定的坐标向导航物体移动。通过引入具有多层节点语义关联的多层拓扑结构,即通过语义特征导航终端设备向导航物体进行移动,由于语义特征不易受场景细节变化的影响,从而提升终端设备导航的准确度以及在多变场景下导航的泛化性。
Smart Images

Figure CN115600053B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal artificial intelligence, and in particular to a navigation method and related equipment. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and fundamental AI theories.
[0003] Currently, robotics technology and related applications have gradually penetrated into people's daily work, production, and life. Among these, mobile robot navigation technology has spurred the rapid development of autonomous driving, leading to the initial deployment of unmanned autonomous mobile chassis in public service scenarios such as parks, shopping malls, restaurants, and hospitals. Simultaneously, popular home application products like robotic vacuum cleaners have emerged and entered countless households. The successful promotion of robotic vacuum cleaners signifies the formal integration of mobile robots into the home service scenario.
[0004] Therefore, how to achieve precise navigation that mobile robots cannot perform is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a navigation method and related equipment. By introducing a multi-layered topology with multi-level node semantic relationships and performing navigation based on this multi-layered topology, the accuracy of robot navigation is improved.
[0006] The first aspect of this application provides a navigation method applicable to navigation scenarios such as homes, shopping malls, and airports. This method can be executed by a cloud device or by components of the cloud device (e.g., processors, chips, or chip systems). The method includes: receiving a navigation request sent by a terminal device, the navigation request indicating a navigation object for the terminal device; determining the coordinates of intermediate nodes based on a first multi-layer topology and the semantic features of the navigation object; the first multi-layer topology includes a first-layer structure and a second-layer structure, the first-layer structure including multiple first nodes and multiple first descriptors corresponding to the multiple first nodes, the second-layer structure including multiple second nodes and multiple second descriptors corresponding to the multiple second nodes, each of the multiple first nodes indicating a first object, the first descriptors describing the semantic features of the first object indicated by the corresponding first node, each of the multiple second nodes being associated with at least one group of first nodes, the first node group including one or more first nodes, the one or more first nodes being related to the position information of their respective corresponding first objects, each of the multiple second descriptors describing the semantic features of the first object indicated by the associated first node group; the coordinates of the intermediate nodes being used to guide the terminal device to move towards the navigation object; and sending the coordinates to the terminal device. Furthermore, the aforementioned positional relationship can be information such as the coordinates of the first object, or it can be the relative positional relationship between the first object and other first objects. The navigation request can be a category label or image of the navigation object sent by the robot; specific details are not limited here. Of course, an intermediate node can be any point on the path between the terminal device's location and the navigation object. An intermediate node can also indicate the navigation object; that is, determining the coordinates of the intermediate node determines the coordinates of the navigation object.
[0007] In this embodiment, after receiving a navigation request from a terminal device, the cloud device uses a first multi-layer topology structure to determine coordinates and sends the coordinates to the terminal device, enabling the terminal device to move towards the navigation object based on the coordinates determined by the first multi-layer topology structure. By introducing a multi-layer topology structure with semantic associations among multiple nodes, that is, by navigating the terminal device towards the navigation object through semantic features, the accuracy of the terminal device's navigation and its generalization in changing scenarios are improved because semantic features are not easily affected by changes in scene details.
[0008] Optionally, in one possible implementation of the first aspect, the aforementioned first descriptor is further used to describe the association between the corresponding first node and at least one other first node among a plurality of first nodes.
[0009] In this possible implementation, the first descriptor can describe the relationships between multiple nodes. These relationships are learned and can refer to explicit relationships (e.g., adjacency relationships) or implicit relationships (e.g., user habits, such as the habit of using a water cup and a water dispenser simultaneously). This is beneficial for the accuracy of subsequent semantic navigation.
[0010] Optionally, in one possible implementation of the first aspect, the aforementioned association relationships are used to describe at least one of the following: category relationship, functional relationship, matching relationship, bearing relationship, positional relationship, membership relationship, etc., between at least two first objects. Specifically, a category relationship means that at least two first objects belong to the same or similar categories; a functional relationship means that at least two first objects have the same or similar functions; a matching relationship means that at least two first objects often perform a certain function together in practical applications; a bearing relationship means that at least two first objects have a bearing and being-bearing relationship; and a positional relationship means that at least two first objects have a spatial relationship, specifically, the adjacent relationship between objects within a certain range (e.g., 1 meter). For example: if one first object is a red table and the other is a black table, then the two first objects have a category relationship (i.e., both are tables). If both first objects belong to kitchen utensils, then the two first objects have a functional or category relationship. If the two first objects are a table and a chair, or a car and a parking lot, respectively, then the two first objects have a matching relationship. If the two first objects are vegetables and a refrigerator, respectively, then the two first objects have a bearing relationship. If the two primary objects are stationery and a table, then the two primary objects have a positional relationship (generally, stationery is placed on the table). It's understandable that this relationship can also represent user habits, preferences, etc.
[0011] In this possible implementation, each first descriptor is used to represent the semantic features of the association relationship between the corresponding first node and at least one other first node among multiple first nodes. Navigation through the first multi-layer topology structure constructed by object association relationships is more generalizable.
[0012] Optionally, in one possible implementation of the first aspect, the above steps further include: predicting a third node based on location information; predicting a third descriptor of the third node based on multiple first descriptors or multiple second descriptors; determining the coordinates of the intermediate node based on the semantic features of the first multi-layer topology and the navigation object, including: updating the first multi-layer topology based on the third node and the third descriptor to obtain a second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology; and determining the coordinates based on the semantic features of the second multi-layer topology and the navigation object. Wherein, if the third node is a first-layer node, the third descriptor of the third node is predicted based on multiple first descriptors. If the third node is a second-layer node, the third descriptor of the third node is predicted based on multiple second descriptors. The third node can be understood as a potential connected region.
[0013] In this possible implementation, the third node is predicted through location information, the third node is used to update the first multi-layer topology to obtain the second multi-layer topology, and the second multi-layer topology is used for navigation. This allows for feature prediction of unknown areas, improving the robot's exploration of unknown areas.
[0014] Optionally, in one possible implementation of the first aspect, the above step of predicting the third node based on location information includes: predicting the third node based on the Vinograph corresponding to the location information. Specifically, the Vinograph includes vertices, edges, and a convex hull. Using the extended edges of the Vinograph as potential connected components, if a node at one end has not yet generated semantic features, a sampling point is selected on it as the third node, representing an unknown node.
[0015] In this possible implementation, the third node is predicted by the edges of the Veno graph, making the predicted third node more reasonable.
[0016] Optionally, in one possible implementation of the first aspect, the above steps: determining the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object include: determining multiple candidate nodes based on the Upper Confidence Interval (UCT) algorithm, wherein the multiple candidate nodes are nodes in the first multi-layer topology; calculating the probability of each candidate node among the multiple candidate nodes as an intermediate node based on the distance, so as to obtain the coordinates of the intermediate node, wherein the distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes, and the intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to a first threshold.
[0017] In this possible implementation, UCT technology is used to construct a Monte Carlo tree search from the established first multi-layer topology. By utilizing the similarity between the semantic features of each layer node and the semantic features of the navigation object, the robot can autonomously balance the use of known information and the exploration of unknown information to carry out goal-driven reasoning-based exploration and navigation.
[0018] Optionally, in one possible implementation of the first aspect, the above step of: calculating the probability of each candidate node among multiple candidate nodes as an intermediate node based on the distance includes: calculating the probability of each candidate node among multiple candidate nodes as an intermediate node based on similarity, access count and distance, so as to obtain the coordinates of the intermediate node, where similarity is the similarity between the semantic features corresponding to multiple candidate nodes and the semantic features of the navigation object, access count is the number of times the terminal device accesses each candidate node, distance is the distance that the terminal device needs to travel to reach multiple candidate nodes, and the intermediate node is the node among multiple candidate nodes whose probability is greater than or equal to a first threshold.
[0019] In this possible implementation, the value of a node is determined by the path. If the mobile robot explores the same node too many times and has not found the navigation object, the mobile robot will abandon the current area and go to the unknown area to explore, thus shortening the time to move to the navigation object.
[0020] Optionally, in one possible implementation of the first aspect, the above step of calculating the probability of each candidate node among multiple candidate nodes as an intermediate node based on similarity, number of visits, and distance includes: calculating the value of each candidate node among multiple candidate nodes using the following formula:
[0021]
[0022] Where i represents one of a plurality of candidate nodes, V(i) represents the value of the candidate node, ω represents the similarity, and L dis Let m represent the distance, j represent the total number of child nodes of the current branch, N represent the total number of visits to the candidate node and its branch child nodes, n represent the number of times the terminal device visits the candidate node, and c1 and c2 are adjustment coefficients.
[0023] In this possible implementation, a backtracking problem is introduced, where the value of a node is determined by the number of times it is visited and the distance between nodes. If the mobile robot explores the same node too many times and has not yet found the navigation object, the mobile robot abandons the current area and goes to an unknown area to explore, thus shortening the time to reach the navigation object.
[0024] Optionally, in one possible implementation of the first aspect, the navigation request further includes location information and / or scale information of multiple first objects. The scale information includes at least one of the following: the number of first objects, the number of rooms containing the first objects, and the area of the region containing the first objects. The scale information is used to determine the number of layers in the first multi-layer topology. This scale information can be related to the number of objects, or it can be related to a range or area. For example, the more objects there are, the more layers the first topology can have. In a navigation scenario within a house, it can also be related to the number of rooms.
[0025] In this possible implementation, the number of layers in the first multi-layer topology can be determined by the scale information of the environment in which the first objects are located. The number of layers can be set according to actual needs, and iterative spatial segmentation and clustering techniques can be used to perform semantic-spatial association to establish a multi-level topology, which is beneficial for accurate robot navigation.
[0026] Optionally, in one possible implementation of the first aspect, the aforementioned first multi-layer topology further includes a third-layer structure, which includes a plurality of fourth nodes and a plurality of fourth descriptors corresponding to the plurality of fourth nodes; each of the plurality of fourth nodes is associated with at least one second node group, and the second node group includes one or more second nodes; each of the plurality of fourth descriptors is used to describe the semantic features of the second node indicated by the second node group. The plurality of second nodes have an association relationship.
[0027] In this possible implementation, the first multi-layer topology may include a three-layer structure. It is understood that the number of layers in the first multi-layer topology is set according to actual needs, and the descriptors corresponding to the nodes of the previous layer are determined based on the node group of this layer related to the nodes of the previous layer, which helps to improve the accuracy of the navigation robot.
[0028] Optionally, in one possible implementation of the first aspect, the above steps further include: acquiring position information of multiple first objects; and constructing a first multi-layer topology based on the position information.
[0029] In this possible implementation, a multi-layered topology is built in the cloud to reduce the computing power consumption of terminal devices (i.e., mobile robots).
[0030] Optionally, in one possible implementation of the first aspect, the above steps: constructing a first multi-layer topology based on location information, include: constructing a multi-layer structure based on location information, the multi-layer structure including first-layer nodes and second-layer nodes; aggregating multiple first descriptors to obtain a second descriptor; associating multiple first descriptors, second descriptors, and the multi-layer structure to obtain the first multi-layer topology.
[0031] In this possible implementation, a multi-layer structure can be constructed first using location information, and then the first multi-layer topology can be obtained by associating each layer node with its corresponding descriptor.
[0032] Optionally, in one possible implementation of the first aspect, the above steps further include: receiving a first multi-layer topology structure sent by the terminal device.
[0033] In this possible implementation, the terminal constructs a multi-layer topology, collecting data from the surrounding environment while updating the multi-layer topology, which improves the efficiency of multi-layer topology construction and updating.
[0034] Optionally, in one possible implementation of the first aspect, the first descriptor is obtained based on the initial descriptor of the first object, the initial descriptor is determined based on the position information of the first object, and each of the multiple initial descriptors is used to represent the semantic features of the corresponding first object; the first descriptor represents the semantic features of the association between the multiple first objects.
[0035] In this possible implementation, semantic features that can be used to represent the relationships between multiple first objects are obtained through multiple initial descriptors. This improves the accuracy and generalization of navigation based on the semantic features of subsequent navigable objects.
[0036] Optionally, in one possible implementation of the first aspect, the first descriptor is obtained based on multiple initial descriptors through a first network, which is used to obtain semantic features representing the relationship between multiple first objects.
[0037] In this possible implementation, a first descriptor that can represent the semantic features of the relationship between multiple first objects is generated based on the individual initial descriptors of multiple first objects, and the first descriptor is used to represent the first object, which is beneficial to the accuracy of subsequent navigation based on the semantic features of the navigation object.
[0038] Optionally, in one possible implementation of the first aspect, the second descriptor is obtained through a second network based on multiple first descriptors, and the second network is used to obtain shared descriptors representing multiple first nodes associated with the second node.
[0039] In this possible implementation, a shared descriptor that can represent the node group associated with the second node is generated based on multiple first descriptors, which is beneficial to the accuracy of subsequent navigation based on the semantic features of the navigation object.
[0040] Optionally, in one possible implementation of the first aspect, the aforementioned third descriptor is obtained through a third network based on multiple first descriptors or multiple second descriptors, the third network being used to predict the descriptor of the third node based on the descriptors of nodes at the same level.
[0041] In this possible implementation, the third descriptor of the third node is predicted by using the descriptors of nodes at the same level. The descriptors of nodes at the same level can represent the semantic features of multiple objects having a relationship. That is, the third descriptor predicted by the relationship between objects can more accurately identify the semantic features of the third node.
[0042] The second aspect of this application provides a navigation method applicable to navigation scenarios such as homes, shopping malls, and airports. This method can be executed by a terminal device or by components of the terminal device (such as a processor, chip, or chip system). The terminal device can be a mobile robot (such as a robotic vacuum cleaner, a transport robot, a guided robot, etc.). The method includes: receiving a user's movement command, the movement command indicating movement towards a navigation object; determining the coordinates of intermediate nodes based on a first multi-layer topology and the semantic features of the navigation object; the first multi-layer topology includes a first layer structure and a second layer structure, the first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes, the second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes, each of the multiple first nodes indicates a first object, the first descriptor is used to describe the semantic features of the first object indicated by the corresponding first node, each of the multiple second nodes is associated with at least one group of first nodes, the group of first nodes includes one or more first nodes, the one or more first nodes are related to the position information of their respective corresponding first objects, each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated group of first nodes; moving towards the navigation object based on the coordinates, of course, the intermediate node can be any point on the path between the location of the terminal device and the navigation object, the intermediate node can also indicate the navigation object, that is, determining the coordinates of the intermediate node determines the coordinates of the navigation object.
[0043] In this embodiment, after receiving a movement command, the terminal device can determine the coordinates based on the first multi-layer topology and move towards the navigation object based on the coordinates. By introducing a multi-layer topology with multi-layer node semantic association, that is, moving towards the navigation object through semantic features, the accuracy of robot navigation and the generalization of navigation in changing scenarios are improved because semantic features are not easily affected by changes in scene details.
[0044] Optionally, in one possible implementation of the second aspect, the first descriptor described above is also used to describe the association between the corresponding first node and at least one other first node among a plurality of first nodes.
[0045] In this possible implementation, the first descriptor can describe the relationships between multiple nodes. These relationships are learned and can refer to explicit relationships (e.g., adjacency relationships) or implicit relationships (e.g., user habits, such as the habit of using a water cup and a water dispenser simultaneously). This is beneficial for the accuracy of subsequent semantic navigation.
[0046] Optionally, in one possible implementation of the second aspect, the aforementioned association relationships are used to describe at least one of the following: category relationship, functional relationship, matching relationship, and positional relationship between at least two first objects. Specifically, a category relationship means that at least two first objects belong to the same or similar categories; a functional relationship means that at least two first objects have the same or similar functions; a matching relationship means that at least two first objects often perform a certain function together in practical applications; a carrying relationship means that at least two first objects have a carrying and being carried relationship; and a positional relationship means that at least two first objects have a spatial relationship, specifically referring to the adjacency relationship between objects within a certain range (e.g., 1 meter). For example: if one first object is a red table and the other is a black table, then the two first objects have a category relationship (i.e., both are tables). If both first objects belong to kitchen utensils, then the two first objects have a functional or category relationship. If the two first objects are a table and a chair, or a car and a parking lot respectively, then the two first objects have a matching relationship. If the two first objects are vegetables and a refrigerator respectively, then the two first objects have a carrying relationship. If the two primary objects are stationery and a table, then the two primary objects have a positional relationship (generally, stationery is placed on the table). It's understandable that this relationship can also represent user habits, preferences, etc.
[0047] In this possible implementation, each first descriptor is used to represent the semantic features of the association relationship between the corresponding first node and at least one other first node among multiple first nodes. Navigation through the first multi-layer topology structure constructed by object association relationships is more generalizable.
[0048] Optionally, in one possible implementation of the second aspect, the above steps further include: predicting a third node based on location information; predicting a third descriptor of the third node based on multiple first descriptors or multiple second descriptors; determining the coordinates of the intermediate node based on the semantic features of the first multi-layer topology and the navigation object, including: updating the first multi-layer topology based on the third node and the third descriptor to obtain a second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology; and determining the coordinates based on the semantic features of the second multi-layer topology and the navigation object. Wherein, if the third node is a first-layer node, the third descriptor of the third node is predicted based on multiple first descriptors. If the third node is a second-layer node, the third descriptor of the third node is predicted based on multiple second descriptors. The third node can be understood as a potential connected region.
[0049] In this possible implementation, the third node is predicted through location information, the third node is used to update the first multi-layer topology to obtain the second multi-layer topology, and the second multi-layer topology is used for navigation. This allows for feature prediction of unknown areas, improving the robot's exploration of unknown areas.
[0050] Optionally, in one possible implementation of the second aspect, the above step of predicting the third node based on location information includes: predicting the third node based on the Vinograph corresponding to the location information. Specifically, the Vinograph includes vertices, edges, and a convex hull. Using the extended edges of the Vinograph as potential connected components, if a node at one end has not yet generated semantic features, a sampling point is selected on it as the third node, representing an unknown node.
[0051] In this possible implementation, the third node is predicted by the edges of the Veno graph, making the predicted third node more reasonable.
[0052] Optionally, in one possible implementation of the second aspect, the above steps: determining the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object include: determining multiple candidate nodes based on the Upper Confidence Interval (UCT) algorithm, wherein the multiple candidate nodes are nodes in the first multi-layer topology; calculating the probability of each candidate node among the multiple candidate nodes as an intermediate node based on the distance, so as to obtain the coordinates of the intermediate node, wherein the distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes, and the intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to a first threshold.
[0053] In this possible implementation, UCT technology is used to construct a Monte Carlo tree search from the established first multi-layer topology. By utilizing the similarity between the semantic features of each layer node and the semantic features of the navigation object, the robot can autonomously balance the use of known information and the exploration of unknown information to carry out goal-driven reasoning-based exploration and navigation.
[0054] Optionally, in one possible implementation of the second aspect, the above step of calculating the probability of each candidate node among multiple candidate nodes as an intermediate node based on the distance includes: calculating the probability of each candidate node among multiple candidate nodes as an intermediate node based on similarity, access count and distance, so as to obtain the coordinates of the intermediate node, where similarity is the similarity between the semantic features corresponding to multiple candidate nodes and the semantic features of the navigation object, access count is the number of times the terminal device accesses each candidate node, distance is the distance that the terminal device needs to travel to reach multiple candidate nodes, and the intermediate node is the node among multiple candidate nodes whose probability is greater than or equal to a first threshold.
[0055] In this possible implementation, the value of a node is determined by the path. If the mobile robot explores the same node too many times and has not found the navigation object, the mobile robot will abandon the current area and go to the unknown area to explore, thus shortening the time to move to the navigation object.
[0056] Optionally, in one possible implementation of the second aspect, the above step of calculating the probability of each candidate node among multiple candidate nodes as an intermediate node based on similarity, number of visits, and distance includes: calculating the value of each candidate node among multiple candidate nodes using the following formula:
[0057]
[0058] Where i represents one of a plurality of candidate nodes, V(i) represents the value of the candidate node, ω represents the similarity, and L dis Let m represent the distance, j represent the total number of child nodes of the current branch, N represent the total number of visits to the candidate node and its branch child nodes, n represent the number of times the terminal device visits the candidate node, and c1 and c2 are adjustment coefficients.
[0059] In this possible implementation, a backtracking problem is introduced, where the value of a node is determined by the number of times it is visited and the distance between nodes. If the mobile robot explores the same node too many times and has not yet found the navigation object, the mobile robot abandons the current area and goes to an unknown area to explore, thus shortening the time to reach the navigation object.
[0060] Optionally, in one possible implementation of the second aspect, the first multi-layer topology structure described above further includes a third-layer structure, which includes a plurality of fourth nodes and a plurality of fourth descriptors corresponding to the plurality of fourth nodes; each of the plurality of fourth nodes is associated with at least one second node group, and the second node group includes one or more second nodes; each of the plurality of fourth descriptors is used to describe the semantic features of the second node indicated by the second node group.
[0061] In this possible implementation, the first multi-layer topology may include a three-layer structure. It is understood that the number of layers in the first multi-layer topology is set according to actual needs, and the descriptors corresponding to the nodes of the previous layer are determined based on the node group of this layer related to the nodes of the previous layer, which helps to improve the accuracy of the navigation robot.
[0062] Optionally, in one possible implementation of the second aspect, the above steps further include: acquiring position information of multiple first objects; and constructing a first multi-layer topology based on the position information.
[0063] In this possible implementation, the terminal constructs a multi-layer topology structure, collects information about the surrounding environment, and updates the multi-layer topology structure simultaneously, which improves the efficiency of multi-layer topology construction and updating.
[0064] Optionally, in one possible implementation of the first aspect, the above steps: constructing a first multi-layer topology based on location information, include: constructing a multi-layer structure based on location information, the multi-layer structure including first-layer nodes and second-layer nodes; aggregating multiple first descriptors to obtain a second descriptor; associating multiple first descriptors, second descriptors, and the multi-layer structure to obtain the first multi-layer topology.
[0065] In this possible implementation, a multi-layer structure can be constructed first using location information, and then the first multi-layer topology can be obtained by associating each layer node with its corresponding descriptor.
[0066] Optionally, in one possible implementation of the second aspect, the above steps further include: sending first information to the cloud device, the first information being used to obtain a first multi-layer topology; and receiving the first multi-layer topology sent by the cloud device.
[0067] In this possible implementation, a multi-layered topology is built in the cloud to reduce the computing power consumption of terminal devices (i.e., mobile robots).
[0068] Optionally, in one possible implementation of the first aspect, the first descriptor is obtained based on the initial descriptor of the first object, the initial descriptor is determined based on the position information of the first object, and each of the multiple initial descriptors is used to represent the semantic features of the corresponding first object; the first descriptor represents the semantic features of the association between the multiple first objects.
[0069] In this possible implementation, semantic features that can be used to represent the relationships between multiple first objects are obtained through multiple initial descriptors. This improves the accuracy and generalization of navigation based on the semantic features of subsequent navigable objects.
[0070] Optionally, in one possible implementation of the first aspect, the first descriptor is obtained based on multiple initial descriptors through a first network, which is used to obtain semantic features representing the relationship between multiple first objects.
[0071] In this possible implementation, a first descriptor that can represent the semantic features of the relationship between multiple first objects is generated based on the individual initial descriptors of multiple first objects, and the first descriptor is used to represent the first object, which is beneficial to the accuracy of subsequent navigation based on the semantic features of the navigation object.
[0072] Optionally, in one possible implementation of the first aspect, the second descriptor is obtained through a second network based on multiple first descriptors, and the second network is used to obtain shared descriptors representing multiple first nodes associated with the second node.
[0073] In this possible implementation, a shared descriptor that can represent the node group associated with the second node is generated based on multiple first descriptors, which is beneficial to the accuracy of subsequent navigation based on the semantic features of the navigation object.
[0074] Optionally, in one possible implementation of the first aspect, the aforementioned third descriptor is obtained through a third network based on multiple first descriptors or multiple second descriptors, the third network being used to predict the descriptor of the third node based on the descriptors of nodes at the same level.
[0075] In this possible implementation, the third descriptor of the third node is predicted by using the descriptors of nodes at the same level. The descriptors of nodes at the same level can represent the semantic features of multiple objects having a relationship. That is, the third descriptor predicted by the relationship between objects can more accurately identify the semantic features of the third node.
[0076] The third aspect of this application provides a cloud device that can be applied to navigation scenarios such as homes, shopping malls, and airports. The method can be executed by the cloud device or by components of the cloud device (such as processors, chips, or chip systems). The cloud device includes: a receiving unit for receiving a navigation request sent by a terminal device, the navigation request indicating a navigation object for the terminal device; a determining unit for determining the coordinates of intermediate nodes based on a first multi-layer topology and the semantic features of the navigation object; the first multi-layer topology includes a first layer structure and a second layer structure, the first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes, the second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes, each of the multiple first nodes indicates a first object, the first descriptor describes the semantic features of the first object indicated by the corresponding first node, each of the multiple second nodes is associated with at least one group of first nodes, the group of first nodes includes one or more first nodes, the one or more first nodes are related to the position information of their respective corresponding first objects, each of the multiple second descriptors describes the semantic features of the first object indicated by the associated group of first nodes; the coordinates of the intermediate nodes are used to guide the terminal device to move toward the navigation object; and a sending unit for sending the coordinates to the terminal device.
[0077] Optionally, in one possible implementation of the third aspect, the aforementioned first descriptor is also used to describe the association between the corresponding first node and at least one other first node among a plurality of first nodes.
[0078] Optionally, in one possible implementation of the third aspect, the aforementioned association is used to describe at least one of the following: category relationship, functional relationship, matching relationship, and positional relationship between at least two first objects.
[0079] Optionally, in one possible implementation of the third aspect, the aforementioned cloud device further includes: a prediction unit for predicting a third node based on location information; the prediction unit is also used to predict a third descriptor of the third node based on multiple first descriptors or multiple second descriptors; a determination unit specifically used to update a first multi-layer topology based on the third node and the third descriptor to obtain a second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology; and a determination unit specifically used to determine coordinates based on the second multi-layer topology and the semantic features of the navigation object.
[0080] Optionally, in one possible implementation of the third aspect, the prediction unit described above is specifically used to predict the third node based on the Venn diagram corresponding to the location information.
[0081] Optionally, in one possible implementation of the third aspect, the aforementioned determining unit is specifically used to determine multiple candidate nodes based on the Upper Confidence Interval (UCT) algorithm, wherein the multiple candidate nodes are nodes in the first multi-layer topology; the determining unit is specifically used to calculate the probability of each candidate node among the multiple candidate nodes being an intermediate node based on the distance, so as to obtain the coordinates of the intermediate node, wherein the distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes, and the intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to a first threshold.
[0082] Optionally, in one possible implementation of the third aspect, the aforementioned determining unit is specifically used to calculate the probability of each candidate node among multiple candidate nodes being an intermediate node based on similarity, access count, and distance, so as to obtain the coordinates of the intermediate node. The similarity is the similarity between the semantic features corresponding to the multiple candidate nodes and the semantic features of the navigation object. The access count is the number of times the terminal device accesses each candidate node. The distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes. The intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to the first threshold.
[0083] Alternatively, in one possible implementation of the third aspect, the determining unit described above is specifically used to calculate the value of each candidate node among a plurality of candidate nodes using the following formula:
[0084]
[0085] Where i represents one of a plurality of candidate nodes, V(i) represents the value of the candidate node, ω represents the similarity, and L dis Let m represent the distance, j represent the total number of child nodes of the current branch, N represent the total number of visits to the candidate node and its branch child nodes, n represent the number of times the terminal device visits the candidate node, and c1 and c2 are adjustment coefficients.
[0086] Optionally, in one possible implementation of the third aspect, the navigation request further includes location information and / or scale information of a plurality of first objects, the scale information including at least one of the number of the plurality of first objects, the number of rooms in which the plurality of first objects are located, and the area of the region in which the plurality of first objects are located, the scale information being used to determine the number of layers of the first multi-layer topology.
[0087] Optionally, in one possible implementation of the third aspect, the first multi-layer topology structure described above further includes a third-layer structure, which includes a plurality of fourth nodes and a plurality of fourth descriptors corresponding to the plurality of fourth nodes; each of the plurality of fourth nodes is associated with at least one second node group, and the second node group includes one or more second nodes; each of the plurality of fourth descriptors is used to describe the semantic features of the second node indicated by the second node group.
[0088] Optionally, in one possible implementation of the third aspect, the aforementioned receiving unit is further configured to acquire the location information of multiple first objects; the cloud device further includes: a building unit configured to build a first multi-layer topology based on the location information.
[0089] Optionally, in one possible implementation of the third aspect, the aforementioned construction unit is specifically used to construct a multi-layer structure based on location information, the multi-layer structure including first-layer nodes and second-layer nodes; the construction unit is specifically used to aggregate multiple first descriptors to obtain a second descriptor; the construction unit is specifically used to associate multiple first descriptors, second descriptors and multi-layer structures to obtain a first multi-layer topology.
[0090] Optionally, in one possible implementation of the third aspect, the aforementioned receiving unit is further configured to receive the first multi-layer topology structure sent by the terminal device.
[0091] The fourth aspect of this application provides a terminal device that can be applied to navigation scenarios such as homes, shopping malls, and airports. The method can be executed by the terminal device or by a component of the terminal device (such as a processor, chip, or chip system). The terminal device can be a mobile robot (such as a sweeping robot, a handling robot, a guiding robot, etc.). The terminal device includes: a receiving unit for receiving a user's movement command, the movement command indicating movement toward a navigation object; a determining unit for determining the coordinates of intermediate nodes based on a first multi-layer topology and the semantic features of the navigation object; the first multi-layer topology includes a first layer structure and a second layer structure, the first layer structure including multiple first nodes and multiple first descriptors corresponding to the multiple first nodes, the second layer structure including multiple second nodes and multiple second descriptors corresponding to the multiple second nodes, each of the multiple first nodes indicating a first object, the first descriptors describing the semantic features of the first object indicated by the corresponding first node, each of the multiple second nodes being associated with at least one group of first nodes, the first node group including one or more first nodes, the one or more first nodes being related to the position information of their respective corresponding first objects, each of the multiple second descriptors describing the semantic features of the first object indicated by the associated group of first nodes; and a moving unit for moving toward the navigation object based on coordinates.
[0092] Optionally, in one possible implementation of the fourth aspect, the first descriptor described above is also used to describe the association between the corresponding first node and at least one other first node among a plurality of first nodes.
[0093] Optionally, in one possible implementation of the fourth aspect, the aforementioned association is used to describe at least one of the following: category relationship, functional relationship, matching relationship, and positional relationship between at least two first objects.
[0094] Optionally, in one possible implementation of the fourth aspect, the terminal device further includes: a prediction unit for predicting a third node based on location information; the prediction unit is also used to predict a third descriptor of the third node based on multiple first descriptors or multiple second descriptors; a determination unit specifically used to update a first multi-layer topology based on the third node and the third descriptor to obtain a second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology; and a determination unit specifically used to determine coordinates based on the second multi-layer topology and the semantic features of the navigation object.
[0095] Optionally, in one possible implementation of the fourth aspect, the prediction unit described above is specifically used to predict the third node based on the Venn diagram corresponding to the location information.
[0096] Optionally, in one possible implementation of the fourth aspect, the aforementioned determining unit is specifically used to determine multiple candidate nodes based on the Upper Confidence Interval (UCT) algorithm, wherein the multiple candidate nodes are nodes in the first multi-layer topology; the determining unit is specifically used to calculate the probability of each candidate node among the multiple candidate nodes being an intermediate node based on the distance, so as to obtain the coordinates of the intermediate node, wherein the distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes, and the intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to a first threshold.
[0097] Optionally, in one possible implementation of the fourth aspect, the aforementioned determining unit is specifically used to calculate the probability of each candidate node among multiple candidate nodes being an intermediate node based on similarity, access count, and distance, so as to obtain the coordinates of the intermediate node. The similarity is the similarity between the semantic features corresponding to multiple candidate nodes and the semantic features of the navigation object. The access count is the number of times the terminal device accesses each candidate node. The distance is the distance that the terminal device needs to travel to reach multiple candidate nodes. The intermediate node is the node among multiple candidate nodes whose probability is greater than or equal to the first threshold.
[0098] Alternatively, in one possible implementation of the fourth aspect, the determining unit described above is specifically used to calculate the value of each candidate node among a plurality of candidate nodes using the following formula:
[0099]
[0100] Where i represents one of a plurality of candidate nodes, V(i) represents the value of the candidate node, ω represents the similarity, and L dis Let m represent the distance, j represent the total number of child nodes of the current branch, N represent the total number of visits to the candidate node and its branch child nodes, n represent the number of times the terminal device visits the candidate node, and c1 and c2 are adjustment coefficients.
[0101] Optionally, in one possible implementation of the fourth aspect, the first multi-layer topology structure described above further includes a third-layer structure, which includes a plurality of fourth nodes and a plurality of fourth descriptors corresponding to the plurality of fourth nodes; each of the plurality of fourth nodes is associated with at least one second node group, and the second node group includes one or more second nodes; each of the plurality of fourth descriptors is used to describe the semantic features of the second node indicated by the second node group.
[0102] Optionally, in one possible implementation of the fourth aspect, the receiving unit described above is further configured to acquire position information of a plurality of first objects; the terminal device further includes: a construction unit configured to construct a first multi-layer topology based on the position information.
[0103] Optionally, in one possible implementation of the fourth aspect, the aforementioned construction unit is specifically used to construct a multi-layer structure based on location information, the multi-layer structure including first-layer nodes and second-layer nodes; the construction unit is specifically used to aggregate multiple first descriptors to obtain a second descriptor; the construction unit is specifically used to associate multiple first descriptors, second descriptors and multi-layer structures to obtain a first multi-layer topology.
[0104] Optionally, in one possible implementation of the fourth aspect, the terminal device further includes: a sending unit for sending first information to the cloud device, the first information being used to obtain a first multi-layer topology; and a receiving unit for receiving the first multi-layer topology sent by the cloud device.
[0105] The fifth aspect of this application provides a cloud device that performs the methods in the first aspect or any possible implementation thereof.
[0106] The sixth aspect of this application provides a terminal device that performs the method in the second aspect or any possible implementation thereof.
[0107] A seventh aspect of this application provides a cloud device, including: a processor coupled to a memory for storing programs or instructions, wherein when the programs or instructions are executed by the processor, the cloud device implements the methods described in the first aspect or any possible implementation thereof.
[0108] The eighth aspect of this application provides a terminal device, including: a processor coupled to a memory for storing programs or instructions, wherein when the program or instructions are executed by the processor, the terminal device implements the methods of the second aspect or any possible implementation thereof.
[0109] The ninth aspect of this application provides a navigation system that includes the cloud device of the fifth aspect and / or the terminal device of the sixth aspect, or includes the cloud device of the seventh aspect and / or the terminal device of the eighth aspect.
[0110] The tenth aspect of this application provides a computer-readable medium having a computer program or instructions stored thereon, which, when run on a computer, cause the computer to perform the methods of the first aspect or any possible implementation thereof, or cause the computer to perform the methods of the second aspect or any possible implementation thereof.
[0111] The eleventh aspect of this application provides a computer program product that, when executed on a computer, causes the computer to perform the methods in the aforementioned first aspect or any possible implementation thereof, or causes the computer to perform the methods in the aforementioned second aspect or any possible implementation thereof.
[0112] The technical effects of the third, fifth, seventh, ninth, tenth, and eleventh aspects or any one of the possible implementations can be found in the first aspect or the technical effects of different possible implementations of the first aspect, and will not be repeated here.
[0113] The technical effects of the fourth, sixth, eighth, ninth, tenth, and eleventh aspects or any one of the possible implementations can be found in the second aspect or the technical effects of different possible implementations of the second aspect, and will not be repeated here.
[0114] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: After receiving the navigation request from the terminal device, the cloud device uses the first multi-layer topology structure to determine the coordinates and sends the coordinates to the terminal device, so that the terminal device can move towards the navigation object according to the coordinates determined by the first multi-layer topology structure. By introducing a multi-layer topology structure with multi-layer node semantic association, that is, by navigating the terminal device to the navigation object through semantic features, since semantic features are not easily affected by changes in scene details, the accuracy of the terminal device navigation and the generalization of navigation in changing scenarios are improved. Attached Figure Description
[0115] Figure 1 The system architecture diagram provided in this application is shown below.
[0116] Figure 2 A schematic diagram of a chip hardware structure is provided in this application;
[0117] Figure 3 A flowchart illustrating the navigation method provided in this application;
[0118] Figure 4 A polygonal convex hull structure diagram determined based on environmental information is provided for this application;
[0119] Figure 5 A Vinio diagram based on environmental information provided for this application;
[0120] Figure 6 A schematic diagram of a multi-layer topology provided in this application;
[0121] Figure 7 A schematic diagram of an association structure of the first network, the second network, and the third network provided in this application;
[0122] Figure 8 A flowchart illustrating the construction of a multi-layer topology provided in this application;
[0123] Figure 9 A schematic diagram of a tree search structure based on a multi-layer topology provided for this application;
[0124] Figure 10 Another flowchart illustrating the navigation method provided in this application;
[0125] Figure 11 , Figure 14 Example diagrams of environmental information provided in this application;
[0126] Figure 12 , Figure 13 , Figure 15 , Figure 16Several schematic diagrams illustrating the movement of the mobile robot to the next operating node provided in this application;
[0127] Figure 17 Another flowchart illustrating the navigation method provided in this application;
[0128] Figure 18 Another flowchart illustrating the navigation method provided in this application;
[0129] Figures 19-23 Several structural diagrams of the navigation device provided in this application. Detailed Implementation
[0130] This application provides a navigation method and related equipment. By introducing a multi-layered topology with multi-level node semantic relationships and performing navigation based on this multi-layered topology, the accuracy of robot navigation is improved.
[0131] The technical solutions of the embodiments of the present invention will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0132] To facilitate understanding, the relevant terms and concepts mainly involved in the embodiments of this application will be introduced below.
[0133] 1. Neural Networks
[0134] Neural networks can be composed of neural units, which can refer to units such as X. s The arithmetic unit that takes an intercept of 1 as input can output the following:
[0135]
[0136] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For X sThe weights are denoted by b, where b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input to the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.
[0137] 2. Deep Neural Networks
[0138] A deep neural network (DNN), also known as a multilayer neural network, can be understood as a neural network with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three types based on their layer positions: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. However, deep neural networks may not necessarily include hidden layers; this is not a limitation here.
[0139] The function of each layer in a deep neural network can be expressed mathematically. To describe it: From a physical perspective, the work of each layer in a deep neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / decrease; 2. Magnification / scaling; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are... Completed, operation 4 is performed by Complete, operation 5 is implemented by α(). The term "space" is used here because the object being classified is not a single thing, but a class of things; space refers to the set of all individuals within this class of things. Here, W is the weight vector, where each value represents the weight of a neuron in that layer of the neural network. This vector W determines the spatial transformation from the input space to the output space, as described above; that is, the weights W of each layer control how the space is transformed. The purpose of training a deep neural network is to ultimately obtain the weight matrix of all layers of the trained neural network (a weight matrix formed by the vectors W from many layers). Therefore, the training process of a neural network is essentially learning how to control spatial transformation, more specifically, learning the weight matrix.
[0140] 3. Convolutional Neural Networks
[0141] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as performing convolution between the same trainable filter and an input image or a convolutional feature map. A convolutional layer is a layer of neurons in a CNN that performs convolutional processing on the input signal. In a convolutional layer of a CNN, a neuron may only be connected to some of its neighboring neurons. A convolutional layer typically contains several feature maps, each composed of rectangularly arranged neural units. Neural units within the same feature map share weights, which are the convolutional kernel. Shared weights can be understood as the way image information is extracted regardless of location. The underlying principle is that the statistical information of one part of an image is the same as that of other parts. This means that image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations in an image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0142] Convolutional kernels can be initialized as matrices of random size. During the training of the convolutional neural network, the kernels can learn to acquire appropriate weights. Furthermore, the direct benefit of shared weights is reducing the connections between layers of the convolutional neural network, while also reducing the risk of overfitting. The separation network, recognition network, detection network, depth estimation network, and other networks in the embodiments of this application can all be CNNs.
[0143] 4. Recurrent Neural Networks
[0144] In traditional neural network models, layers are fully connected, but nodes within each layer are unconnected. However, this type of ordinary neural network cannot solve many problems. For example, predicting the next word in a sentence, because words in a sentence are not independent; the preceding words are usually needed. Recurrent neural networks (RNNs) refer to a sequence where the current output is related to the previous outputs. Specifically, the network memorizes previous information, storing it in its internal state, and applies it to the calculation of the current output.
[0145] 5. Loss Function
[0146] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value and update the weight vector of each layer based on the difference. (Of course, there's usually an initialization process before the first update, where parameters are pre-configured for each layer.) For example, if the network's prediction is too high, the weight vector is adjusted to predict a lower value, and this adjustment continues until the network can predict the target value accurately. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss.
[0147] 6. Word Embedding
[0148] Word embedding can also be called "vectorization", "vector mapping", or "embedding". Formally, word embedding uses a dense vector to represent an object.
[0149] 7. Point cloud data
[0150] Point cloud data, or simply point cloud, refers to a set of points that represent the spatial distribution and surface characteristics of a target under the same spatial reference frame. After obtaining the spatial coordinates of each sampling point on the surface of an object, what is obtained is a set of points, called a point cloud.
[0151] In the embodiments of this application, point cloud data is used to characterize the three-dimensional coordinate value of each point in the point cloud data in the reference coordinate system; in addition, in some embodiments of this application, point cloud data can also be fused with the pixels of RGB images. Therefore, in some embodiments of this application, point cloud data can also be used to characterize the pixel value of each point in the point cloud data and the three-dimensional coordinate value of each point in the reference coordinate system.
[0152] 8. A priori and posterior
[0153] Prior knowledge generally refers to knowledge or experience acquired beforehand. Posterior knowledge refers to the conditional probability obtained after considering and providing relevant evidence or data. In the embodiments of this application, prior knowledge specifically refers to known data before considering observational data.
[0154] 9. Topology
[0155] Topology is an abstract structure that reflects the properties of geometric figures or spaces that remain unchanged after continuous changes in shape. It generally considers only the relationships between objects without regard to their shape and size. In this article, it specifically refers to the graphical structure describing the relationships between objects.
[0156] 10. Semantics
[0157] The literal meaning is the meaning of data. In the embodiments of this application, semantics specifically refers to higher-level data that is distinct from geometric scale coordinates and other data that conform to human logical thinking. It can also be understood as a collection of representative data, generally represented by multi-dimensional vectors.
[0158] 11. Voronoi Graph
[0159] A Vino diagram, also known as a Thiessen polygon or Dirichlet diagram, is a graph consisting of a set of continuous polygons formed by the perpendicular bisectors of lines connecting two adjacent points.
[0160] 12. Nearest Neighbor Search
[0161] Nearest neighbor search (NNS) is an optimization problem that finds the nearest point in a scale space. Commonly used NNS methods include r-disc and k-Nearest Neighbors (KNN). r-disc represents the method of finding the nearest neighbor within a circular region of radius r, while KNN represents the method of finding the k nearest neighbors.
[0162] The system architecture provided in the embodiments of this application is described below.
[0163] See appendix Figure 1This invention provides a system architecture 100. As shown in the system architecture 100, a data acquisition device 160 is used to acquire training data. In this embodiment, the training data includes at least one of first training data, second training data, and third training data. The first training data corresponds to a first network, the second training data corresponds to a second network, and the third training data corresponds to a third network. Further, for the first network, the training data includes multiple first training descriptors and multiple second training descriptors. The first training descriptor represents a single semantic feature of a training object, and the second training descriptor represents the actual semantic features between the training object and other objects (or, in other words, the second training descriptor can be used to describe the actual semantic features between the training object and its surrounding objects). For the second network, the training data includes multiple third training descriptors and a fourth training descriptor. The third training descriptor represents the actual semantic features between the training object and other objects, and the fourth training descriptor is obtained by aggregating multiple third training descriptors. In other words, the fourth training descriptor can be understood as the semantic features corresponding to a group of training objects composed of multiple training objects, and the number of fourth training descriptors can be one or more. For the third network, the training data includes a fifth training descriptor and a sixth training descriptor. The fifth training descriptor represents a single semantic feature of the training object, and the sixth training descriptor represents the semantic feature corresponding to the training object in a known region. Alternatively, the fifth training descriptor represents the actual semantic feature corresponding to a group of training objects, and the sixth training descriptor represents the actual semantic feature corresponding to a group of training objects in a known region. The training data is stored in database 130, and training device 120 trains the target model / rule 101 based on the training data maintained in database 130. The following is a simplified description of how training device 120 obtains the target model / rule 101 based on the training data: The first network is trained using the first training data as input, with the goal of the first loss function being less than a certain threshold. The first loss function represents the difference between the descriptor output by the first network and the second training descriptor. The second network is trained using the second training data as input (or, as can be understood, the input of the second network is the output of the first network), with the goal of the second loss function being less than a certain threshold. The second loss function represents the difference between the descriptor output by the second network and the fourth training descriptor. The third network is trained using third training data as input (or, as can be understood, the input to the third network can be the output of the first or second network), with the goal of obtaining a third loss function value less than a certain threshold. The third loss function is used to represent the difference between the descriptor output by the third network and the sixth training descriptor. This target model / rule 101 can be used to implement the navigation method provided in the embodiments of this application.The target model / rule 101 in this embodiment can specifically be a first network, a second network, and a third network. It should be noted that in practical applications, the training data maintained in the database 130 may not all originate from the data acquisition device 160; it may also be received from other devices. Furthermore, it should be noted that the training device 120 may not necessarily train the target model / rule 101 entirely based on the training data maintained in the database 130; it may also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.
[0164] The target model / rule 101 trained using training device 120 can be applied to different systems or devices, such as... Figure 1 The execution device 110 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, AR / VR, vehicle terminal, etc., or it can be a server or cloud platform. (See attached...) Figure 1 In this embodiment, the execution device 110 is equipped with an I / O interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140. The input data may include at least one of multiple initial descriptors, multiple first descriptors, and second descriptors in this application embodiment. It may be input by the user or come from the database. The specific details are not limited here.
[0165] The preprocessing module 113 is used to preprocess the input data received by the I / O interface 112. In this embodiment, the preprocessing module 113 can be used to perform operations such as pruning the vector dimension or size of the input data.
[0166] During the preprocessing of input data by the execution device 110, or during the calculation module 111 of the execution device 110 performing calculations and other related processes, the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0167] Finally, I / O interface 112 returns the processing result, such as the descriptor obtained over the network as described above, to client device 140, thereby providing it to the user.
[0168] It is worth noting that the training device 120 can generate corresponding target models / rules 101 based on different training data for different objectives or tasks. The corresponding target models / rules 101 can be used to achieve the above objectives or complete the above tasks, thereby providing the user with the required results.
[0169] In the appendix Figure 1In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 140 can automatically send input data to I / O interface 112. If user authorization is required for the client device 140 to automatically send input data, the user can set the corresponding permissions in the client device 140. The user can view the output results of the execution device 110 on the client device 140, which can be presented in various forms such as display, sound, or animation. The client device 140 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 140, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130.
[0170] It is worth noting that, attached Figure 1 This is merely a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationships between the devices, components, modules, etc. shown in the diagram do not constitute any limitation. For example, in the attached diagram... Figure 1 In this case, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 may also be placed in the execution device 110.
[0171] like Figure 1 As shown, a target model / rule 101 is obtained by training the training device 120. In this embodiment of the application, the target model / rule 101 may be at least one of a first network, a second network, and a third network.
[0172] The following describes a chip hardware structure provided by an embodiment of this application.
[0173] Figure 2 A chip hardware structure provided in this embodiment of the invention includes a neural network processor 20. This chip can be configured as follows: Figure 1 The execution device 110 shown is used to perform the calculations of the calculation module 111. This chip can also be placed in, for example... Figure 1 The training device 120 shown is used to complete the training work of the training device 120 and output the target model / rule 101.
[0174] The neural network processor 20 can be any processor suitable for large-scale XOR operations, such as a neural network processing unit (NPU), tensor processing unit (TPU), or graphics processing unit (GPU). Taking an NPU as an example: the neural network processor NPU40 is mounted as a coprocessor on the main central processing unit (CPU) (host CPU), and tasks are assigned by the main CPU. The core of the NPU is the arithmetic circuit 203, and the controller 204 controls the arithmetic circuit 203 to retrieve data from the memory (weight memory or input memory) and perform operations.
[0175] In some implementations, the arithmetic circuit 203 internally includes multiple process engines (PEs). In some implementations, the arithmetic circuit 203 is a two-dimensional pulsating array. The arithmetic circuit 203 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 203 is a general-purpose matrix processor.
[0176] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 202 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 201 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 208.
[0177] The vector computation unit 207 can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector computation unit 207 can be used for network computation in non-convolutional / non-FC layers of neural networks, such as pooling, batch normalization, local response normalization, etc.
[0178] In some implementations, the vector computation unit 207 can store the processed output vector into a unified buffer 206. For example, the vector computation unit 207 can apply a nonlinear function to the output of the arithmetic circuit 203, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit 207 generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as activation input to the arithmetic circuit 203, for example, for use in subsequent layers of a neural network.
[0179] The unified memory 206 is used to store input data and output data.
[0180] The weight data is directly transferred from the external memory to the input memory 201 and / or the unified memory 206 through the direct memory access controller 205 (DMAC), the weight data in the external memory is stored in the weight memory 202, and the data in the unified memory 206 is stored in the external memory.
[0181] The bus interface unit (BIU) 210 is used to enable interaction between the main CPU, DMAC and instruction fetch memory 209 via the bus.
[0182] The instruction fetch buffer 209, which is connected to the controller 204, is used to store the instructions used by the controller 204.
[0183] The controller 204 is used to call the instructions cached in the instruction memory 209 to control the operation of the computing accelerator.
[0184] Generally, the unified memory 206, input memory 201, weight memory 202, and instruction fetch memory 209 are all on-chip memories, while the external memory is memory outside the NPU. This external memory can be double data rate synchronous dynamic random access memory (DDR SDRAM), high bandwidth memory (HBM), or other readable and writable memory.
[0185] The navigation method provided in this application embodiment can be applied to robot navigation in scenarios such as homes, hotels, restaurants, shopping malls, airports, hospitals, scenic spots, and factories. When the above scenarios change (e.g., objects move or light changes), the navigation terminal device moves towards the navigation object through semantic features. Since semantic features are not easily affected by changes in scene details, the accuracy of the terminal device navigation and the generalization of navigation in changing scenarios are improved.
[0186] It is understandable that the above scenarios are just examples. In actual applications, there may be other application scenarios, which are not limited here.
[0187] The navigation method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0188] Please see Figure 3 One embodiment of the navigation method provided in this application can be applied to navigation scenarios such as homes, shopping malls, and airports. This method can be executed by a navigation device or by components of the navigation device (such as a processor, chip, or chip system). This embodiment includes steps 301 to 304.
[0189] The navigation device in this application embodiment can be a terminal device (e.g., a mobile robot) or a cloud device (or control device) that controls the movement of the mobile robot. If the navigation device is a mobile robot, the mobile robot can be a robot capable of moving and performing certain tasks in indoor environments (e.g., homes, hospitals, airports, shopping malls, and factory workshops). Specifically, the mobile robot can include a sweeping robot, a handling robot, a guiding robot, etc. The following description only uses the example of a mobile robot as the terminal device.
[0190] Step 301: Receive the user's movement command.
[0191] Navigation devices can receive movement commands from users, which instruct them to move towards the navigation object. The specific form of these movement commands can be instructions, category labels for the navigation object, or images of the navigation object, etc., and is not limited here.
[0192] If the navigation device is a mobile robot, it can obtain movement commands from user input. If the navigation device is a cloud-based device, it can directly receive user input movement commands or receive movement commands forwarded by other devices (such as mobile robots, relay devices, etc.), without further limitations here.
[0193] Optionally, the navigation device acquires environmental information, which can be information from a publicly available dataset or information obtained by scanning the surrounding environment through a mobile robot; the specific method is not limited here. This environmental information includes multiple objects and the positional relationships between them. Acquiring the navigation objects can also be understood as acquiring the semantic features of the navigation objects.
[0194] Optionally, the environmental information also includes scale information, which is used to determine the number of layers in the first multi-layer topology. This scale information can be related to the number of objects, or to factors such as range or area. For example, the more objects there are, the more layers the first topology can have. In a navigation scenario inside a building, it could also be related to the number of rooms.
[0195] In one possible implementation, the navigation device is a mobile robot. The mobile robot scans the surrounding environment to obtain perceptual information, which is then processed to obtain environmental information. The environmental information can be represented in at least one of the following formats: point cloud, red-green-blue (RGB) image, depth image, etc. This environmental information is equivalent to the information acquired by the perception module in the mobile robot, which perceives and identifies the external environment through sensors, cameras, and other devices. Sensors can include at least one of the following: global positioning system (GPS), wheel speedometer, inertial measurement unit (IMU), radio, radio frequency identification (RFID), radar (e.g., laser ranging radar), or camera.
[0196] Optionally, the aforementioned depth image can be obtained directly by the navigation device, or it can be obtained through a depth estimation network and an RGB image, or it can be sent by receiving from other devices (such as a mobile robot), or it can be obtained from a database. The specific method is not limited here.
[0197] In another possible implementation, the navigation device is a cloud device. The cloud device can obtain environmental information from a database or public dataset, or it can obtain environmental information by receiving environmental information sent by other devices such as mobile robots. The specific method is not limited here.
[0198] Step 302: Construct the first multi-layer topology. This step is optional.
[0199] The first multi-layer topology in this embodiment can be constructed by a navigation device or sent from other devices; no specific limitation is made here.
[0200] Optionally, after acquiring environmental information, the navigation device can construct a multi-layer structure based on the environmental information. This multi-layer structure includes at least one layer; if it has two layers, the multi-layer structure includes a first layer and a second layer. The first layer includes multiple first nodes, and the second layer includes multiple second nodes. Each of the multiple first nodes indicates a first object. Further, multiple first descriptors can be aggregated to obtain second descriptors, and the multiple first descriptors and second descriptors can be associated with the multi-layer structure to obtain a first multi-layer topology. The first descriptor is used to describe the semantic features of the first object indicated by the corresponding first node. Each of the multiple second nodes is associated with at least one group of first nodes. The group of first nodes includes one or more first nodes, and the one or more first nodes are related to the position information of their respective corresponding first objects. Each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated group of first nodes.
[0201] Of course, the multi-layer structure can also include a third-layer structure, which includes multiple fourth nodes and multiple fourth descriptors corresponding to the multiple fourth nodes; each of the multiple fourth nodes is associated with at least one second node group, and the second node group includes one or more second nodes; each of the multiple fourth descriptors is used to describe the semantic features of the second node indicated by the second node group. The multiple second nodes are associated with each other.
[0202] The process of the navigation device constructing the first multi-layer topology can be understood as follows: constructing a multi-layer structure based on environmental information, obtaining the descriptors corresponding to each node in the multi-layer structure, and then associating the descriptors with the multi-layer structure to obtain the first multi-layer topology.
[0203] Below, we will first describe how navigation devices construct multi-layer topology structures (taking three layers as an example) based on location information from environmental information. For example, we will describe a multi-layer structure diagram including three layers of nodes (first node - object node, second node - block node, and fourth node - region node).
[0204] Optionally, assume that the environmental information acquired by the navigation device is in the form of a color depth image (red, green, blue depth, RGBD) (i.e., RGB image and depth image). The RGB image and depth image are projected onto a 2D plane object block point cloud, and the point cloud is clustered using methods such as density-based spatial clustering of applications with noise (DBSCAN) to generate corresponding polygonal convex hulls (e.g., ...). Figure 4 (As shown). Optionally, the Veno diagram is drawn using the geometric center of the polygon's convex hull as the block node (e.g., ...). Figure 5 As shown, a Vinograph includes vertices, connecting edges, and convex hulls. Vertices correspond to regions, convex hulls correspond to blocks separated by connecting edges, and connecting edges correspond to connected regions. The vertices of the Vinograph are designated as third-layer nodes (region nodes), the polygonal convex hull as second-layer nodes (block nodes), and each first object within the polygonal convex hull as a first-layer node (object node). Region nodes are bound to a set of convex hulls (i.e., multiple block nodes) within a certain range using a nearest neighbor search method {such as the aforementioned r-disc or k-nearest neighbors (KNN) algorithm} to determine the dependency relationship between region nodes and block nodes. The first objects included in the polygonal convex hull are determined based on their positions, thus determining the dependency relationship between block nodes and object nodes. A multi-layered graph is determined based on the dependency relationships of region nodes, block nodes, and object nodes. For example, a local graph of a multi-layered graph can be shown as follows: Figure 6 As shown, region node A has a subordinate relationship with block nodes A1, A2, and A3, and block node A1 has a subordinate relationship with object nodes a1, a2, a3, and a4. This is understandable. Figure 6 The multi-layer structure diagram shown is only a partial multi-layer structure diagram. The connection relationship of each layer node is described using only the region node A, block nodes A1, A2, A3, and object nodes a1, a2, a3, and a4 as examples.
[0205] Additionally, prediction nodes (Ghosts) (i.e., third nodes) can be added to the object layer (i.e., the layer containing object nodes), block layer (i.e., the layer containing block nodes), and / or region layer (i.e., the layer containing region nodes). The environmental information does not include one or more first objects associated with this third node. The following description uses adding prediction nodes to the region layer as an example. For environments with large areas, since the environmental information acquired by the mobile robot is not complete information about the actual environment, in order to enable the robot to navigate in unknown areas, the extended edges of the aforementioned Vinograph are used as potential connected regions to determine prediction nodes. This can also be understood as using connected regions that the mobile robot cannot observe as prediction nodes. For example, Figure 5 As shown, if one end of an edge in a Vinograph lacks observation data, the end without observation data can be selected as the prediction node. Furthermore, as... Figure 6As shown, assuming there is no observation data at one end of an edge between A and B, a prediction node can be added between region node A and region node B in the region layer of the multi-layer structure graph. Of course, the mobile robot can correct (delete or add) the prediction node based on subsequent motion; that is, if both ends of the edge have observation data, the prediction node can be removed. Additionally, if a region is still blank and it is unknown whether it has been scanned, a prediction node can be added to the blank region if the blank region is larger than a threshold. It is understood that the conditions for adding prediction nodes in this embodiment can include other conditions besides the edges of the Venn diagram or blank regions where it is uncertain whether they have been scanned; these are not specifically limited here.
[0206] The multi-layered structure in the embodiments of this application can be a Veno diagram or other graph used to transform environmental information into a multi-layered structure.
[0207] Next, we will describe the operation of the navigation device in obtaining the descriptors of each node in the multi-layer structure.
[0208] Optionally, after acquiring environmental information, the navigation device can determine multiple initial descriptors for multiple first objects based on the environmental information. Then, based on the multiple initial descriptors, a first network is used to obtain multiple first descriptors. This first network is used to obtain semantic features representing the relationships between the multiple first objects.
[0209] Optionally, the navigation device can identify environmental information through template matching or classifiers to obtain multiple first objects and their positional relationships (the positional relationship can be the coordinates of the first object or the relative positional relationship between the first object and other first objects). Then, the semantic relationships (also known as object distribution relationships) between the multiple objects are abstracted and refined through manual or automatic encoding to generate vectorized semantic feature descriptors (i.e., initial descriptors: representing the individual semantic features of a single first object). These initial descriptors corresponding to multiple first objects are then directly or indirectly labeled and stored using hash tables or function fitting.
[0210] Optionally, multiple initial descriptors are input into a first network to obtain multiple first descriptors, with the number of initial descriptors corresponding one-to-one with the number of first descriptors. In this embodiment, the first network can be an unsupervised trained neural network such as a graph sampling and aggregation (GraphSAGE) algorithm. This first network is used to fuse and refine the semantic features of a single first object with the features of its surrounding objects, aggregating them into semantic features (i.e., first descriptors) that characterize the distribution relationship between the current first object and its surrounding objects. These first descriptors represent the semantic features of the association between the corresponding first node and at least one other first node among multiple first nodes. Optionally, the first descriptor can describe the association between multiple nodes, which is learned and can refer to explicit associations (e.g., adjacency relationships) or implicit associations (e.g., user habits, such as the habit of using a water cup and a water dispenser simultaneously), which is beneficial for the accuracy of subsequent semantic navigation.
[0211] It is understood that the aforementioned relationships are used to describe at least one of the following relationships between at least two first objects: category relationship, functional relationship, matching relationship, bearing relationship, positional relationship, and membership relationship. Specifically, a category relationship means that at least two first objects belong to the same or similar categories; a functional relationship means that at least two first objects have the same or similar functions; a matching relationship means that at least two first objects often work together to perform a certain function in practical applications; a bearing relationship means that at least two first objects have a bearing and being-bearing relationship; and a positional relationship means that at least two first objects have a spatial relationship, specifically referring to their adjacency within a certain range (e.g., 1 meter). For example: if one first object is a red table and the other is a black table, then the two first objects have a category relationship (i.e., both are tables). If both first objects belong to kitchen utensils, then the two first objects have a functional or category relationship. If the two first objects are a table and a chair, or a car and a parking lot respectively, then the two first objects have a matching relationship. If the two first objects are vegetables and a refrigerator respectively, then the two first objects have a bearing relationship. If the two primary objects are stationery and a table, then the two primary objects have a positional relationship (generally, stationery is placed on the table). Additionally, this relationship can also indicate user habits, preferences, etc.
[0212] Optionally, edge weights can be introduced into the aggregation function and loss function during the training of the first network. These edge weights are used to ensure high semantic feature similarity between similar objects. For example, consider a mouse and a keyboard as the first objects. Since the mouse and keyboard may be classified into the same category during subsequent object classification, edge weights are introduced here to ensure high semantic feature similarity between the mouse and keyboard.
[0213] Additionally, to ensure rapid convergence of the first network, the first descriptors of objects of the same type can be averaged (or weighted averaged) to obtain the updated first descriptor. For example, the semantic features of three cups can be averaged to obtain the first descriptor of a cup. Of course, the "same type" here needs to be determined based on the navigation target or actual needs. If the navigation target is a red cup, the semantic features of multiple red cups can be averaged to obtain the first descriptor of a red cup.
[0214] Optionally, the first multi-layer topology also includes a third node and its third descriptor. The environmental information perceived by the mobile robot does not include one or more objects corresponding to the third node. It is understood that the first multi-layer topology may include two, three, or more layers of nodes, depending on actual needs, and is not limited here. Alternatively, the first multi-layer topology can be updated based on the third node and its third descriptor to obtain a second multi-layer topology. The second multi-layer topology is then used as the first multi-layer topology for subsequent steps.
[0215] After acquiring the multi-layer structure and corresponding descriptors, the navigation device can associate multiple first descriptors with object nodes (i.e., first objects) in the object layer. The first descriptors corresponding to multiple first objects included in the polygonal convex hull are input into a second network (the second network is used to obtain shared descriptors representing multiple first nodes associated with second nodes) to obtain the shared semantic features of the polygonal convex hull (i.e., second descriptors, where one polygonal convex hull corresponds to one second descriptor), and these second descriptors are associated with block nodes in the block layer. The second descriptors of multiple block nodes belonging to region nodes are input into the second network to obtain the shared semantic features of the region nodes, and these shared semantic features are associated with the region nodes to obtain the multi-layer topology. If the multi-layer structure graph includes prediction nodes in the region layer, the second descriptors corresponding to multiple blocks belonging to the prediction nodes can be input into a third network to obtain the semantic features of the prediction nodes (i.e., third descriptors). Of course, if the prediction nodes are added in the block layer, the first descriptors corresponding to multiple first objects belonging to the prediction nodes are input into the third network to obtain the semantic features of the prediction nodes.
[0216] It's understandable that the above example uses a multi-layered structure with three layers of nodes. With two layers, the nodes could include an object layer and a block layer, meaning the operations regarding region nodes could be omitted. With more layers, the hierarchical relationship between the current layer node and the previous layer node can be further determined. The shared semantic features of the previous layer node can be obtained through the semantic features of the current layer node and the second network. Then, the semantic features of each layer node are associated with their corresponding semantic features, resulting in a multi-layered topology. By iteratively integrating information from more distant objects around the current object (from local to global, bottom-up, etc.), the scope of aggregated information is expanded.
[0217] Optionally, based on the creation and updating of the Venn diagram, the nodes of each layer in the multi-layer topology are incrementally expanded, and the corresponding semantic features are associated. The relative position information of the nodes given by the Venn diagram can also be saved.
[0218] Optionally, if a new first object is detected during the robot's movement, the mobile robot can update the multi-layer topology based on the new first object and perform the following steps using the updated topology. Alternatively, the mobile robot can send the updated multi-layer topology to the cloud device, or send the new first object (or new environmental information) to the cloud device, which will then update the multi-layer topology.
[0219] The first network has already been described and will not be repeated here. The second and third networks in this step are described below. The second network can be a self-supervised trained graph convolutional neural network (GCN), used to perform graph convolution on the fully connected subgraphs corresponding to the feature nodes of specified object groups, thereby aggregating shared semantic features corresponding to each object group. The third network is a supervised trained recurrent graph neural network (GraphRNN), which predicts the semantic features of unknown objects / blocks / regions based on the semantic features of the observed and refined object nodes / block nodes / region nodes. In this embodiment, the first, second, and third networks can be used together to obtain the semantic features corresponding to the nodes of each layer in the multi-layer structure graph. For example, the output of the first network can be used as the input of the second network, and the output of the first or second network can be used as the input of the third network. Furthermore, to facilitate subsequent calculation and navigation of object similarity, the three networks output semantic features of a unified paradigm; for example, the outputs of all three networks are 100-dimensional feature sequences.
[0220] Optionally, to emphasize the hierarchical relationships between objects and regions, the second network can incorporate both classification and hierarchical loss components into its loss function during training, simultaneously constraining them. The third network can employ cross-entropy during training to minimize the error between the predicted and actual semantic features.
[0221] Optionally, the structure of the third network can be a combination of CNN and gated recurrent unit (GRU), or a combination of CNN and long short-term memory (LSTM), or other structures, which are not limited here.
[0222] For example, the structure and connection of the first network, the second network, and the third network are as follows: Figure 7 As shown. For example Figure 7 As shown in the third network diagram, the third network predicts the semantic features of node 4 based on the shared semantic features of nodes 1, 2, and 3. Here, nodes 1, 2, 3, and 4 can represent block nodes or region nodes. It is understandable that... Figure 7 The network structure shown is just an example. In practical applications, the structure of the first, second, and third networks can take other forms, which are not limited here.
[0223] For example, the flowchart for constructing the first multi-layer topology is as follows: Figure 8 As shown.
[0224] Step 303: Determine the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object.
[0225] After acquiring the first multi-layer topology, the navigation device can determine multiple candidate nodes in the first multi-layer topology based on a search algorithm (or, in other words, construct and expand a search tree based on the search algorithm). It then determines the probability (hereinafter referred to as value, where higher value represents higher probability) of each candidate node as an intermediate node based on the similarity between the candidate nodes and the navigation object. The navigation object can be user-inputted. The similarity between each candidate node and the navigation object can be understood as the dot product of the semantic feature vectors corresponding to the nodes (i.e., the similarity between the semantic features corresponding to multiple nodes and the semantic features of the navigation object), used to calculate the similarity between nodes. Based on the similarity, the number of times nodes are visited in the first multi-layer topology, and the movement distance, the value of nodes in the multi-layer topology is calculated. Nodes with high value (either the highest value node in the multi-layer topology or nodes with a value above a certain threshold) are selected as intermediate nodes (i.e., the next movement node). After the mobile robot navigates to the next movement node, it uses that node as the current node and repeats the process. Figure 3The steps shown continue until the navigation object is reached. That is, the navigation device can continuously acquire environmental information, update the Venn diagram (nodes, node position information, etc.), update the first multi-layer topology, calculate the value of candidate nodes, and select intermediate nodes based on the value until the navigation object is reached. It can be understood that an intermediate node can be understood as any node on the path between the mobile robot's location and the navigation object; of course, an intermediate node can also indicate the navigation object.
[0226] The search algorithm in this application embodiment can be a random search algorithm, a graph-based path search algorithm, or a Monte Carlo tree search (MCTS) algorithm, etc., and is not specifically limited here. MCTS can also incorporate the upper confidence bound (UCB) to obtain the upper confidence bound apply to tree (UCT) algorithm. In other words, MCTS is actually the underlying framework, and UCB is a method used to dynamically update the value of tree nodes; when applied to MCTS, it becomes UCT.
[0227] Optionally, the value of each candidate node among multiple candidate nodes (which can also be understood as the probability of it being an intermediate node) is calculated using the following formula:
[0228]
[0229] Where i represents one of the multiple candidate nodes, V(i) represents the value of that candidate node, ω represents the similarity, and L dis Let m represent the movement distance, j represent the number of child nodes of the current branch, N represent the total number of visits to the node and its branch child nodes, n represent the number of visits to the candidate node, and c1 and c2 are adjustment coefficients set according to actual needs. The first term in the formula can be understood as the average value of the current candidate node, the second term as the first penalty term (related to the number of explorations), and the third term as the second penalty term (related to the exploration distance). The purpose of introducing the second penalty term is to encourage the mobile robot to explore unexplored areas. That is, if the mobile robot explores the same node too many times without finding a navigation object, its value will gradually decrease, making the overall value of candidate nodes with excessive exploration lower than other candidate nodes, thus causing the mobile robot to abandon the current area and explore unknown areas.
[0230] It is understandable that the above formula is just an example. In practical applications, the above formula may have other forms, which are not limited here.
[0231] For example, the movement of a mobile robot at the regional layer and the above Figure 6 The first multi-layer topology shown is used as an example for description, such as... Figure 9 As shown, node A is the root node, and nodes B, C, and the prediction node are... Figure 6 The region node in the middle region layer. In addition, node A can also connect to nodes A1, A2, A3, and other nodes in various layers that are associated with the current node, but this is not limited here.
[0232] Step 304: Move towards the navigation object based on the coordinates of the intermediate node.
[0233] Optionally, if the navigation device is a mobile robot, the mobile robot can generate a 2D scale navigation map after acquiring RGBD images. This map can be understood as a local dynamic map that changes according to the movement of the mobile robot. The 3D scale navigation map can be understood as an obstacle avoidance map. After the mobile robot determines the coordinates of intermediate nodes, it can move according to these coordinates and the 2D scale navigation map.
[0234] Optionally, if the navigation device is a cloud device, after the cloud device determines the coordinates of the intermediate node, it can send the coordinates of the intermediate node to the mobile robot, and the mobile robot moves towards the navigation object according to the coordinates of the intermediate node and the 2D scale navigation map.
[0235] Understandably, if the navigation device is a cloud-based device controlling the mobile robot, then the cloud device executes the steps in this embodiment. After determining the next motion node based on the search algorithm and multi-layer topology, the cloud device sends the coordinates of the next motion node to the mobile robot, and the mobile robot moves according to the coordinates of the next motion node. This process is repeated until the mobile robot moves to the navigation object.
[0236] In this embodiment, on the one hand, by constructing a multi-layered topology and adding prediction nodes (i.e., third nodes) within it, and ensuring that semantic features are not easily affected by changes in scene details, the understanding of the scene itself is enhanced, improving the generalization of navigation in varied and unknown scenarios. On the other hand, there is no need to use a complex sensor fusion system for the establishment and maintenance of scale maps, and the storage and updating costs of the multi-layered topology are low. Furthermore, a penalty term is introduced when calculating the value of candidate nodes. If the mobile robot explores the same node too many times without finding the navigation object, the mobile robot abandons the current area and explores an unknown area, shortening the time to reach the navigation object.
[0237] Please see Figure 10This application provides another flowchart of the navigation method. Incremental multi-layer topology modeling is performed using prior knowledge of object distribution relationships, and the value of nodes in the multi-layer topology is calculated based on the semantic features of the navigation objects. The tree structure is expanded with the robot's node as the root node, and predictive nodes are added to potential connected regions for semantic feature prediction. An exploration penalty term is constructed based on the number of node explorations and the robot's mileage, and a UCB (Unified Knowledge Base) is introduced to guide the robot in reasoning and making decisions to complete the navigation sub-objective. The observed data is the aforementioned environmental information; (a) network is the aforementioned first network; (b) network is the aforementioned second network; and (c) network is the aforementioned third network. The environmental prior knowledge graph is equivalent to the aforementioned initial descriptors and the semantic associations between multiple initial descriptors.
[0238] For example, to more intuitively illustrate the navigation method in this application's embodiments, the following is a brief description of the mobile robot's movement process at the area layer, using a home scenario as an example. Please refer to... Figures 11 to 13 This is a diagram illustrating a mobile robot moving to a certain motion node in an embodiment of this application. Figure 11 The left side shows the RGB image captured by the mobile robot. Figure 11 The right side of the image shows a depth image captured by the mobile robot. Figure 12 This is a structural diagram of the actual environment in which the mobile robot is currently located. The gray areas represent regions observed by the robot (or observed areas), and the black areas represent regions not observed by the robot (or unobserved areas). Assume the navigation object is a chest of drawers. Figure 13 The right side displays the region node, the current location of the mobile robot, and the next moving node determined based on the multi-layer topology and search algorithm. Figure 13 Block nodes 0, 1, and 2 on the left are equivalent to the block nodes associated with the area node where the mobile robot is currently located. Please refer to [link / reference]. Figures 14 to 16 This is a diagram illustrating a motion node where a mobile robot moves to a navigation object in an embodiment of this application. Figure 14 The left side shows the RGB image captured by the mobile robot. Figure 14 The right side of the image shows a depth image captured by the mobile robot. Figure 15 This is a structural diagram of the actual environment in which the mobile robot is currently located. The gray areas represent regions observed by the robot (or observed areas), and the black areas represent regions not observed by the robot (or unobserved areas). Assume the navigation object is a chest-of-drawer. Figure 16 The right side displays the region node, the current location of the mobile robot, and the next moving node determined based on the multi-layer topology and search algorithm. Figure 16 Block nodes 29 and 25 on the left are equivalent to block nodes associated with the current location of the mobile robot. During the movement of the mobile robot, the exploration value of nodes in the multi-layer topology is calculated based on the semantic features corresponding to the navigation object; the tree structure is expanded with the robot's current node as the root node, and prediction nodes are added to potential connected regions for semantic feature prediction; an exploration penalty term (i.e., the third term in the aforementioned formula) is constructed based on the number of explorations of the current node and the mobile robot's mileage, and UCB is introduced to guide the mobile robot to complete the reasoning decision for the navigation sub-objective (i.e., the next movement node).
[0239] To more intuitively demonstrate the beneficial effects of the navigation method provided in this application embodiment compared to existing navigation methods, a quantitative comparison is made between the navigation method provided in this application embodiment {or relational reasoning and voronoi local graph planning for target-driven navigation (ReVoLT)} and existing navigation methods on a dataset (the robot needs to complete the task within a 500-step action step length; finding the navigation object is considered a success). The existing navigation methods compared include: Random exploration, end-to-end reinforcement learning {e.g., RGBD + decentralized distributed proximal policy optimization (DD-PPO)}, active neural simultaneous localization and mapping (active neural SLAM), and object goal navigation using goal-oriented semantic exploration (semEXP). The comparison results are shown in Table 1.
[0240] Table 1
[0241] method SR (%) SPL DTS (meter) Random 0 0 10.3298 RGBD+DD-PPO 6.2 0.021 9.3162 active neural SLAM 32.1 0.119 7.056 semEXP 36.0 0.144 6.733 ReVoLT-i-small 66.7 0.256 0.9762 ReVoLT-i 62.5 0.102 1.0511 ReVoLT-c 85.7 0.070 0.0253
[0242] The three evaluation metrics are average success rate (SR ≤ 1), average successful path optimization rate (SPL ≤ 1), and distance to target at the end of the simulation (DTS). This embodiment can complete the exploration path efficiently in multiple target search tasks, and all three metrics are superior to other methods. Table 1 includes two seed modes of the navigation method provided in this embodiment. ReVoLT-i and ReVoLT-c represent whether the topology map memory is reset when exploring different navigation objects. ReVoLT-i means that the memory is cleared and a multi-layer topology structure is rebuilt each time; ReVoLT-c means that the existing memory can be reused for the same environment, and the multi-layer topology structure can be incrementally built and updated. ReVoLT-i-small means that, consistent with other methods, the navigation object types are limited to 6 types, and the multi-layer topology structure is rebuilt each time; non--small (i.e., ReVoLT-i and ReVoLT-c) means that the navigation object types cover the 21 object types in the dataset.
[0243] The following is a detailed description of the interaction implementation between cloud devices and terminal devices:
[0244] Please see Figure 17 Another embodiment of the navigation device in this application can be applied to navigation scenarios such as homes, shopping malls, and airports. This embodiment includes steps 1701 to 1705.
[0245] Step 1701: The terminal device sends a navigation request to the cloud device.
[0246] Optionally, the terminal device may receive a user's movement command, which instructs the user to move toward the navigation object.
[0247] The terminal device sends a navigation request to the cloud device. This navigation request is used to indicate the navigation object of the terminal device. The navigation request can be an instruction, a category label of the navigation object, or an image of the navigation object, etc., and there are no specific restrictions here.
[0248] Step 1702: The cloud device constructs the first multi-layer topology.
[0249] In this embodiment, the specific steps for the cloud device to construct the coordinates of the first multi-layer topology are the same as described above. Figure 3 Step 302 in the illustrated embodiment is similar and will not be repeated here.
[0250] Step 1703: The cloud device determines the coordinates of the intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object.
[0251] In this embodiment, the specific steps by which the cloud device determines the coordinates of intermediate nodes based on the first multi-layer topology and the semantic features of the navigation object are the same as described above. Figure 3Step 303 in the illustrated embodiment is similar and will not be repeated here.
[0252] Optionally, the cloud device can also continuously receive environmental information collected by the terminal device, and continuously update the first multi-layer topology based on the environmental information and the predicted third node, and use the updated first multi-layer topology to determine the coordinates of the intermediate nodes. The descriptions of the environmental information and the third node can be found in the preceding text. Figure 3 The corresponding descriptions are not repeated here.
[0253] Step 1704: The cloud device sends the coordinates of the intermediate node to the terminal device.
[0254] Once the cloud device determines the coordinates of the intermediate node, it can send the coordinates of the intermediate node to the terminal device.
[0255] Step 1705: The terminal device moves toward the target object based on the coordinates of the intermediate node.
[0256] In this embodiment, the specific steps for the terminal device to move towards the target object based on the coordinates of the intermediate node are the same as described above. Figure 3 Step 304 in the illustrated embodiment is similar and will not be repeated here.
[0257] In this embodiment, on the one hand, semantic features are less susceptible to changes in scene details. By introducing a multi-layer topology with multi-level node semantic associations and performing navigation based on this multi-layer topology, the accuracy of mobile robot navigation and its generalization in varying scenarios are improved. On the other hand, there is no need to use a complex sensor fusion system for the establishment and maintenance of scale maps, and the storage and updating costs of the multi-layer topology are low. Furthermore, the complex establishment of the first multi-layer topology and the determination of intermediate nodes are performed by cloud devices, which can reduce the computing power and storage space of terminal devices.
[0258] Please see Figure 18 Another embodiment of the navigation device in this application can be applied to navigation scenarios such as homes, shopping malls, and airports. This embodiment includes steps 1801 to 1805.
[0259] Step 1801: The terminal device sends a navigation request to the cloud device.
[0260] In this embodiment, the specific steps for the terminal device to send a navigation request to the cloud device are the same as described above. Figure 17 Step 1701 in the illustrated embodiment is similar and will not be repeated here.
[0261] Step 1802: The cloud device constructs the first multi-layer topology.
[0262] In this embodiment, the specific steps for the cloud device to construct the coordinates of the first multi-layer topology are the same as described above. Figure 3 Step 302 in the illustrated embodiment is similar and will not be repeated here.
[0263] Step 1803: The cloud device sends the first multi-layer topology structure to the terminal device.
[0264] After the cloud device constructs the first layer topology, it can send the first layer topology to the terminal device.
[0265] Step 1804: The terminal device determines the coordinates of the intermediate node based on the semantic features of the first multi-layer topology and the navigation object.
[0266] In this embodiment, the specific steps for the terminal device to determine the coordinates of intermediate nodes based on the first multi-layer topology and the semantic features of the navigation object are the same as described above. Figure 3 Step 303 in the illustrated embodiment is similar and will not be repeated here.
[0267] Optionally, the terminal device can continuously update the first multi-layer topology based on the collected environmental information and the predicted third node, and use the updated first multi-layer topology to determine the coordinates of the intermediate nodes. The environmental information and the description of the third node can be referred to the aforementioned... Figure 3 The corresponding descriptions are not repeated here.
[0268] Step 1805: The terminal device moves toward the target object based on the coordinates of the intermediate node.
[0269] In this embodiment, the specific steps for the terminal device to move towards the target object based on the coordinates of the intermediate node are the same as described above. Figure 3 Step 304 in the illustrated embodiment is similar and will not be repeated here.
[0270] In this embodiment, on the one hand, semantic features are less susceptible to changes in scene details. By introducing a multi-layer topology with multi-level node semantic associations and performing navigation based on this multi-layer topology, the accuracy of mobile robot navigation and its generalization in varying scenarios are improved. On the other hand, there is no need to use a complex sensor fusion system for the establishment and maintenance of scale maps, and the storage and updating costs of the multi-layer topology are low. Furthermore, the complex establishment of the first multi-layer topology is performed by cloud devices, which can reduce the computing power and storage space of terminal devices.
[0271] Of course, the interaction process between terminal devices and cloud devices includes the above. Figure 17 and Figure 18Beyond the illustrated process, other possible implementation methods exist, such as: the terminal device constructs a first multi-layer topology structure and sends it to the cloud device. The cloud device determines the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology structure and the navigation object, and sends the coordinates of the intermediate nodes back to the terminal device. The terminal device then moves towards the target object based on the coordinates of the intermediate nodes. In this method, the terminal device can update the multi-layer topology structure in real time based on collected environmental information and use the updated multi-layer topology structure to determine intermediate nodes. This reduces the data transmission overhead between the terminal device and the cloud device for updating the topology structure, thus improving the time it takes for the terminal device to move to the navigation object.
[0272] In this application embodiment, no limitation is made on the method by which the terminal device and the cloud device jointly perform navigation.
[0273] The navigation method in the embodiments of this application has been described above. The navigation device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 19 One embodiment of the navigation device in this application includes: a perception module, an environmental prior knowledge utilization module, a semantic space association and topology construction module, and an inference and decision-making module. The perception module is used for multi-dimensional (2D, 3D) data acquisition (images / depth / point clouds) and the detection and recognition of objects in the environment. The environmental prior knowledge utilization module is used for storing and encoding environmental prior knowledge, querying and reproducing posterior features, and predicting features in unknown areas. The semantic space association and topology construction module is used for spatial partitioning and clustering, and constructing a multi-layer topology structure based on the spatial association results. The inference and decision-making module is used for creating and updating Monte Carlo trees, calculating the UCB value of tree nodes, and determining the next moving node.
[0274] Please see Figure 20 One embodiment of the navigation device in this application can be applied to navigation scenarios such as homes, shopping malls, and airports. This navigation device can be a cloud-based device. The navigation device includes:
[0275] The receiving unit 2001 is used to receive a navigation request sent by the terminal device, the navigation request being used to indicate the navigation object of the terminal device;
[0276] The determining unit 2002 is used to determine the coordinates of intermediate nodes based on the semantic features of a first multi-layer topology and the navigation object. The first multi-layer topology includes a first-layer structure and a second-layer structure. The first-layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes. The second-layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes. Each of the multiple first nodes indicates a first object. The first descriptors are used to describe the semantic features of the first object indicated by the corresponding first node. Each of the multiple second nodes is associated with at least one group of first nodes. The group of first nodes includes one or more first nodes. The one or more first nodes are related to the position information of their respective corresponding first objects. Each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated group of first nodes. The coordinates of the intermediate nodes are used by the navigation terminal device to move towards the navigation object.
[0277] The sending unit 2003 is used to send coordinates to the terminal device.
[0278] Optionally, the navigation device in this embodiment further includes: a prediction unit 2004, used to predict a third node based on location information;
[0279] Optionally, the prediction unit 2004 is also used to predict the third descriptor of the third node based on multiple first descriptors or multiple second descriptors;
[0280] Optionally, the navigation device in this embodiment further includes a construction unit 2005, used to construct a first multi-layer topology based on location information.
[0281] In this embodiment, the operations performed by each unit in the navigation device are the same as those described above. Figure 3 , Figure 17 and Figure 18 The operations performed by the cloud device in the illustrated embodiment are similar and will not be described in detail here.
[0282] In this embodiment, on the one hand, by introducing a multi-layer topology structure with multi-level node semantic associations, the terminal device is guided to move towards the navigation object through semantic features. Since semantic features are not easily affected by changes in scene details, the accuracy of the terminal device's navigation and its generalization ability in variable scenarios are improved. On the other hand, the multi-layer topology structure is constructed by the construction unit 2005, and a prediction node (i.e., a third node) is added to the multi-layer topology structure by the prediction unit 2004. The semantic features are not easily affected by changes in scene details, enhancing the understanding of the scene itself and improving the generalization ability of navigation in variable and unknown scenarios. Furthermore, there is no need to use a complex sensor fusion system for the establishment and maintenance of scale maps, resulting in lower storage and update costs for the multi-layer topology structure. Additionally, a penalty term is introduced when calculating node value. If the mobile robot explores the same node too many times without finding the navigation object, the mobile robot abandons the current area and explores an unknown area, shortening the time to reach the navigation object.
[0283] Please see Figure 21 One embodiment of the navigation device in this application can be applied to navigation scenarios such as homes, shopping malls, and airports. The navigation device can be a terminal device (e.g., a mobile robot). The navigation device includes:
[0284] The receiving unit 2101 is used to receive the user's movement command, which is used to instruct movement toward the navigation object;
[0285] The determining unit 2102 is used to determine the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object; the first multi-layer topology includes a first layer structure and a second layer structure, the first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes, the second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes, each of the multiple first nodes indicates a first object, the first descriptor is used to describe the semantic features of the first object indicated by the corresponding first node, each of the multiple second nodes is associated with at least one first node group, the first node group includes one or more first nodes, the one or more first nodes are related to the position information of their respective corresponding first objects, and each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated first node group;
[0286] The moving unit 2103 is used to move towards the navigation object based on coordinates.
[0287] Optionally, the navigation device in this embodiment further includes: a prediction unit 2104, used to predict a third node based on location information;
[0288] Optionally, the prediction unit 2104 is also used to predict the third descriptor of the third node based on multiple first descriptors or multiple second descriptors;
[0289] Optionally, the navigation device in this embodiment further includes a construction unit 2105, used to construct a first multi-layer topology based on location information.
[0290] Optionally, the navigation device in this embodiment further includes: a sending unit 2106, used to send first information to the cloud device, the first information being used to obtain a first multi-layer topology;
[0291] In this embodiment, the operations performed by each unit in the navigation device are the same as those described above. Figure 3 , Figure 17 and Figure 18 The operations performed by the terminal device in the illustrated embodiment are similar and will not be described again here.
[0292] In this embodiment, a multi-layered topology with semantic associations among nodes is introduced. This means that movement to the navigation object is driven by semantic features. Since these semantic features are less affected by changes in scene details, the accuracy of robot navigation and its generalization ability in varying scenarios are improved. Furthermore, the multi-layered topology is constructed by the construction unit 2105, and a prediction node (i.e., a third node) is added to the topology by the prediction unit 2104. The semantic features are also less affected by changes in scene details, enhancing the understanding of the scene itself and improving the generalization ability of navigation in varied and unknown scenarios. Additionally, there is no need to use a complex sensor fusion system for building and maintaining scale maps, resulting in lower storage and update costs for the multi-layered topology. Moreover, a penalty term is introduced when calculating node value. If the mobile robot explores the same node too many times without finding the navigation object, it abandons the current area and explores an unknown area, shortening the time to reach the navigation object.
[0293] See Figure 22 This application provides a schematic diagram of another navigation device. This navigation device can be a mobile robot or a cloud-based device for a mobile robot. The navigation device may include a processor 2201, a memory 2202, and a communication interface 2203. The processor 2201, memory 2202, and communication interface 2203 are interconnected via lines. The memory 2202 stores program instructions and data.
[0294] The aforementioned are stored in memory 2202 Figure 3 , Figure 17 and Figure 18 In the corresponding implementation, the steps executed by the cloud device include program instructions and data.
[0295] Processor 2201, for performing the aforementioned Figure 3 , Figure 17 and Figure 18 The steps performed by the cloud device are shown in any of the embodiments illustrated.
[0296] Communication interface 2203 can be used to receive and send data, and to perform the aforementioned tasks. Figure 3 , Figure 17 and Figure 18 The steps related to acquiring, sending, and receiving in any of the embodiments shown.
[0297] In one implementation, the navigation device may include, relative to... Figure 22 More or fewer components are merely illustrative in this application and are not intended to limit the scope of the application.
[0298] See Figure 23 This application provides a schematic diagram of another navigation device. This navigation device can be a mobile robot. Specifically, it can be a sweeping robot, a handling robot, a guiding robot, etc., without limitation. Specifically, the navigation device includes: a receiver 2301, a transmitter 2302, a processor 2303, and a memory 2304 (wherein, the number of processors 2303 can be one or more). Figure 23 (Taking a processor as an example), the processor 2303 may include an application processor 23031 and a communication processor 23032. In some embodiments of this application, the receiver 2301, transmitter 2302, processor 2303, and memory 2304 may be connected via a bus or other means.
[0299] Memory 2304 may include read-only memory and random access memory, and provides instructions and data to processor 2303. A portion of memory 2304 may also include non-volatile random access memory (NVRAM). Memory 2304 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0300] The processor 2303 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.
[0301] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 2303. The processor 2303 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by the integrated logic circuitry in the hardware of the processor 2303 or by instructions in software form. The processor 2303 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or microcontroller, a vision processing unit (VPU), a tensor processing unit (TPU), or other processors suitable for AI computation. It may further include application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 2303 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 2304, and processor 2303 reads information from memory 2304 and, in conjunction with its hardware, completes the steps of the above method.
[0302] Receiver 2301 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 2302 can be used to output digital or character information through the first interface; transmitter 2302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 2302 may also include a display device such as a display screen.
[0303] In one implementation, the navigation device may include, relative to... Figure 23 More or fewer components are merely illustrative in this application and are not intended to limit the scope of the application.
[0304] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0305] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0306] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented wholly or partially through software, hardware, firmware, or any combination thereof.
[0307] When an integrated unit is implemented using software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0308] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
Claims
1. A navigation method, characterized in that, The method includes: Receive a navigation request sent by a terminal device, the navigation request being used to indicate the navigation object of the terminal device; The coordinates of intermediate nodes are determined based on a first multi-layer topology and the semantic features of the navigation object. The first multi-layer topology includes a first layer structure and a second layer structure. The first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes. The second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes. Each of the multiple first nodes indicates a first object. The first descriptor is used to describe the semantic features of the first object indicated by the corresponding first node. Each of the multiple second nodes is associated with at least one group of first nodes. The first group of first nodes includes one or more first nodes, and the one or more first nodes are related to the position information of their respective first objects. Each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated first group of first nodes. The coordinates of the intermediate nodes are used to navigate the terminal device to move towards the navigation object. Send the coordinates to the terminal device; The determination of the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object includes: Multiple candidate nodes are determined based on the Upper Confidence Interval (UCT) algorithm, and these multiple candidate nodes are nodes in the first multi-layer topology. The probability of each candidate node among the multiple candidate nodes being the intermediate node is calculated based on similarity, number of visits, and distance, so as to obtain the coordinates of the intermediate node. The similarity is the similarity between the multiple candidate nodes and the navigation object. The number of visits is the number of visits to the multiple first nodes in the first multi-layer topology. The distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes. The intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to a first threshold.
2. The method according to claim 1, characterized in that, The first descriptor is also used to describe the association between the corresponding first node and at least one other first node among the plurality of first nodes.
3. The method according to claim 2, characterized in that, The association relationship is used to describe at least one of the following: category relationship, functional relationship, matching relationship, and positional relationship between at least two first objects.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Predict the third node based on the location information; Predict the third descriptor of the third node based on the plurality of first descriptors or the plurality of second descriptors; The determination of the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object includes: The first multi-layer topology is updated based on the third node and the third descriptor to obtain the second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology. The coordinates are determined based on the second multi-layer topology and the semantic features of the navigation object.
5. The method according to claim 4, characterized in that, The prediction of the third node based on the location information includes: The third node is predicted based on the Venn diagram corresponding to the location information.
6. The method according to any one of claims 1 to 3, characterized in that, The navigation request also includes location information and / or scale information of multiple first objects. The scale information includes at least one of the number of multiple first objects, the number of rooms where the multiple first objects are located, and the area of the region where the multiple first objects are located. The scale information is used to determine the number of layers of the first multi-layer topology.
7. The method according to any one of claims 1 to 3, characterized in that, The first multi-layer topology also includes a third layer structure, which includes a plurality of fourth nodes and a plurality of fourth descriptors corresponding to the plurality of fourth nodes; each of the plurality of fourth nodes is associated with at least one second node group, the second node group including one or more second nodes; each of the plurality of fourth descriptors is used to describe the semantic features of the second node indicated by the second node group.
8. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the position information of multiple first objects; The first multi-layer topology is constructed based on the location information.
9. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Receive the first multi-layer topology structure sent by the terminal device.
10. A navigation method, characterized in that, The method includes: Receive a user's movement command, the movement command being used to instruct movement toward a navigation object; The coordinates of intermediate nodes are determined based on a first multi-layer topology and the semantic features of the navigation object. The first multi-layer topology includes a first layer structure and a second layer structure. The first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes. The second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes. Each of the multiple first nodes indicates a first object. The first descriptor is used to describe the semantic features of the first object indicated by the corresponding first node. Each of the multiple second nodes is associated with at least one group of first nodes. The first group of first nodes includes one or more first nodes. The one or more first nodes are related to the position information of their respective corresponding first objects. Each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated group of first nodes. Move towards the navigation object based on the coordinates; The determination of the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object includes: Multiple candidate nodes are determined based on the Upper Confidence Interval (UCT) algorithm, and these multiple candidate nodes are nodes in the first multi-layer topology. The probability of each candidate node among the multiple candidate nodes being the intermediate node is calculated based on similarity, number of visits, and distance, so as to obtain the coordinates of the intermediate node. The similarity is the similarity between the multiple candidate nodes and the navigation object. The number of visits is the number of visits to the multiple first nodes in the first multi-layer topology. The distance is the distance that the terminal device needs to travel to reach the multiple candidate nodes. The intermediate node is the node among the multiple candidate nodes whose probability is greater than or equal to a first threshold.
11. The method according to claim 10, characterized in that, The first descriptor is also used to describe the association between the corresponding first node and at least one other first node among the plurality of first nodes.
12. The method according to claim 10, characterized in that, The method further includes: Predict the third node based on the location information; Predict the third descriptor of the third node based on the plurality of first descriptors or the plurality of second descriptors; The determination of the coordinates of intermediate nodes based on the semantic features of the first multi-layer topology and the navigation object includes: The first multi-layer topology is updated based on the third node and the third descriptor to obtain the second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology. The coordinates are determined based on the second multi-layer topology and the semantic features of the navigation object.
13. The method according to claim 12, characterized in that, The prediction of the third node based on the location information includes: The third node is predicted based on the Venn diagram corresponding to the location information.
14. The method according to any one of claims 10 to 13, characterized in that, The first multi-layer topology also includes a third-layer structure, which includes a plurality of fourth nodes and a plurality of fourth descriptors corresponding to the plurality of fourth nodes; Each of the plurality of fourth nodes is associated with at least one group of second nodes, the group of second nodes including one or more second nodes; Each of the plurality of fourth descriptors is used to describe the semantic features of the second node indicated by the second node group.
15. The method according to any one of claims 10 to 13, characterized in that, The method further includes: Obtain the position information of multiple first objects; The first multi-layer topology is constructed based on the location information.
16. The method according to any one of claims 10 to 13, characterized in that, The method further includes: Send first information to the cloud device, the first information being used to obtain the first multi-layer topology; Receive the first multi-layer topology structure sent by the cloud device.
17. A cloud device, characterized in that, The cloud device includes: A receiving unit is configured to receive a navigation request sent by a terminal device, wherein the navigation request is used to indicate the navigation object of the terminal device; A determining unit is configured to determine the coordinates of intermediate nodes based on a first multi-layer topology and the semantic features of the navigation object. The first multi-layer topology includes a first layer structure and a second layer structure. The first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes. The second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes. Each of the multiple first nodes indicates a first object. The first descriptors are used to describe the semantic features of the first object indicated by the corresponding first node. Each of the multiple second nodes is associated with at least one group of first nodes. The first group of first nodes includes one or more first nodes, and the one or more first nodes are related to the position information of their respective corresponding first objects. Each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated group of first nodes. The coordinates of the intermediate nodes are used to navigate the terminal device to move towards the navigation object. A sending unit is configured to send the coordinates to the terminal device; The determining unit is specifically used to determine multiple candidate nodes based on the Upper Confidence Interval (UCT) algorithm, wherein the multiple candidate nodes are nodes in the first multi-layer topology. The determining unit is specifically used to calculate the probability of each candidate node among the plurality of candidate nodes being the intermediate node based on similarity, number of visits, and distance, so as to obtain the coordinates of the intermediate node. The similarity is the similarity between the plurality of candidate nodes and the navigation object. The number of visits is the number of visits to the plurality of first nodes in the first multi-layer topology structure. The distance is the distance that the terminal device needs to travel to reach the plurality of candidate nodes. The intermediate node is the node among the plurality of candidate nodes whose probability is greater than or equal to a first threshold.
18. The device according to claim 17, characterized in that, The first descriptor is also used to describe the association between the corresponding first node and at least one other first node among the plurality of first nodes.
19. The device according to claim 17 or 18, characterized in that, The cloud device also includes: A prediction unit is used to predict a third node based on the location information; The prediction unit is further configured to predict the third descriptor of the third node based on the plurality of first descriptors or the plurality of second descriptors; The determining unit is specifically used to update the first multi-layer topology based on the third node and the third descriptor to obtain a second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology. The determining unit is specifically used to determine the coordinates based on the second multi-layer topology and the semantic features of the navigation object.
20. A terminal device, characterized in that, The terminal device includes: A receiving unit is configured to receive a user's movement command, the movement command being used to instruct movement toward a navigation object; A determining unit is configured to determine the coordinates of intermediate nodes based on a first multi-layer topology and the semantic features of the navigation object. The first multi-layer topology includes a first layer structure and a second layer structure. The first layer structure includes multiple first nodes and multiple first descriptors corresponding to the multiple first nodes. The second layer structure includes multiple second nodes and multiple second descriptors corresponding to the multiple second nodes. Each of the multiple first nodes indicates a first object. The first descriptors are used to describe the semantic features of the first object indicated by the corresponding first node. Each of the multiple second nodes is associated with at least one group of first nodes. The first group of first nodes includes one or more first nodes. The one or more first nodes are related to the position information of their respective corresponding first objects. Each of the multiple second descriptors is used to describe the semantic features of the first object indicated by the associated group of first nodes. A movement unit, used to move toward the navigation object based on the coordinates; The determining unit is specifically used to determine multiple candidate nodes based on the Upper Confidence Interval (UCT) algorithm, wherein the multiple candidate nodes are nodes in the first multi-layer topology. The determining unit is specifically used to calculate the probability of each candidate node among the plurality of candidate nodes being the intermediate node based on similarity, number of visits, and distance, so as to obtain the coordinates of the intermediate node. The similarity is the similarity between the plurality of candidate nodes and the navigation object. The number of visits is the number of visits to the plurality of first nodes in the first multi-layer topology structure. The distance is the distance that the terminal device needs to travel to reach the plurality of candidate nodes. The intermediate node is the node among the plurality of candidate nodes whose probability is greater than or equal to a first threshold.
21. The device according to claim 20, characterized in that, The first descriptor is also used to describe the association between the corresponding first node and at least one other first node among the plurality of first nodes.
22. The device according to claim 20 or 21, characterized in that, The terminal device also includes: A prediction unit is used to predict a third node based on the location information; The prediction unit is further configured to predict the third descriptor of the third node based on the plurality of first descriptors or the plurality of second descriptors; The determining unit is specifically used to update the first multi-layer topology based on the third node and the third descriptor to obtain a second multi-layer topology, wherein the third node belongs to the first layer structure and / or the second layer structure of the second multi-layer topology. The determining unit is specifically used to determine the coordinates based on the second multi-layer topology and the semantic features of the navigation object.
23. A navigation device, characterized in that, include: A processor coupled to a memory for storing programs or instructions that, when executed by the processor, cause the navigation device to perform the method as claimed in any one of claims 1 to 9, or the method as claimed in any one of claims 10 to 16.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 9, or cause the computer to perform the method as described in any one of claims 10 to 16.
25. A computer program product, characterized in that, When the computer program product is executed on a computer, it causes the computer to perform the method as described in any one of claims 1 to 9, or causes the computer to perform the method as described in any one of claims 10 to 16.