Hierarchical semantic reasoning-based autonomous navigation intelligent method and related device

Through a method based on hierarchical semantic reasoning, the room-level and object-level co-occurrence and spatial exploration value are comprehensively evaluated, and the efficiency and adaptability of existing navigation methods in sparse targets and complex environments are solved, and efficient and accurate object navigation is achieved.

CN120252686AActive Publication Date: 2025-07-04BEIHANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510419097.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing zero-sample object navigation method has insufficient success rate and redundant paths in sparse target scenarios, and is susceptible to environmental noise in complex home scenarios, and has poor adaptability to dynamic environments.

Method used

Using a method based on hierarchical semantic reasoning, the best prospective point is determined through a comprehensive evaluation of room-level co-occurrence scores, object-level co-occurrence scores, and spatial exploration scores, and the navigation path is optimized in combination with an adaptive weighting mechanism.

Benefits of technology

It improves navigation efficiency and environmental adaptability, reduces repeated exploration and path oscillation, and achieves efficient exploration of different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120252686A_ABST
    Figure CN120252686A_ABST
Patent Text Reader

Abstract

The invention provides an autonomous navigation intelligent method based on hierarchical semantic reasoning and a related device. Determining a plurality of look-ahead points to be explored from the exploration area; for each look-ahead point, obtaining a plurality of exploration scores of the look-ahead point; wherein the plurality of exploration scores comprise a room-level co-occurrence score, an object-level co-occurrence score and a space exploration score; according to the respective weights of the plurality of exploration scores, obtaining a comprehensive score of the look-ahead point; and determining the look-ahead point with the highest comprehensive score as the optimal look-ahead point for the exploration equipment to continue to explore. Therefore, the method accurately evaluates the local environment of the look-ahead point and the spatial co-occurrence of the target object from two levels of room level and article level, introduces a spatial exploration score, balances the semantic reasoning result and exploration efficiency through an adaptive weighting mechanism, and can achieve the detection and updating of the article information without data collection and training, thereby improving the detection efficiency. The navigation efficiency of the exploration equipment and the adaptability to different environments can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of object navigation. Specifically, it relates to an autonomous navigation intelligent method and related device based on hierarchical semantic reasoning. Background Art

[0002] The current mainstream methods in the zero-shot object navigation field can be divided into three technical routes: geometric exploration, vision-language model-based methods, and large language model-based methods. Among them, the geometric exploration-based methods include the Frontier-Based Exploration (FBE) method and the navigation exploration method based on the Voronoi diagram. These methods explore the environment by constructing an occupancy map to generate boundary regions. Although pure geometric strategies can ensure the basic exploration efficiency, they completely ignore semantic information, resulting in a significant blindness in target search, and the success rate is less than 40% in sparse target scenarios. The vision-language model-based methods include the CLIP on Wheels (CoW) method and the Vision-Language Frontier Mapping (VLFM) method. These methods achieve semantic mapping through open-vocabulary detection or language-driven value maps. However, their single-layer semantic reasoning mechanism only relies on target category label matching and does not model the spatial co-occurrence relationship between objects, making them vulnerable to environmental noise interference and generating path oscillations in complex home scenarios, specifically manifested as repeated exploration of already explored locations. The large language model-based methods include the L3MVN and OpenFMNav methods. These methods attempt to infer the target location using a pre-trained common sense knowledge base, but are limited by the lag in static knowledge update and cloud processing delay, and it is difficult to meet the real-time requirements of dynamic environments.

[0003] Therefore, how to improve the navigation efficiency of exploration devices and their adaptability to different environments has become an urgent problem to be solved. Summary of the Invention

[0004] To overcome all the deficiencies in the prior art, the present application provides an autonomous navigation intelligent method and related device based on hierarchical semantic reasoning, specifically including:

[0005] In a first aspect, the present application provides an autonomous navigation intelligent method based on hierarchical semantic reasoning, and the method includes:

[0006] Determine multiple look-ahead points to be explored from the exploration area;

[0007] For each of the forward-looking points, obtain multiple exploration scores of the forward-looking point, where the multiple exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighborhood objects within a preset first range from the forward-looking point and the target object appear in multiple candidate rooms simultaneously. The object-level co-occurrence score represents the probability that the target object and all neighborhood objects appear simultaneously. The space exploration score represents the value of evaluating the exploration of the exploration device from the current position to the forward-looking point only from spatial factors;

[0008] Obtain the comprehensive score of the forward-looking point according to the weights of the multiple exploration scores;

[0009] Determine the forward-looking point with the highest comprehensive score as the best forward-looking point for the exploration device to continue to explore.

[0010] In a second aspect, the present application provides an autonomous navigation intelligent device based on hierarchical semantic reasoning. The device includes:

[0011] A forward-looking point determination module, configured to determine multiple forward-looking points to be explored from the exploration area;

[0012] A forward-looking point evaluation module, configured to, for each of the forward-looking points, obtain multiple exploration scores of the forward-looking point, where the multiple exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighborhood objects within a preset first range from the forward-looking point and the target object appear in multiple candidate rooms simultaneously. The object-level co-occurrence score represents the probability that the target object and all neighborhood objects appear simultaneously. The space exploration score represents the value of evaluating the exploration of the exploration device from the current position to the forward-looking point only from spatial factors;

[0013] The forward-looking point evaluation module is further configured to obtain the comprehensive score of the forward-looking point according to the weights of the multiple exploration scores;

[0014] A forward-looking point screening module, configured to determine the forward-looking point with the highest comprehensive score as the best forward-looking point for the exploration device to continue to explore.

[0015] In a third aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the autonomous navigation intelligent method based on hierarchical semantic reasoning.

[0016] In a fourth aspect, the present application provides an electronic device, which includes a processor and a memory. The memory stores a computer program, which, when executed by the processor, implements the autonomous navigation intelligent method based on hierarchical semantic reasoning.

[0017] Compared with the prior art, the present application has the following beneficial effects:

[0018] The present application provides an autonomous navigation intelligent method based on hierarchical semantic reasoning and related devices. Multiple forward-looking points to be explored are determined from the exploration area; for each forward-looking point, multiple exploration scores of the forward-looking point are obtained; wherein, the multiple exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score; according to the respective weights of the multiple exploration scores, a comprehensive score of the forward-looking point is obtained; the forward-looking point with the highest comprehensive score is determined as the best forward-looking point for the exploration device to continue to explore. In this way, the method accurately evaluates the spatial co-occurrence of the local environment of the forward-looking point and the target object from two levels of room level and item level, and at the same time introduces a space exploration score, balances the semantic reasoning result and the exploration efficiency through an adaptive weighting mechanism, and moreover, can detect and update item information without data collection and training. Therefore, it can improve the navigation efficiency of the exploration device and the adaptability to different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of the autonomous navigation intelligent method based on hierarchical semantic reasoning provided by the embodiment of the present application;

[0021] Figure 2 It is a complete principle schematic diagram of the autonomous navigation intelligent method based on hierarchical semantic reasoning provided by the embodiment of the present application;

[0022] Figure 3 It is a schematic structural diagram of the autonomous navigation intelligent device based on hierarchical semantic reasoning provided by the embodiment of the present application;

[0023] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application (hereinafter simply referred to as "these embodiments") clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0025] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application that is claimed, but merely represents selected embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.

[0026] It should be noted that: like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0027] In the description of this application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use. It is only for the convenience of describing this application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.

[0028] In addition, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0029] In addition, terms such as "horizontal", "vertical", "overhanging", etc. do not mean that the components are required to be absolutely horizontal or overhanging, but may be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but may be slightly inclined.

[0030] In the description of the present application, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0031] Based on the above statement, as introduced in the background art, although the method based on geometric exploration can achieve basic environment exploration through boundary region segmentation, it completely depends on the geometric features of the occupancy map and ignores the relevance between the target object and the scene semantics. For example, when the target object is occluded or distributed in a sparse area, due to the lack of a semantic guidance mechanism, the average search success rate of such methods decreases, and the path redundancy is relatively high.

[0032] Although the method based on the vision-language model can build a semantic map, its single-layer semantic reasoning mechanism only generates a navigation path through target category label matching, without modeling the spatial co-occurrence relationship between objects and the room-level semantic distribution features. This leads to the situation that when searching for a "microwave oven" in a kitchen scene, the existing methods may mislead to the location where there are semantically similar objects. For example, the location where the "oven" is located. However, in some cases, the "microwave oven" and the "oven" are not placed together, and this method does not memorize the explored locations. When the "oven" is seen within the field of view again, it will go to the location where the "oven" is located again for exploration, resulting in the phenomenon of path oscillation.

[0033] The method based on the large language model is limited by the lag in the update of the static common sense knowledge base and the cloud processing delay. For example, OpenFMNav needs to upload the perception data to the large language model in the cloud for processing, resulting in decision-making delays.

[0034] Therefore, the existing methods generally have the problem of flat semantic reasoning. They neither establish a hierarchical association between the room-level semantic distribution and the target object, nor have a collaborative optimization mechanism for semantic clues and geometric exploration strategies, resulting in the navigation efficiency and adaptability being difficult to meet the actual deployment requirements.

[0035] Based on the discovery of the above technical problems, after creative labor, the following technical solutions are proposed to solve or improve the above problems. It should be noted that the defects existing in the above solutions in the prior art are the results obtained after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be regarded as contributions made to the present application in the process of invention and creation, and should not be understood as the technical content known to those skilled in the art.

[0036] In view of the discovery of the above problems, this embodiment provides an autonomous navigation intelligent method based on hierarchical semantic reasoning. As Figure 1 shown, the method includes:

[0037] S1. Determine multiple look-ahead points to be explored from the exploration area.

[0038] S2. For each look-ahead point, obtain multiple exploration scores of the look-ahead point.

[0039] Among them, the multiple exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighborhood objects within a preset first range from the look-ahead point and the target object appear in multiple candidate rooms simultaneously. The object-level co-occurrence score represents the probability that the target object and all neighborhood objects appear simultaneously. The space exploration score represents the value of evaluating the exploration of the exploration device from the current position to the look-ahead point only from the spatial factor.

[0040] S3. According to the weights of the multiple exploration scores respectively, obtain the comprehensive score of the look-ahead point.

[0041] S4. Determine the look-ahead point with the highest comprehensive score as the best look-ahead point for the exploration device to continue to explore.

[0042] In this way, the method accurately evaluates the spatial co-occurrence of the local environment of the look-ahead point and the target object from two levels of the room level and the item level. At the same time, a space exploration score is introduced, and the semantic reasoning result and the exploration efficiency are balanced through an adaptive weighting mechanism. Moreover, the detection and update of item information can be realized without data collection and training. Therefore, the navigation efficiency of the exploration device and the adaptability to different environments can be improved.

[0043] For the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in this embodiment, the electronic device for implementing this method can be, but is not limited to, a host computer in communication connection with the exploration device and a controller integrated with the exploration device body. Among them, the host computer can be, but is not limited to, a mobile terminal, a tablet computer, a laptop computer, a desktop computer, etc., as long as it can provide sufficient computing power for the navigation of the exploration device.

[0044] It should also be understood that the exploration device in this embodiment can be, but is not limited to, a robot (for example, an autonomous mobile robot, a quadruped robot, a wheeled robot, a tracked robot), a drone (for example, a quadrotor drone, a fixed-wing drone), etc.

[0045] Exemplarily, taking a wheeled robot as an example, when a task of searching for a "TV set" in a house is given to the wheeled robot, after the wheeled robot enters the house, it will use a depth camera to scan the environment and construct a two-dimensional or three-dimensional map in combination with odometer data; then, the wheeled robot constructs a navigation route based on a built-in search strategy to explore in the house until it finds the "TV set" located at a certain position in the house. Of course, the exploration area in this application is not limited to indoor scenarios, but can also be areas such as parking lots, office buildings, campuses, industrial parks, etc.

[0046] To make the solution provided in this embodiment clearer, the following uses the controller of the wheeled robot as the electronic device for implementing this method to Figure 1 elaborate on each step in the method in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship can be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application. Continuing to refer to Figure 1 , the method includes:

[0047] S1, determining multiple look-ahead points to be explored from the exploration area.

[0048] In this embodiment, during the process of the controller controlling the wheeled robot to move in the house, the semantic map of the house is dynamically updated. This semantic map describes the objects in each local area in the two-dimensional plane corresponding to the house. For example, the areas occupied by the sink, wardrobe, wall, counter, chair, box, and table are marked in the two-dimensional plane.

[0049] To achieve the above purpose, the controller can generate a semantic segmentation mask with semantic confidence for the color image collected by the wheeled robot's camera through a pre-trained semantic segmentation model (for example, Semantic-Segmentation-Anything). For example, each pixel in the semantic segmentation mask of the wardrobe corresponds to the probability that the pixel belongs to the wardrobe. During the research process, it was found that the pixels in the image obtained by the camera imaging have different degrees of confidence according to their different positions in the camera's field of view, specifically manifested as the segmentation accuracy at positions closer to the optical axis of the field of view being higher than that at positions farther from the optical axis of the field of view. Therefore, in this embodiment, when constructing the semantic map, the semantic confidence segmented by the semantic segmentation model is optimized using the angle between each pixel and the optical axis. Among them, the confidence generated by the angle between each pixel and the optical axis is called the FOV confidence, and the expression is:

[0050]

[0051] In the formula, θ represents the angle between the pixel and the optical axis. For the pixels located at the edge position of the color image, the corresponding θ is larger. θ fov represents the horizontal field of view angle of the camera. The controller multiplies the FOV confidence of each pixel by the semantic confidence of that pixel to obtain the optimized semantic confidence.

[0052] Then, the controller constructs a 3D semantic point cloud using the optimized semantic confidence, depth image, and camera parameters, filters out the irrelevant points, and finally projects it onto the two-dimensional plane of the indoor area to obtain the semantic map. For the convenience of subsequent processing, the controller also uses the odometer data to convert these projected points into the global coordinate system. Based on this global coordinate system, during the exploration of the wheeled robot, the newly observed images are continuously projected onto the two-dimensional plane, and the semantic probabilities of the existing projected points in the two-dimensional plane are averaged to update the semantic map, so as to represent the probability of the existence of object categories at each pixel position in the entire environment.

[0053] Based on the semantic map obtained from the above embodiments, the controller generates the look-ahead points to be explored according to the boundary between the explored area and the unexplored area in the semantic map. Specifically, the controller can specify the midpoints of these boundaries as the look-ahead points. Therefore, as the wheeled robot continuously explores in the house, the set of look-ahead points will be continuously updated. Each movement will update the explored and unexplored areas, thus updating the set of look-ahead points; when the set of look-ahead points is empty, it means that the entire scene environment has been fully explored.

[0054] Based on the description of the look-ahead points in the above embodiments, the following continues to describe Figure 1 the step S2 in

[0055] S2. For each look-ahead point, obtain multiple exploration scores of the look-ahead point.

[0056] Among them, the multiple exploration scores include the room-level co-occurrence score, the object-level co-occurrence score, and the space exploration score. The room-level co-occurrence score represents the probability that all neighboring objects within a preset first range from the look-ahead point and the target object appear simultaneously in multiple candidate rooms. The object-level co-occurrence score represents the probability that the target object and all neighboring objects within the preset first range from the look-ahead point appear simultaneously. The space exploration score represents the value of evaluating the exploration of the exploration device from the current position to the look-ahead point only from the spatial factor. The target object is the object that the wheeled robot needs to find. For example, the instruction issued to the wheeled robot is "find the TV in the house", and the target object is the TV in the instruction.

[0057] It can be understood that for each look-ahead point, this embodiment evaluates it from three aspects. Next, these three aspects will be described in detail, that is, the acquisition methods of the room-level co-occurrence score, the object-level co-occurrence score, and the space exploration score will be introduced in detail. It should be understood that for the above two co-occurrence scores, this embodiment uses the semantic similarity between names for evaluation. Therefore, as an alternative implementation, step S2 may include:

[0058] S2-1, obtaining the first room co-occurrence score between the target object and multiple candidate rooms.

[0059] The above multiple candidate rooms can be "living room", "bedroom", "bathroom", "office", "kitchen". This embodiment uses a text embedding model to calculate the correlation between each neighborhood object and multiple candidate rooms to obtain the first room co-occurrence score.

[0060] Exemplarily, assume that the target object in this embodiment is "TV set". Then the first room co-occurrence score between the TV set and "living room", "bedroom", "bathroom", "office", "kitchen" can be calculated through the following expression.

[0061] R t =[S o-r (t,r1),S o-r (t,r2),…,S o-r (t,r m )]

[0062] In the formula, m represents the number of multiple candidate rooms. S o-r (t,r k ) represents calculating the semantic similarity between the TV set and the kth candidate room, and the expression is:

[0063]

[0064] It can be understood that the controller converts "living room", "bedroom", "bathroom", "office", "kitchen" into text embeddings and then calculates the semantic similarity with the text embedding of "TV set" respectively to obtain five semantic similarities. Then, the five semantic similarities are summed to obtain the first room co-occurrence score between the "TV set" and "living room", "bedroom", "bathroom", "office", "kitchen".

[0065] Based on the description of the first room co-occurrence score in the above embodiment, step S2 further includes:

[0066] S2-2, obtaining the second room co-occurrence score between each neighborhood object and multiple candidate rooms.

[0067] S2-3. Aggregate the second room co-occurrence scores of all neighborhood objects to obtain the aggregated room score of all neighborhood objects.

[0068] In this embodiment, each look-ahead point is represented as f i , and taking the position where the look-ahead point is located as the center of the circle, determine a circular area with an empirical value r as the radius as the neighborhood of the look-ahead point. The controller of the wheeled robot determines all neighborhood objects located in the circular area from the semantic map according to the position of the circular area, and represents them with the set O fk . Each neighborhood object is represented as o j . And represent the above 5 candidate rooms with the set R, and each candidate room is represented as r k , that is, r k ∈R. Then the second room co-occurrence score between the target object and multiple candidate rooms can be calculated through the following expression:

[0069]

[0070]

[0071] It can be understood that for each neighborhood object, the controller of the wheeled robot calculates the semantic similarity between the neighborhood object and each candidate room, so as to obtain the second room co-occurrence score of the neighborhood object.

[0072] Finally, the controller of the wheeled robot calculates the mean value of these second room co-occurrence scores as the aggregated room score. Of course, the controller of the wheeled robot can also perform weighted summation on these second room co-occurrence scores according to preset weights to obtain the aggregated room score.

[0073] S2-4. Determine the similarity between the first room co-occurrence score and the aggregated room score as the room-level co-occurrence score of the look-ahead point.

[0074] In this embodiment, the cosine similarity is used to measure the similarity between the first room co-occurrence score and the aggregated room score, and the expression is as follows:

[0075]

[0076] In this way, the room-level co-occurrence score can effectively identify the look-ahead points located in the room type most relevant to the target object, and can guide the wheeled robot towards the area with higher semantic relevance to the target object, thereby improving the exploration efficiency.

[0077] Based on the description of the room-level co-occurrence score in the above embodiments, this embodiment also captures the fine-grained spatial relationship between the neighborhood objects near the look-ahead point and the target object through the object-level co-occurrence score. Specifically, for each look-ahead point, the set of objects around it is analyzed to evaluate the direct spatial co-occurrence score with the target object. For this purpose, the same text embedding method as in the room-level calculation is used to capture the functional spatial context. Therefore, step S2 also includes:

[0078] S2-5, obtaining the object co-occurrence score between the target object and each neighborhood object.

[0079] S2-6, aggregating the object co-occurrence scores of all neighborhood objects to obtain the object-level co-occurrence score of the look-ahead point.

[0080] For the object co-occurrence score between the target object and each neighborhood object, it can be calculated through the following expression:

[0081] S o (o i ,t) = cos(E(o i ), E(t))

[0082] Then, these object co-occurrence scores are aggregated through the following expression to obtain the object-level co-occurrence score:

[0083]

[0084] Based on the description of the object-level co-occurrence score in the above embodiments, it is also found in the research process that in the initial exploration stage where limited observations are available, it is difficult to explore target navigation simply relying on semantic reasoning. Therefore, this embodiment introduces a spatial exploration score driven by exploration based on geometric cues, mainly based on density gain and distance penalty. Therefore, step S2 also includes:

[0085] S2-7, determining the number of neighborhood look-ahead points according to the position of the look-ahead point.

[0086] Among them, the neighborhood look-ahead point refers to other look-ahead points within a preset second range from the look-ahead point.

[0087] S2-8, determining the distance between the look-ahead point and the current position of the exploration device.

[0088] S2-9, normalizing the number and the distance, and aggregating the normalized number and the normalized distance to obtain the spatial exploration score.

[0089] The above three steps will be described in detail below in conjunction with specific expressions. In this embodiment, the number of neighborhood look-ahead points is used as the density gain index, which can provide look-ahead points with greater potential for new area discovery. For each look-ahead point, the controller normalizes the number of neighborhood look-ahead points within a specified radius as the density gain index. The normalization method is as follows:

[0090]

[0091] In the formula, D(f i ) represents calculating the number of neighborhood look-ahead points within the preset second range of the look-ahead point f i . represents the minimum number of neighborhood look-aheads among all look-ahead points, represents the maximum number of neighborhood look-ahead points among all look-ahead points.

[0092] As a supplement to the density gain index, this embodiment also introduces a distance penalty to suppress an overly long exploration path length. For each look-ahead point, the controller calculates the Euclidean distance from the current position of the wheeled robot to this look-ahead point and normalizes it to obtain the distance penalty. Among them, the normalization method is as follows:

[0093]

[0094] In the formula, dist(f i , a) represents calculating the distance from the current position a of the wheeled robot to the look-ahead point f i , represents the minimum distance, represents the maximum distance.

[0095] Considering that a higher value represents a larger distance, which will guide the wheeled robot to preferentially explore nearby look-ahead points, because if a target object is detected, this will minimize the path length. Therefore, by combining these exploration-driven valuation metrics, the two are aggregated through the following expression to obtain the spatial exploration score:

[0096] S e (f i ) = w·G density (f i ) + (1 - P distance (f i ))·w

[0097] In the formula, w represents the weight factor of the two components. Therefore, this aggregation method balances the two density gain indices and the distance penalty. Among them, the density gain index encourages the exploration of clusters of look-ahead points with greater potential for new discoveries, while the distance supplement term 1 - P distance (f i)First, the wheeled robot is made to prioritize nearby look-ahead points to minimize the path length.

[0098] In the above embodiments, the multiple exploration scores for each look-ahead point were described. Next, the description continues for Figure 1 step S3 in

[0099] S3. Obtain the comprehensive score of the look-ahead point according to the weights of the multiple exploration scores respectively.

[0100] It should be understood that the multiple exploration scores of each look-ahead point correspond one-to-one with multiple score types. For example, the room-level co-occurrence score of each look-ahead point corresponds to one score type, the object-level co-occurrence score corresponds to one score type, and the space exploration score corresponds to one score type, meaning that each look-ahead point has a total of 3 score types. Instead of presetting fixed weights for each score type to weight the multiple exploration scores of each look-ahead point, this embodiment provides an adaptive weight mechanism for dynamically evaluating the weights of multiple score types to identify the best look-ahead point that is most promising for efficient navigation. Therefore, before step S3, the method further includes:

[0101] S2.1. Normalize the multiple exploration scores of each look-ahead point to obtain the multiple normalized scores of each look-ahead.

[0102] As an alternative implementation, for each score type, the controller sorts the multiple look-ahead points according to the magnitude of the normalized score corresponding to the score type to obtain the sequence position of each look-ahead point corresponding to the score type; respectively according to the sequence position of each look-ahead point corresponding to the score type, obtain the normalized score of each look-ahead point corresponding to the score type. The corresponding expression is:

[0103]

[0104] In the formula, S r (f i ), S o (f i ), S e (f i ) represent the room-level co-occurrence score, object-level co-occurrence score, and space exploration score of the look-ahead point f i in sequence, and rank(S r (f i )), rank(S o (f i )), rank(S e (f i )) represent the sequence positions of the look-ahead point f iThe sequence position when sorted by room-level co-occurrence score, the sequence position when sorted by object-level co-occurrence score, and the sequence position when sorted by spatial exploration score; F represents the total number of look-ahead points.

[0105] Exemplarily, assume there are a total of 10 look-ahead points. These 10 look-ahead points are sorted according to the room-level co-occurrence score to obtain the sequence position of each of the 10 look-ahead points. Then, the sequence position of each look-ahead point is substituted into the above expression, so as to normalize the room-level co-occurrence score of each look-ahead point. For the object-level co-occurrence scores and spatial exploration scores of these 10 look-ahead points respectively, the normalization method is the same, and this embodiment will not elaborate further.

[0106] In this way, the above conversion method retains the relative order and at the same time amplifies the numerical differences between the look-ahead points, which is of great significance when combining scores of different dimensions with the potential compressed value range.

[0107] S2.2, for each type of score, calculate the variance between the normalized scores corresponding to multiple look-ahead points and the score type to obtain the score variance of the score type.

[0108] Exemplarily, continuing with the above 10 look-ahead points as an example, based on the normalized room-level co-occurrence scores of these 10 look-ahead points, a score variance can be calculated, which is called the room-level variance. Similarly, based on the normalized object-level co-occurrence scores of 10 look-ahead points, a second score variance can be calculated, which is called the object-level variance; based on the normalized spatial exploration scores of 10 look-ahead points, a third score variance can be calculated, which is called the space-level variance.

[0109] S2.3, based on the score variances of multiple score types, obtain the weights of the exploration scores corresponding to each score type.

[0110] In this regard, in this embodiment, the controller of the wheeled robot sums up the score variances of multiple score types to obtain the sum of variances; the ratio between the score variance of each score type and the sum of variances is respectively determined as the weight of the exploration score corresponding to each score type. Based on the above expression of the normalized score, the expression of the weight of the exploration score corresponding to each score type is:

[0111]

[0112] In the formula, are the room-level variance, object-level variance, and space-level variance in sequence, w r , w o , w eThe weights of the room-level co-occurrence score, the object-level co-occurrence score, and the spatial exploration score of each prospective point are obtained in turn. Based on the weights obtained by the above expressions, the normalized room-level co-occurrence score, object-level co-occurrence score, and spatial exploration score of each prospective point can be weighted and fused by the following expression:

[0113]

[0114] In the formula, S(f i ) represents the look-ahead point f i The comprehensive score of .

[0115] In the above embodiment, the comprehensive score is described. Figure 1 Step S4 in the description is as follows:

[0116] S4, determining the forward-looking point with the highest comprehensive score as the best forward-looking point for the exploration device to continue exploring.

[0117] In this way, after the robot reaches each look-ahead point, it determines the best look-ahead point through the above implementation method and continues to explore until the target object is found. In summary, in the above implementation method, the local environment of the look-ahead point and the spatial co-occurrence of the target object can be accurately evaluated at the room level and the object level at the same time, so as to select the look-ahead point that is most relevant to the target object for exploration. In addition, the look-ahead points are scored based on the exploration, and the semantic reasoning results and exploration efficiency of each look-ahead point are balanced through an adaptive weighting mechanism, so as to select the best look-ahead point for navigation of the target object. Moreover, the method is zero-sample and open vocabulary during implementation, which means that the location of any specified object can be detected without data collection and training.

[0118] In order to have a more comprehensive and clear understanding of the above embodiments, Figure 2 The whole embodiment is described more intuitively. Figure 2 As shown in the figure, after the wheeled robot receives the instruction to find a TV, the controller performs semantic segmentation on the taken color image to obtain the semantic segmentation mask of the color image. Then, the controller combines the depth map to project the semantic segmentation mask onto the two-dimensional plane corresponding to the indoor area to obtain a semantic map, and determines multiple forward-looking points from the semantic map. Next, the controller performs hierarchical semantic reasoning on each forward-looking point to obtain a room-level co-occurrence score and an object-level co-occurrence score; and performs exploration-driven numerical evaluation on the forward-looking point to obtain a spatial exploration score. Finally, the controller adaptively weights the above three scores for each forward-looking point, and selects the one with the highest score as the best frontier point based on the weighted comprehensive score.

[0119] Based on the same inventive concept as the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in this embodiment, this embodiment also provides an autonomous navigation intelligent device based on hierarchical semantic reasoning, which includes at least one software function module that can be stored in a memory or solidified in an electronic device in the form of software. The processor in the electronic device is used to execute the executable module stored in the memory. For example, the software function modules and computer programs included in the device. Please refer to Figure 3 , functionally speaking, the device may include:

[0120] A forward-looking point determination module 11 is used to determine a plurality of forward-looking points to be explored from the exploration area;

[0121] The forward-looking point evaluation module 12 is used to obtain multiple exploration scores of the forward-looking point for each forward-looking point, wherein the multiple exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighboring objects within a preset first range from the forward-looking point and the target object appear in multiple candidate rooms at the same time. The object-level co-occurrence score represents the probability that the target object and all neighboring objects appear at the same time. The space exploration score represents the value of evaluating the exploration device from the current position to the forward-looking point only from the spatial factor.

[0122] The prospective point evaluation module 12 is further used to obtain a comprehensive score of the prospective point according to the respective weights of the plurality of exploration scores;

[0123] The forward-looking point screening module 13 is used to determine the forward-looking point with the highest comprehensive score as the best forward-looking point for the exploration device to continue exploring.

[0124] In this embodiment, the forward-looking point determination module 11 is used to implement Figure 1 In step S1, the forward point evaluation module 12 is used to implement Figure 1 S2, S3, the forward point screening module 13 is used to achieve Figure 1 Therefore, for the detailed description of each of the above modules, please refer to the specific implementation of the corresponding steps, which will not be described in detail in this embodiment.

[0125] Since the invention concept is the same as that of the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in this embodiment, the autonomous navigation intelligent device based on hierarchical semantic reasoning can also implement other steps or sub-steps of the method through the above modules.

[0126] Optionally, the prospective point evaluation module 12 is further specifically used for:

[0127] Obtaining the first room co-occurrence scores between the target object and multiple candidate rooms;

[0128] Obtain the second room co-occurrence scores between each neighborhood object and multiple candidate rooms;

[0129] Aggregate the second room co-occurrence scores of all neighborhood objects to obtain the aggregated room scores of all neighborhood objects;

[0130] Determine the similarity between the first room co-occurrence score and the aggregated room score as the room-level co-occurrence score of the look-ahead point.

[0131] Optionally, the look-ahead point evaluation module 12 is further specifically configured to:

[0132] Obtain the object co-occurrence scores between the target object and each neighborhood object;

[0133] Aggregate the object co-occurrence scores of all neighborhood objects to obtain the object-level co-occurrence score of the look-ahead point.

[0134] Optionally, the look-ahead point evaluation module 12 is further specifically configured to:

[0135] Determine the number of neighborhood look-ahead points according to the position of the look-ahead point, where the neighborhood look-ahead point represents other look-ahead points within a preset second range from the look-ahead point;

[0136] Determine the distance between the look-ahead point and the current position of the exploration device;

[0137] Normalize the number and the distance, and aggregate the normalized number and the normalized distance to obtain the spatial exploration score.

[0138] Optionally, multiple exploration scores of each look-ahead point correspond to multiple score types one by one. Before obtaining the comprehensive score of the look-ahead point according to the respective weights of the multiple exploration scores, the look-ahead point evaluation module 12 is further configured to:

[0139] Normalize the multiple exploration scores of each look-ahead point to obtain multiple normalized scores for each look-ahead;

[0140] For each score type, calculate the variance between the multiple normalized scores corresponding to the score type of the look-ahead to obtain the score variance of the score type;

[0141] Obtain the weights of the exploration scores corresponding to each score type according to the score variances of the multiple score types.

[0142] Optionally, the look-ahead point evaluation module 12 is further specifically configured to:

[0143] Sum the score variances of the multiple score types to obtain the sum of variances;

[0144] Respectively determine the ratio between the score variance of each score type and the sum of variances as the weight of the exploration score corresponding to each score type.

[0145] Optionally, the look-ahead point evaluation module 12 is further specifically configured to:

[0146] For each score type, sort the multiple look-ahead points according to the magnitude of the normalized score corresponding to the score type, and obtain the sequence position corresponding to each look-ahead point and the score type;

[0147] Respectively, according to the sequence position corresponding to each look-ahead point and the score type, obtain the normalized score corresponding to each look-ahead point and the score type.

[0148] In addition, in each embodiment of the present application, each functional module may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0149] It should also be understood that if the above implementation manner is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0150] Therefore, this embodiment also provides a storage medium, and this storage medium is a computer-readable storage medium. This storage medium stores a computer program, and when the computer program is executed by a processor, it implements the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in this embodiment. Among them, the storage medium may be various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0151] An electronic device for implementing the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in this embodiment. As Figure 4 shown, the electronic device may include a processor 22 and a memory 120. And, the memory 21 stores a computer program, and the processor realizes the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in this embodiment by reading and executing the computer program corresponding to the above implementation manner in the memory 21.

[0152] Continue to refer to Figure 4, the electronic device further includes a communication unit 23. Each of the memory 21, the processor 22, and the communication unit 23 is directly or indirectly electrically connected to each other through a system bus 24 to achieve data transmission or interaction.

[0153] Among them, the memory 21 can be an information recording device based on any electronic, magnetic, optical or other physical principles, and is used to record execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.

[0154] In some embodiments, the volatile memory can be a random access memory (RAM); in some embodiments, the non-volatile memory can be a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, etc.; in some embodiments, the storage drive can be a disk drive, a solid state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.

[0155] The communication unit 23 is used to send and receive data through a network. In some embodiments, the network may include a wired network, a wireless network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, etc., or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the service request processing system may be connected to the network through the access point to exchange data and / or information.

[0156] The processor 22 may be an integrated circuit chip with signal processing capabilities, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the above-mentioned processor may include a central processing unit (CPU), an application specific integrated circuit (ASIC), an application specific instruction-set processor (ASIP), a graphics processing unit (GPU), a physics processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computing (RISC), or a microprocessor, etc., or any combination thereof.

[0157] It can be understood that Figure 4The structure shown is only illustrative. The electronic device 100 may also have more or fewer components than Figure 4 shown, or have a different configuration from Figure 4 shown. Figure 4 Each component shown may be implemented by hardware, software, or a combination thereof.

[0158] It should be understood that the devices and methods disclosed in the above embodiments may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0159] As described above, these are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An autonomous navigation intelligent method based on hierarchical semantic reasoning, characterized in that, The method includes: Determining a plurality of forward-looking points to be explored from the exploration area; For each of the forward-looking points, obtaining a plurality of exploration scores of the forward-looking point, where the plurality of exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighborhood objects within a preset first range from the forward-looking point and the target object appear in multiple candidate rooms simultaneously. The object-level co-occurrence score represents the probability that the target object and all neighborhood objects appear simultaneously. The space exploration score represents the value of evaluating the exploration device's movement from the current position to the forward-looking point only from spatial factors; Obtaining a comprehensive score of the forward-looking point according to the weights of the plurality of exploration scores; Determining the forward-looking point with the highest comprehensive score as the best forward-looking point for the exploration device to continue exploring.

2. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that Obtaining the room-level co-occurrence score of the forward-looking point includes: Obtaining a first room co-occurrence score between the target object and the multiple candidate rooms; Obtaining a second room co-occurrence score between each neighborhood object and the multiple candidate rooms; Aggregating the second room co-occurrence scores of all neighborhood objects to obtain an aggregated room score of all neighborhood objects; Determining the similarity between the first room co-occurrence score and the aggregated room score as the room-level co-occurrence score of the forward-looking point.

3. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that, Obtaining the object-level co-occurrence score of the forward-looking point includes: Obtaining an object co-occurrence score between the target object and each neighborhood object; Aggregating the object co-occurrence scores of all neighborhood objects to obtain the object-level co-occurrence score of the forward-looking point.

4. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that Obtaining the space exploration score of the forward-looking point includes: Determining the number of neighborhood forward-looking points according to the position of the forward-looking point, where the neighborhood forward-looking points refer to other forward-looking points within a preset second range from the forward-looking point; Determining the distance between the forward-looking point and the current position of the exploration device; Normalizing the number and the distance, and aggregating the normalized number and the normalized distance to obtain the space exploration score.

5. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that, The multiple exploration scores of each forward-looking point correspond to multiple score types one by one. Before obtaining the comprehensive score of the forward-looking point according to the weights of the multiple exploration scores, the method further includes: Normalizing the multiple exploration scores of each forward-looking point to obtain multiple normalized scores of each forward-looking point; For each score type, calculating the variance between the multiple normalized scores corresponding to the score type of the multiple forward-looking points to obtain the score variance of the score type; Obtaining the weights of the exploration scores corresponding to each score type according to the score variances of the multiple score types.

6. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 5, wherein Obtaining the weights of the exploration scores corresponding to each score type according to the score variances of the multiple score types includes: Summing the score variances of the multiple score types to obtain the sum of variances; Respectively determining the ratio between the score variance of each score type and the sum of variances as the weight of the exploration score corresponding to each score type.

7. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 5, characterized in that Normalize the multiple exploration scores of each of the said look-ahead points to obtain multiple normalized scores for each of the said look-aheads, including: For each type of score, sort the multiple look-ahead points according to the magnitudes of the normalized scores corresponding to the score type to obtain the sequence position corresponding to each of the said look-ahead points and the score type; Respectively, according to the sequence position corresponding to each of the said look-ahead points and the score type, obtain the normalized score corresponding to each of the said look-ahead points and the score type.

8. An autonomous navigation intelligent device based on hierarchical semantic reasoning, characterized in that, The said device includes: A look-ahead point determination module, configured to determine multiple look-ahead points to be explored from the exploration area; A look-ahead point evaluation module, configured to, for each of the said look-ahead points, obtain multiple exploration scores of the look-ahead point, wherein the multiple exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score, the room-level co-occurrence score representing the probability that all neighborhood objects within a preset first range from the look-ahead point and the target object appear in multiple candidate rooms simultaneously, the object-level co-occurrence score representing the probability that the target object and all the neighborhood objects appear simultaneously, and the space exploration score representing the value of evaluating the exploration device's movement from the current position to the look-ahead point for exploration only from spatial factors; The said look-ahead point evaluation module is further configured to obtain a comprehensive score of the look-ahead point according to the weights of the multiple exploration scores respectively; A look-ahead point screening module, configured to determine the look-ahead point with the highest comprehensive score as the best look-ahead point for the exploration device to continue to explore.

9. A storage medium, characterized in that, The said storage medium stores a computer program, which, when executed by a processor, implements the autonomous navigation intelligent method based on hierarchical semantic reasoning according to any one of claims 1-7.

10. An electronic device, characterized in that, The said electronic device includes a processor and a memory, the memory stores a computer program, which, when executed by a processor, implements the autonomous navigation intelligent method based on hierarchical semantic reasoning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional point cloud semantic segmentation method for optimizing boundary

    CN115409989A

  • Robot target navigation method using multi-clue semantic matching and related device

    CN119268696A

  • Passable region exploration method and apparatus, storage medium, and electronic apparatus

    WO2022110853A1