Autonomous navigation intelligent method based on hierarchical semantic reasoning and related device

By employing a hierarchical semantic reasoning-based autonomous navigation method, which combines room-level and object-level co-occurrence scores and uses an adaptive weighting mechanism to select the optimal look-ahead point, the efficiency and adaptability issues of existing navigation methods in sparse target scenarios are resolved, achieving efficient navigation and item information updates.

CN120252686BActive Publication Date: 2026-02-13BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510419097.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2026-02-13
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Existing zero-sample object navigation methods have insufficient success rates in sparse target scenarios and are susceptible to environmental noise interference, path oscillation, and difficulty in adapting to dynamic environments, resulting in insufficient navigation efficiency and adaptability.

Method used

An autonomous navigation method based on hierarchical semantic reasoning is adopted. By acquiring multiple exploration scores, including room-level co-occurrence scores, object-level co-occurrence scores, and spatial exploration scores, an adaptive weighting mechanism is used to balance semantic reasoning results and exploration efficiency to determine the optimal look-ahead point.

Benefits of technology

It improves the navigation efficiency of exploration equipment and its adaptability to different environments, enables the detection and updating of zero-sample item information, and reduces path redundancy and repeated exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120252686B_ABST
    Figure CN120252686B_ABST
Patent Text Reader

Abstract

The application provides an autonomous navigation intelligent method based on hierarchical semantic reasoning and related devices. A plurality of prospective points to be explored are determined from an exploration area; for each prospective point, a plurality of exploration scores of the prospective point are obtained; wherein the plurality of exploration scores include a room-level co-occurrence score, an object-level co-occurrence score and a space exploration score; according to respective weights of the plurality of exploration scores, a comprehensive score of the prospective point is obtained; and the prospective point with the highest comprehensive score is determined as the best prospective point for the exploration device to continue to explore. In this way, the method accurately evaluates the spatial co-occurrence of the local environment of the prospective point and the target object from two levels of room level and object level, introduces a space exploration score, balances the semantic reasoning result and the exploration efficiency through an adaptive weighting mechanism, and realizes the detection and update of the object information without data collection and training, thereby improving the navigation efficiency of the exploration device and the adaptability to different environments.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of object navigation, in particular, to an autonomous navigation intelligent method based on hierarchical semantic reasoning and related device. BACKGROUND

[0002] The mainstream methods in the current zero-shot object navigation field can be divided into three technical routes: geometric exploration, vision-language model-based methods, and large language model-based methods. Among them, the geometric exploration-based methods include frontier-based exploration (FBE) methods and Voronoi diagram-based navigation exploration methods, which explore the environment by constructing a boundary area based on the occupancy map. Although the pure geometric strategy can guarantee the basic exploration efficiency, it completely ignores the semantic information, resulting in significant blindness in target search, and the success rate is less than 40% in sparse target scenes. The vision-language model-based methods include the wheel-based CLIP (CLIP on Wheels, CoW) method and the vision-language frontier mapping (VLFM) method. These methods realize semantic mapping through open vocabulary detection or language-driven value mapping, but their single-layer semantic reasoning mechanism only relies on target category label matching and does not model the spatial co-occurrence relationship between objects, which is easily disturbed by environmental noise in complex home scenes, resulting in path oscillation, specifically repeating the exploration of places that have been explored. The large language model-based methods include the L3MVN and OpenFMNav methods. These methods attempt to use a pre-trained common sense knowledge base to infer the target position, but are limited by the lagging static knowledge update and cloud processing delay, making it difficult to adapt to real-time needs in dynamic environments.

[0003] Therefore, how to improve the navigation efficiency of the exploration device and its adaptability to different environments has become a problem that needs to be solved. SUMMARY

[0004] In order to overcome all the deficiencies in the prior art, the present application provides an autonomous navigation intelligent method based on hierarchical semantic reasoning and related device, specifically including:

[0005] In the first aspect, the present application provides an autonomous navigation intelligent method based on hierarchical semantic reasoning, which comprises:

[0006] determining a plurality of forward-looking points to be explored from the exploration area;

[0007] For each of the prospective points, a plurality of exploration scores of the prospective point are obtained, wherein the plurality of exploration scores comprise a room-level co-occurrence score, an object-level co-occurrence score, and a spatial exploration score, the room-level co-occurrence score representing a probability that all neighborhood objects within a preset first range from the prospective point and a target object simultaneously appear in a plurality of candidate rooms, the object-level co-occurrence score representing a probability that the target object and the all neighborhood objects simultaneously appear, and the spatial exploration score representing a value of exploration of the exploration device from a current position to the prospective point only from a spatial factor;

[0008] A comprehensive score of the prospective point is obtained according to respective weights of the plurality of exploration scores.

[0009] The prospective point with the highest comprehensive score is determined as a best prospective point to which the exploration device continues to explore.

[0010] In a second aspect, the present application provides an autonomous navigation intelligent device based on hierarchical semantic reasoning, the device comprising:

[0011] A prospective point determination module is configured to determine a plurality of prospective points to be explored from an exploration area.

[0012] A prospective point evaluation module is configured to obtain, for each of the prospective points, a plurality of exploration scores of the prospective point, wherein the plurality of exploration scores comprise a room-level co-occurrence score, an object-level co-occurrence score, and a spatial exploration score, the room-level co-occurrence score representing a probability that all neighborhood objects within a preset first range from the prospective point and a target object simultaneously appear in a plurality of candidate rooms, the object-level co-occurrence score representing a probability that the target object and the all neighborhood objects simultaneously appear, and the spatial exploration score representing a value of exploration of the exploration device from a current position to the prospective point only from a spatial factor.

[0013] The prospective point evaluation module is further configured to obtain a comprehensive score of the prospective point according to respective weights of the plurality of exploration scores.

[0014] A prospective point screening module is configured to determine the prospective point with the highest comprehensive score as a best prospective point to which the exploration device continues to explore.

[0015] In a third aspect, the present application provides a storage medium storing a computer program, the computer program realizing the autonomous navigation intelligent method based on hierarchical semantic reasoning when executed by a processor.

[0016] In a fourth aspect, the present application provides an electronic device comprising a processor and a memory, the memory storing a computer program, the computer program realizing the autonomous navigation intelligent method based on hierarchical semantic reasoning when executed by the processor.

[0017] Compared with the prior art, the present application has the following beneficial effects:

[0018] The present application provides an autonomous navigation intelligent method based on hierarchical semantic reasoning and related devices. A plurality of prospective points to be explored are determined from an exploration area; for each prospective point, a plurality of exploration scores of the prospective point are obtained; wherein the plurality of exploration scores include a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score; according to respective weights of the plurality of exploration scores, a comprehensive score of the prospective point is obtained; and the prospective point with the highest comprehensive score is determined as the best prospective point for the exploration device to continue to explore. In this way, the method accurately evaluates the spatial co-occurrence of the local environment of the prospective point and the target object from two levels of room level and object level, introduces a space exploration score, balances the semantic reasoning result and the exploration efficiency through an adaptive weighting mechanism, and realizes the detection and updating of the object information without data collection and training, thereby improving the navigation efficiency of the exploration device and the adaptability to different environments. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0020] Figure 1 A flowchart of the autonomous navigation intelligent method based on hierarchical semantic reasoning provided by the embodiments of the present application is shown in the figure.

[0021] Figure 2 A complete principle diagram of the autonomous navigation intelligent method based on hierarchical semantic reasoning provided by the embodiments of the present application is shown in the figure.

[0022] Figure 3 A structure diagram of the autonomous navigation intelligent device based on hierarchical semantic reasoning provided by the embodiments of the present application is shown in the figure.

[0023] Figure 4 A structure diagram of the electronic device provided by the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the embodiments of the present application (hereinafter referred to as the present embodiments) clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some but not all of the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0025] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present application.

[0026] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0027] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship when the product of the present application is usually placed, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" and the like are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0028] In addition, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0029] In addition, the terms "horizontal", "vertical", "overhang" and the like do not mean that the component must be absolutely horizontal or overhanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.

[0030] In the description of the present application, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set", "install", "connect", "connect" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be connected inside two elements. For those skilled in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.

[0031] Based on the above statement, as introduced in the background art, the method based on geometric exploration can realize basic environment exploration through boundary area segmentation, but it completely depends on the geometric features of the occupancy map and ignores the relevance of target objects and scene semantics. For example, when the target object is blocked or distributed in a sparse area, the average search success rate of such methods decreases due to the lack of semantic guidance mechanism, and the path redundancy is high.

[0032] And the method based on visual language model can construct semantic map, but its single-layer semantic reasoning mechanism only generates navigation path through target class label matching, without modeling the spatial co-occurrence relationship between objects and the room-level semantic distribution characteristics. This leads to the fact that when searching for "microwave oven" in the kitchen scene, the existing method may be misled to the location where there is a semantically similar object, for example, the location of "oven". But in some cases, "microwave oven" and "oven" are not placed together, and this method does not remember the explored location, and when "oven" is seen again in the field of view, it will go to the location of "oven" again to explore, resulting in the phenomenon of path oscillation.

[0033] The method based on large language model will be limited by the lag of static common sense knowledge base update and the delay of cloud processing. For example, OpenFMNav needs to upload perception data to the cloud large language model for processing, which has a decision delay.

[0034] Therefore, the existing methods generally have the problem of semantic reasoning flattening, neither establishing the hierarchical association of room-level semantic distribution and target objects, nor lacking the cooperative optimization mechanism of semantic clues and geometric exploration strategy, resulting in the difficulty of navigation efficiency and adaptability to meet the actual deployment requirements.

[0035] Based on the discovery of the above technical problems, after creative labor, the following technical solutions are proposed to solve or improve the above problems. It should be noted that the defects of the above prior art solutions are the result of careful study and practice, therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application to solve the above problems should be considered as contributions to the present application in the process of invention and creation, and should not be understood as technical content known to those skilled in the art.

[0036] In view of the discovery of the above problems, the embodiment provides an autonomous navigation intelligent method based on hierarchical semantic reasoning. As shown in the figure, the method comprises: Figure 1

[0037] S1, determining a plurality of prospective points to be explored from an exploration area.

[0038] S2, for each prospective point, obtaining a plurality of exploration scores of the prospective point.

[0039] Among them, the plurality of exploration scores includes a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighboring objects within a preset first range from the prospective point and the target object appear in multiple candidate rooms at the same time. The object-level co-occurrence score represents the probability that the target object and all neighboring objects appear at the same time. The space exploration score represents the value of the exploration device exploring from the current position to the prospective point only from the spatial factor.

[0040] S3, obtaining a comprehensive score of the prospective point according to the respective weights of the plurality of exploration scores.

[0041] S4, determining the prospective point with the highest comprehensive score as the best prospective point for the exploration device to continue to explore.

[0042] In this way, the method accurately evaluates the spatial co-occurrence of the local environment of the prospective point and the target object from the room level and the object level, and also introduces the space exploration score. Through the adaptive weighting mechanism, the semantic reasoning result and the exploration efficiency are balanced. Moreover, the detection and update of the object information can be realized without data collection and training. Therefore, the navigation efficiency of the exploration device and the adaptability to different environments can be improved.

[0043] For the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in the embodiment, the electronic device implementing the method can be, but is not limited to, a host computer in communication connection with the exploration device and a controller integrated with the exploration device. Among them, the host computer can be, but is not limited to, a mobile terminal, a tablet computer, a laptop computer, a desktop computer, etc., as long as it can provide sufficient computing power for the navigation of the exploration device.

[0044] It should also be understood that the exploration device in the embodiment can be, but is not limited to, a robot (such as an autonomous mobile robot, a quadruped robot, a wheeled robot, a tracked robot), a drone (such as a quadcopter drone, a fixed-wing drone), etc.

[0045] ​For example, taking a wheeled robot as an example, when a wheeled robot is given the task of finding a "television set" inside a house, after entering the house, it will use a depth camera to scan the environment and combine it with odometry data to construct a two-dimensional or three-dimensional map. Then, the wheeled robot will construct a navigation route based on its built-in search strategy to explore the house until it finds the "television set" located in a certain position inside the house. Of course, the exploration area in this application is not limited to indoor scenes, but can also be areas such as parking lots, office buildings, campuses, and industrial parks.

[0046] To make the solution provided in this embodiment clearer, the controller of a wheeled robot is used as the electronic device for implementing this method. Figure 1 Each step in the method is described in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical contextual relationships may be reversed in order or performed simultaneously. Furthermore, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowchart, or remove one or more operations from the flowchart. See also... Figure 1 The method includes:

[0047] S1 identifies multiple forward points to be explored from the exploration area.

[0048] In this embodiment, the controller dynamically updates the semantic map of the room as the wheeled robot moves within it. This semantic map describes the objects in each local area of ​​a two-dimensional plane corresponding to the room. For example, it marks the areas occupied by a sink, wardrobe, wall, counter, chair, box, and table in the two-dimensional plane.

[0049] To achieve the above objectives, the controller can generate semantic segmentation masks with semantic confidence from color images captured by the wheeled robot's camera using a pre-trained semantic segmentation model (e.g., Semantic-Segmentation-Anything). For example, each pixel in the semantic segmentation mask of a wardrobe corresponds to a probability that the pixel belongs to the wardrobe. Research has shown that pixels in the camera-imported image have different levels of confidence depending on their position within the camera's field of view; specifically, the segmentation accuracy is higher for pixels closer to the optical axis than for pixels farther from the optical axis. Therefore, in this embodiment, when constructing the semantic map, the semantic confidence generated by the semantic segmentation model is optimized using the angle between each pixel and the optical axis. The confidence generated by the angle between each pixel and the optical axis is called the FOV confidence, and its expression is:

[0050]

[0051] where θ represents the angle between a pixel and the optical axis, and the larger the corresponding θ of a pixel located at the edge of a color image is. θ fov represents the horizontal field of view of the camera. The controller multiplies the FOV confidence of each pixel with the semantic confidence of that pixel to obtain the optimized semantic confidence.

[0052] Then, the controller constructs a 3D semantic point cloud using the optimized semantic confidence, the depth image and the camera parameters, and filters out irrelevant points in the 3D semantic point cloud. Finally, the controller projects the 3D semantic point cloud onto a two-dimensional plane of the indoor area to obtain a semantic map. In order to facilitate subsequent processing, the controller also converts the projection points into a global coordinate system using the odometer data. Based on the global coordinate system, the controller controls the wheeled robot during exploration to continue projecting the newly observed image onto the two-dimensional plane and updating the semantic map by averaging the semantic probabilities of the projection points in the two-dimensional plane, so as to represent the probability of the existence of an object class at each pixel position in the entire environment.

[0053] Based on the semantic map obtained by the above embodiment, the controller generates a lookahead point to be explored according to the boundary between the explored area and the unexplored area in the semantic map. Specifically, the controller can specify the midpoint of the boundary as the lookahead point. Therefore, as the wheeled robot continues to explore in the indoor area, the lookahead point set will be updated. Each movement updates the explored and unexplored areas, thereby updating the lookahead point set; when the lookahead point set is empty, it indicates that the entire scene environment has been completely explored.

[0054] Based on the above description of the lookahead point, the following will continue to describe step S2 in the above embodiment: Figure 1

[0055] S2, for each lookahead point, obtaining a plurality of exploration scores of the lookahead point.

[0056] The plurality of exploration scores include a room-level co-occurrence score, an object-level co-occurrence score and a spatial exploration score. The room-level co-occurrence score represents the probability that all neighboring objects within a preset first range of the lookahead point and a target object appear in multiple candidate rooms at the same time. The object-level co-occurrence score represents the probability that the target object and all neighboring objects within a preset first range of the lookahead point appear at the same time. The spatial exploration score represents the value of the exploration device exploring from the current position to the lookahead point only from the spatial factor. The target object represents the object to be found by the wheeled robot. For example, the instruction issued to the wheeled robot is “find the TV in the house”, and the target object is the TV in the instruction.

[0057] ​It can be understood that for each prospective point, the embodiment evaluates it from three aspects. Next, the three aspects will be described in detail, i.e., the acquisition method of the room-level co-occurrence score, the object-level co-occurrence score, and the spatial exploration score. It should be understood that for the above two co-occurrence scores, the embodiment uses the semantic similarity between names for evaluation, therefore, as an optional implementation, step S2 can include:

[0058] S2-1, obtaining a first room co-occurrence score between the target object and a plurality of candidate rooms.

[0059] The plurality of candidate rooms can be "living room", "bedroom", "bathroom", "office", and "kitchen". The embodiment uses a text embedding model to calculate the correlation between each neighborhood object and the plurality of candidate rooms to obtain the first room co-occurrence score.

[0060] Exemplarily, assuming that the target object in the embodiment is "television". Then the first room co-occurrence scores between the television and "living room", "bedroom", "bathroom", "office", and "kitchen" can be calculated by the following expression.

[0061] R t =[S o-r (t,r1),S o-r (t,r2),…,S o-r (t,r m )]

[0062] In the formula, m represents the number of the plurality of candidate rooms. S o-r (t,r k ) represents calculating the semantic similarity between the television and the kth candidate room, and the expression is:

[0063]

[0064] It can be understood that the controller converts "living room", "bedroom", "bathroom", "office", and "kitchen" into text embedding respectively, and then calculates the semantic similarity with the text embedding of "television" to obtain five semantic similarities. Then, the five semantic similarities are summed to obtain the first room co-occurrence scores between "television" and "living room", "bedroom", "bathroom", "office", and "kitchen".

[0065] Based on the above embodiment for the first room co-occurrence score, step S2 further includes:

[0066] S2-2, obtaining a second room co-occurrence score between each neighborhood object and a plurality of candidate rooms.

[0067] S2-3, aggregate the second room co-occurrence scores of all the neighborhood objects to obtain an aggregated room score of all the neighborhood objects.

[0068] In this embodiment, each look-ahead point is represented as f i , and the position of the look-ahead point is taken as the center of a circle, a circular region with an empirical value r as the radius is determined as the neighborhood of the look-ahead point. The controller of the wheeled robot determines all the neighborhood objects located in the circular region according to the position of the circular region in the semantic map, and represents them as a set O fk , each neighborhood object is represented as o j . And the above 5 candidate rooms are represented as a set R, each candidate room is represented as r k , i.e. r k ∈R. The second room co-occurrence score between the target object and each candidate room can be calculated by the following expression:

[0069]

[0070]

[0071] It can be understood that, for each neighborhood object, the controller of the wheeled robot calculates the semantic similarity between the neighborhood object and each candidate room, thereby obtaining the second room co-occurrence score of the neighborhood object.

[0072] Finally, the controller of the wheeled robot calculates the mean of these second room co-occurrence scores as the aggregated room score. Of course, the controller of the wheeled robot can also weight-sum these second room co-occurrence scores according to the pre-set weight to obtain the aggregated room score.

[0073] S2-4, determine the similarity between the first room co-occurrence score and the aggregated room score as the room-level co-occurrence score of the look-ahead point.

[0074] In this embodiment, the similarity between the first room co-occurrence score and the aggregated room score is measured by the cosine similarity, and the expression is as follows:

[0075]

[0076] In this way, the look-ahead point located in the room type most relevant to the target object is effectively identified by the room-level co-occurrence score, which can guide the wheeled robot to the area with higher semantic relevance to the target object, thereby improving the exploration efficiency.

[0077] Based on the above embodiments' description of room-level co-occurrence scores, this embodiment also captures fine-grained spatial relationships between neighboring objects near a lookahead and the target object through object-level co-occurrence scores. Specifically, for each lookahead, the surrounding set of objects is analyzed to evaluate the direct spatial co-occurrence score with the target object. For this, the same text embedding method as in room-level computation is used to capture functional spatial context. Therefore, step S2 further includes:

[0078] S2-5, obtain the object co-occurrence score between the target object and each neighboring object.

[0079] S2-6 aggregates the object co-occurrence scores of all neighboring objects to obtain the object-level co-occurrence score of the lookahead.

[0080] The object co-occurrence score between the target object and each neighboring object can be calculated using the following expression:

[0081] S o (o i ,t)=cos(E(o i ),E(t))

[0082] Then, these object co-occurrence scores are aggregated using the following expression to obtain the object-level co-occurrence score:

[0083]

[0084] Based on the above embodiments' explanation of object-level co-occurrence scores, the research also found that in the initial exploration phase with limited observations available, relying solely on semantic reasoning for target navigation is quite difficult. Therefore, this embodiment introduces a spatial exploration score driven by geometric cues, primarily based on density gain and distance penalty. Therefore, step S2 further includes:

[0085] S2-7, Determine the number of neighboring lookaheads based on their location.

[0086] Among them, the neighboring lookahead points represent other lookahead points within a preset second range from the lookahead points.

[0087] S2-8, Determine the distance between the lookout point and the current position of the exploration device.

[0088] S2-9 normalizes the quantity and distance, and aggregates the normalized quantity and distance to obtain the space exploration score.

[0089] The three steps are described in detail below in connection with specific expressions. In this embodiment, the number of neighborhood look-ahead points is taken as the density gain indicator, which can provide more potential look-ahead points for new region discovery. For each look-ahead point, the controller normalizes the number of neighborhood look-ahead points within a specified radius as the density gain indicator. The normalization is as follows:

[0090]

[0091] where D(f i ) denotes the number of neighborhood look-ahead points within a specified radius of look-ahead point f i . denotes the number of neighborhood look-ahead points within a specified radius of look-ahead point f i , denotes the number of neighborhood look-ahead points within a specified radius of look-ahead point f i .

[0092] In addition to the density gain indicator, this embodiment also introduces a distance penalty to suppress excessively high exploration path length. For each look-ahead point, the controller calculates the Euclidean distance from the current position of the wheeled robot to the look-ahead point and normalizes it to obtain the distance penalty. The normalization is as follows:

[0093]

[0094] where dist(f i , a) denotes the distance from the current position a of the wheeled robot to look-ahead point f i , denotes the minimum distance, denotes the maximum distance.

[0095] Considering that a higher value indicates a greater distance, this will guide the wheeled robot to preferentially explore nearby look-ahead points, as this will minimize the path length if a target object is detected. Therefore, by combining these exploration-driven evaluation indicators, the two are aggregated by the following expression to obtain the spatial exploration score:

[0096] S e (f i ) = w · G density (f i ) + (1 - P distance (f i )) · w

[0097] where w denotes the weight factor of the two components. Therefore, this aggregation balances the two density gain indicators and the distance penalty. The density gain indicator encourages exploration of look-ahead point clusters with greater potential for new discoveries, while the distance penalty 1 - P distance (f i) to prioritize nearby lookahead points to minimize path length.

[0098] In the above embodiments, multiple exploration scores for each lookahead point are described, and the following continues to describe Figure 1 S3, obtaining a comprehensive score of the lookahead point according to the weight of each exploration score.

[0099] S3, obtaining a comprehensive score of the lookahead point according to the weight of each exploration score.

[0100] It should be understood that the multiple exploration scores of each lookahead point correspond to multiple score types one by one. For example, the room-level co-occurrence score of each lookahead point corresponds to one score type, the object-level co-occurrence score corresponds to one score type, and the space exploration score corresponds to one score type, which means that each lookahead point has 3 score types. Instead of assigning a preset fixed weight to each score type, the multiple exploration scores of each lookahead point are weighted. The present embodiment provides an adaptive weight mechanism for dynamically evaluating the weights of multiple score types to identify the best lookahead point that is most likely to achieve efficient navigation. Therefore, before step S3, the method further comprises:

[0101] S2.1, normalizing the multiple exploration scores of each lookahead point to obtain multiple normalized scores of each lookahead point.

[0102] As an optional implementation, for each score type, the controller sorts the multiple lookahead points according to the size of the normalized score corresponding to the score type to obtain a sequence position of each lookahead point corresponding to the score type; and obtains a normalized score of each lookahead point corresponding to the score type according to the sequence position of each lookahead point corresponding to the score type. The corresponding expression is:

[0103]

[0104] In the formula, S r (f i ), S o (f i ), and S e (f i ) represent the room-level co-occurrence score, the object-level co-occurrence score, and the space exploration score of the lookahead point f i , respectively, and rank(S r (f i )), rank(S o (f i )), and rank(S e (f i )) represent the sequence position of the lookahead point f ithe sequence position when ranked by the room-level co-occurrence score, the sequence position when ranked by the object-level co-occurrence score, the sequence position when ranked by the spatial exploration score; F represents the number of all the prospective points.

[0105] Exemplarily, assuming that there are 10 prospective points, the 10 prospective points are ranked by the room-level co-occurrence score to obtain the sequence position of each of the 10 prospective points; then, each prospective point sequence position is brought into the above expression, so as to normalize the room-level co-occurrence score of each prospective point. The normalization manner of the object-level co-occurrence score and the spatial exploration score of each of the 10 prospective points is the same, and the present embodiment will not be described in detail.

[0106] In this way, the above conversion manner preserves the relative order while amplifying the numerical difference between the prospective points, which is of great significance when combining scores of different dimensions with a potential compressed value range.

[0107] S2.2, for each score type, calculate the variance between the plurality of prospective points and the normalized score corresponding to the score type, to obtain a score variance of the score type.

[0108] Exemplarily, continuing with the above 10 prospective points as an example, according to the room-level co-occurrence scores of the 10 prospective points after normalization, a score variance can be calculated, which is referred to as a room-level variance. Similarly, according to the object-level co-occurrence scores of the 10 prospective points after normalization, a second score variance can be calculated, which is referred to as an object-level variance; according to the spatial exploration scores of the 10 prospective points after normalization, a third score variance can be calculated, which is referred to as a spatial-level variance.

[0109] S2.3, according to the score variances of the plurality of score types, obtain a weight of the exploration score corresponding to each score type.

[0110] In this regard, in the present embodiment, the controller of the wheeled robot sums the score variances of the plurality of score types to obtain a sum of variances; and determines the ratio between the score variance of each score type and the sum of variances as the weight of the exploration score corresponding to each score type. Based on the expression of the normalized score, the expression of the weight of the exploration score corresponding to each score type is:

[0111]

[0112] In the formula, varroom, varobj, and varspatial represent the room-level variance, the object-level variance, and the spatial-level variance, respectively, and wroom, wobj, and wspatial represent the weights of the room-level exploration score, the object-level exploration score, and the spatial exploration score, respectively. In the formula, varroom, varobj, and varspatial represent the room-level variance, the object-level variance, and the spatial-level variance, respectively, and wroom, wobj, and wspatial represent the weights of the room-level exploration score, the object-level exploration score, and the spatial exploration score, respectively. r o e ​​The weights of the room-level co-occurrence score, the object-level co-occurrence score, and the spatial exploration score of each prospective point are sequentially obtained. Based on the weights obtained by the above expression, the normalized room-level co-occurrence score, the object-level co-occurrence score, and the spatial exploration score of each prospective point can be fused by the following expression:

[0113]

[0114] In the formula, S(f i ) represents the comprehensive score of the prospective point f i .

[0115] The comprehensive score is described in the above embodiment. The step S4 in the above embodiment is described as follows: Figure 1

[0116] S4, the prospective point with the highest comprehensive score is determined as the optimal prospective point for the exploration device to continue to explore.

[0117] Thus, the robot determines the optimal prospective point after reaching each prospective point by the above embodiment, and continues to explore until the target object is found. In summary, in the above embodiment, the spatial co-occurrence of the local environment of the prospective point and the target object can be accurately evaluated from the room level and the object level, so that the prospective point most related to the target object is selected for exploration. In addition, the prospective point is scored based on exploration, and the semantic reasoning result and the exploration efficiency of each prospective point are balanced by the adaptive weighting mechanism, so that the optimal prospective point is selected for navigation of the target object. Moreover, the method is zero-sample and open-vocabulary in the implementation process, which also means that data collection and training are not required for detection of any specified object position.

[0118] To facilitate a more comprehensive and clear understanding of the above embodiment, the following describes the whole embodiment more intuitively. As shown in Figure 2 Figure 2 After receiving the instruction to find the TV, the controller performs semantic segmentation on the color image to obtain a semantic segmentation mask of the color image. Then, the controller projects the semantic segmentation mask into a two-dimensional plane corresponding to the indoor area to obtain a semantic map in combination with the depth map, and determines a plurality of prospective points from the semantic map. Next, the controller performs hierarchical semantic reasoning on each prospective point to obtain a room-level co-occurrence score and an object-level co-occurrence score; and performs exploration-driven numerical evaluation on the prospective point to obtain a spatial exploration score. Finally, the controller adaptively weights the three scores of each prospective point, and selects the prospective point with the highest score as the optimal prospective point according to the weighted comprehensive score.

[0119] ​​Based on the same inventive concept as the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in the embodiment, the embodiment further provides an autonomous navigation intelligent device based on hierarchical semantic reasoning, which comprises at least one software function module stored in the memory in the form of software or solidified in the electronic device. The processor in the electronic device is used to execute the executable module stored in the memory. For example, the software function module and the computer program included in the device. Please refer to Figure 3 Functionally, the device can comprise:

[0120] The prospective point determination module 11 is configured to determine a plurality of prospective points to be explored from the exploration area.

[0121] The prospective point evaluation module 12 is configured to obtain a plurality of exploration scores of the prospective point for each prospective point, wherein the plurality of exploration scores comprises a room-level co-occurrence score, an object-level co-occurrence score, and a space exploration score. The room-level co-occurrence score represents the probability that all neighboring objects within a preset first range from the prospective point and the target object appear in multiple candidate rooms at the same time. The object-level co-occurrence score represents the probability that the target object and all neighboring objects appear at the same time. The space exploration score represents the value of the exploration device exploring from the current position to the prospective point only from the spatial factor.

[0122] The prospective point evaluation module 12 is further configured to obtain a comprehensive score of the prospective point according to the respective weights of the plurality of exploration scores.

[0123] The prospective point screening module 13 is configured to determine the prospective point with the highest comprehensive score as the best prospective point for the exploration device to continue to explore.

[0124] In the embodiment, the prospective point determination module 11 is configured to implement step S1 in Figure 1 The prospective point evaluation module 12 is configured to implement S2 and S3 in Figure 1 The prospective point screening module 13 is configured to implement step S4 in Figure 1 Therefore, the detailed description of each module can be referred to the specific embodiments of the corresponding steps, and the embodiment will not be described again.

[0125] Since the autonomous navigation intelligent device based on hierarchical semantic reasoning has the same inventive concept as the autonomous navigation intelligent method based on hierarchical semantic reasoning provided in the embodiment, the device can also implement other steps or sub-steps of the method through the above modules.

[0126] Optionally, the prospective point evaluation module 12 is further configured to:

[0127] Obtain a first room co-occurrence score between the target object and the plurality of candidate rooms.

[0128] obtaining a second room co-occurrence score between each neighborhood object and each candidate room;

[0129] aggregating the second room co-occurrence scores of all neighborhood objects to obtain an aggregated room score of all neighborhood objects;

[0130] determining a similarity between the first room co-occurrence score and the aggregated room score as a room-level co-occurrence score of the prospective point.

[0131] Optionally, the prospective point evaluation module 12 is further configured to:

[0132] obtaining an object co-occurrence score between the target object and each neighborhood object;

[0133] aggregating the object co-occurrence scores of all neighborhood objects to obtain an object-level co-occurrence score of the prospective point.

[0134] Optionally, the prospective point evaluation module 12 is further configured to:

[0135] determining a number of neighborhood prospective points according to the location of the prospective point, wherein a neighborhood prospective point represents another prospective point within a preset second range from the prospective point;

[0136] determining a distance of the prospective point from the current location of the exploration device;

[0137] normalizing the number and the distance, and aggregating the normalized number and the normalized distance to obtain a spatial exploration score.

[0138] Optionally, the plurality of exploration scores of each prospective point correspond to a plurality of score types, and before obtaining a comprehensive score of the prospective point according to respective weights of the plurality of exploration scores, the prospective point evaluation module 12 is further configured to:

[0139] normalizing the plurality of exploration scores of each prospective point to obtain a plurality of normalized scores of each prospective point;

[0140] for each score type, calculating a score variance between the normalized scores of the plurality of prospective points corresponding to the score type to obtain a score variance of the score type;

[0141] obtaining a weight of an exploration score corresponding to each score type according to the score variances of the plurality of score types.

[0142] Optionally, the prospective point evaluation module 12 is further configured to:

[0143] summing the score variances of the plurality of score types to obtain a sum of variances;

[0144] determining a ratio between the score variance of each score type and the sum of variances as the weight of the exploration score corresponding to each score type.

[0145] Optionally, the forward-looking point evaluation module 12 is further specifically configured to:

[0146] For each score type, the plurality of forward-looking points are sorted according to the size of the normalized score corresponding to the score type, to obtain a sequence position corresponding to each forward-looking point and the score type;

[0147] According to the sequence position corresponding to each forward-looking point and the score type, a normalized score corresponding to each forward-looking point and the score type is obtained.

[0148] In addition, each function module in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0149] It should also be understood that the above embodiments, if implemented in the form of software function modules and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0150] Therefore, the present embodiment also provides a storage medium, which is a computer readable storage medium. The storage medium stores a computer program, and the computer program is executed by a processor to implement the autonomous navigation intelligent method based on hierarchical semantic reasoning provided by the present embodiment. The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0151] The present embodiment provides an electronic device for implementing the autonomous navigation intelligent method based on hierarchical semantic reasoning. As shown in the Figure 4 The electronic device can include a processor 22 and a memory 120. Moreover, the memory 21 stores a computer program, and the processor implements the autonomous navigation intelligent method based on hierarchical semantic reasoning by reading and executing the computer program corresponding to the above embodiments in the memory 21.

[0152] Continuing to refer to Figure 4The electronic device further includes a communication unit 23. The memory 21, the processor 22, and the communication unit 23 are electrically connected to each other directly or indirectly through a system bus 24 to enable data transmission or interaction.

[0153] The memory 21 can be an information recording device based on any electronic, magnetic, optical, or other physical principle for recording execution instructions, data, and the like. In some embodiments, the memory 21 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, and the like.

[0154] In some embodiments, the volatile memory can be a random access memory (RAM); in some embodiments, the non-volatile memory can be a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, and the like; in some embodiments, the storage drive can be a disk drive, a solid state drive, any type of storage disk (such as an optical disk, a DVD, and the like), or a similar storage medium, or a combination thereof, and the like.

[0155] The communication unit 23 is configured to transmit and receive data via a network. In some embodiments, the network can include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Network (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, the network can include one or more network access points. For example, the network can include wired or wireless network access points, such as base stations and / or network switching nodes, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.

[0156] The processor 22 can be an integrated circuit chip that has the ability to process signals and can include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor can include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC), or a microprocessor, etc., or any combination thereof.

[0157] It can be understood that, Figure 4The illustrated structure is merely schematic. The electronic device 100 can also have more or fewer components than shown, or a different configuration of components than shown. Figure 4 The illustrated components can be implemented in hardware, software, or a combination thereof. Figure 4 The illustrated components can be implemented in hardware, software, or a combination thereof. Figure 4 The illustrated components can be implemented in hardware, software, or a combination thereof.

[0158] It should be understood that the apparatus and method disclosed in the above embodiments can also be implemented in other manners. The above described apparatus embodiments are merely exemplary. For example, the flowcharts and block diagrams in the accompanying drawings show the possible implementation modes of the apparatus, method and computer program product according to the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementation modes, the functions noted in the blocks can occur in different orders from those noted in the accompanying drawings. For example, two consecutive blocks can actually be executed in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for implementing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0159] The above describes only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An autonomous navigation intelligent method based on hierarchical semantic reasoning, characterized in that, The method includes: Multiple prospective points to be explored were identified from the exploration area; For each lookout point, multiple exploration scores are obtained for the lookout point, wherein the multiple exploration scores include room-level co-occurrence scores, object-level co-occurrence scores, and spatial exploration scores. The room-level co-occurrence score represents the probability that all neighboring objects and the target object within a preset first range from the lookout point appear simultaneously in multiple candidate rooms. The object-level co-occurrence score represents the probability that the target object and all neighboring objects appear simultaneously. The spatial exploration score represents the value of the exploration device traveling from its current location to the lookout point for exploration based solely on spatial factors. The comprehensive score of the prospective point is obtained based on the weights of the multiple exploration scores. The lookout point with the highest overall score is determined as the optimal lookout point for the exploration device to continue its exploration.

2. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that, Obtaining the room-level co-occurrence score of the lookahead points includes: Obtain the first room co-occurrence score between the target object and the multiple candidate rooms; Obtain the second room co-occurrence score between each of the neighboring objects and the multiple candidate rooms; The second room co-occurrence scores of all the neighboring objects are aggregated to obtain the aggregated room scores of all the neighboring objects; The similarity between the first room co-occurrence score and the aggregated room score is determined as the room-level co-occurrence score of the lookahead.

3. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that, Obtain the object-level co-occurrence score of the lookahead point, including: Obtain the object co-occurrence score between the target object and each of the neighboring objects; The object co-occurrence scores of all the neighboring objects are aggregated to obtain the object-level co-occurrence score of the lookahead point.

4. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that, The spatial exploration score of the prospect point is obtained, including: Based on the location of the look-ahead point, the number of neighboring look-ahead points is determined, wherein the neighboring look-ahead points represent other look-ahead points within a preset second range from the look-ahead point; Determine the distance between the lookout point and the current position of the exploration device; The quantity and the distance are normalized, and the normalized quantity and the normalized distance are aggregated to obtain the space exploration score.

5. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 1, characterized in that, Each prospective point has multiple exploration scores that correspond one-to-one with various score types. Before obtaining the comprehensive score of the prospective point based on the weights of each of the multiple exploration scores, the method further includes: The multiple exploration scores of each prospect point are normalized to obtain multiple normalized scores for each prospect. For each score type, the variance between the plurality of prospective views and the normalized score corresponding to the score type is calculated to obtain the score variance of the score type; Based on the score variances of the various score types, the weights of the exploration scores corresponding to each score type are obtained.

6. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 5, characterized in that, Based on the score variances of the various score types, the weights of the exploration scores corresponding to each score type are obtained, including: The variances of the scores for the various scoring types are summed to obtain the sum of variances. The ratio between the variance of the score for each score type and the sum of the variances is determined as the weight of the exploration score corresponding to each score type.

7. The autonomous navigation intelligent method based on hierarchical semantic reasoning according to claim 5, characterized in that, The exploration scores of each prospect point are normalized to obtain multiple normalized scores for each prospect, including: For each score type, the multiple look-ahead points are sorted according to the size of the normalized score corresponding to the score type to obtain the sequence position of each look-ahead point and the score type. Based on the sequence position corresponding to each look-ahead point and the score type, the normalized score corresponding to each look-ahead point and the score type is obtained.

8. An autonomous navigation intelligent device based on hierarchical semantic reasoning, characterized in that, The device includes: The lookout point determination module is used to identify multiple lookout points to be explored from the exploration area; The lookout point evaluation module is used to obtain multiple exploration scores for each lookout point. The multiple exploration scores include room-level co-occurrence scores, object-level co-occurrence scores, and spatial exploration scores. The room-level co-occurrence score represents the probability that all neighboring objects and the target object within a preset first range from the lookout point appear simultaneously in multiple candidate rooms. The object-level co-occurrence score represents the probability that the target object and all neighboring objects appear simultaneously. The spatial exploration score represents the value of the exploration device traveling from its current location to the lookout point for exploration based solely on spatial factors. The prospect evaluation module is also used to obtain the comprehensive score of the prospect based on the weights of the multiple exploration scores. The lookout point selection module is used to determine the lookout point with the highest comprehensive score as the best lookout point for the exploration device to continue its exploration.

9. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the autonomous navigation intelligent method based on hierarchical semantic reasoning as described in any one of claims 1-7.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the autonomous navigation intelligent method based on hierarchical semantic reasoning as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional point cloud semantic segmentation method for optimizing boundary

    CN115409989A

  • Passable region exploration method and apparatus, storage medium, and electronic apparatus

    WO2022110853A1